Quality Detection Method, Device, Medium and Electronic Device for Text Images

By detecting and adjusting the character scale of text images and using neural network to detect the quality score of text areas, the problem of inaccurate text image quality evaluation is solved, and the accuracy and robustness of detection are improved.

CN113763313BActive Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110484548.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-30
Publication Date
2025-05-30
Estimated Expiration
2041-04-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively evaluate the quality of text images, especially when text line recognition is inaccurate.

Method used

By detecting the character scale of the text image, performing appropriate enlargement or reduction processing, and using neural networks to perform feature extraction and mapping, detecting the quality score of the text area, and weighting summing is performed based on the preset weight to obtain the quality score of the text image.

Benefits of technology

The accuracy of text area detection in text images and the accuracy of quality score prediction are improved, and the robustness of the detection scheme is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113763313B_ABST
    Figure CN113763313B_ABST
Patent Text Reader

Abstract

This application belongs to the field of computer technology, and particularly relates to a method, device, computer-readable medium, and electronic device for quality detection of text images. The method includes: detecting the character scale of the text image; when the character scale of the text image is less than a first preset scale, performing magnification processing on the text image, and when the character scale of the text image is greater than a second preset scale, performing reduction processing on the text image; inputting the text image into a first neural network to detect one or more text regions in the text image; predicting the quality score of the text region to obtain the quality score of the text region; and performing weighted summation processing on the quality scores of the text regions based on preset weights to obtain the quality score of the text image. The embodiments of this application can improve the accuracy of predicting the quality scores of text regions and even text images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and particularly relates to a method, device, computer-readable medium and electronic device for quality detection of text images. Background Art

[0002] With the increasingly wide application of OCR (Optical Character Recognition) technology, the quality of text images collected has received more attention, and text image quality assessment methods have also attracted more extensive interest in the academic and industrial fields.

[0003] Most of the existing image quality assessment methods are for natural scene images and are not suitable for text image quality evaluation. Therefore, a text image quality assessment method needs to be provided.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of this application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The purpose of this application is to provide a method, device, computer-readable medium and electronic device for quality detection of text images, which can at least overcome the technical problems such as inaccurate text line recognition in related technologies to a certain extent.

[0006] Other features and advantages of this application will become apparent through the following detailed description, or will be learned in part through the practice of this application.

[0007] According to one aspect of the embodiments of this application, a method for quality detection of text images is provided, including:

[0008] Detecting the character scale of the text image, where the character scale of the text image is the average scale of the characters in the text image;

[0009] When the character scale of the text image is less than the first preset scale, performing magnification processing on the text image so that the character scale of the text image is within the preset scale range; when the character scale of the text image is greater than the second preset scale, performing reduction processing on the text image so that the character scale of the text image is within the preset scale range, where the preset scale range is a scale range greater than the first preset scale and less than the second preset scale;

[0010] Input the text image into a first neural network, and perform feature extraction and mapping processing on the text image through the first neural network to detect one or more text regions in the text image; wherein, a receptive window is configured in the first neural network, the receptive window is within a preset scale range, and the receptive window is used to move on the text image to perform feature extraction on the text image; the text region is a continuous image region in the text image composed of some or all of the characters;

[0011] Perform quality score prediction on the text regions in the text image to obtain the quality scores of the text regions;

[0012] Obtain preset weights corresponding to each of the text regions, and perform weighted summation processing on the quality scores of the text regions based on the preset weights to obtain the quality score of the text image.

[0013] According to one aspect of the embodiments of the present application, there is provided a quality detection device for a text image, the quality detection device includes:

[0014] A character scale detection module, configured to detect the character scale of the text image, and the character scale of the text image is the average scale of the characters in the text image;

[0015] A scaling module, configured to, when the character scale of the text image is less than a first preset scale, perform magnification processing on the text image so that the character scale of the text image is within the preset scale range, and when the character scale of the text image is greater than a second preset scale, perform reduction processing on the text image so that the character scale of the text image is within the preset scale range, and the preset scale range is a scale range greater than the first preset scale and less than the second preset scale;

[0016] A text region detection module, configured to input the text image into a first neural network, and perform feature extraction and mapping processing on the text image through the first neural network to detect one or more text regions in the text image; wherein, a receptive window is configured in the first neural network, the receptive window is within a preset scale range, and the receptive window is used to move on the text image to perform feature extraction on the text image; the text region is a continuous image region in the text image composed of some or all of the characters;

[0017] A quality score prediction module, configured to perform quality score prediction on the text regions in the text image to obtain the quality scores of the text regions;

[0018] A weighted summation module, configured to obtain preset weights corresponding to each of the text regions, and perform weighted summation processing on the quality scores of the text regions based on the preset weights to obtain the quality score of the text image.

[0019] In some embodiments of the present application, based on the above technical solution, the weighted summation module includes:

[0020] A word count acquisition unit, configured to obtain the word counts of each of the text regions, and perform a summation operation on the word counts of each of the text regions to obtain the total word count of the text image;

[0021] A preset weight calculation unit, configured to respectively obtain the ratios of the word counts of each of the text regions to the total word count of the text image, and use the ratios as the preset weights corresponding to each of the text regions.

[0022] In some embodiments of the present application, based on the above technical solution, the text region is a text line, and the text line is a continuous image region in the text image that contains one or more serially arranged characters. The quality score prediction module includes:

[0023] An aspect ratio detection unit, configured to detect the text lines in the text image and the aspect ratios of the text lines. The aspect ratio is the ratio of the length and height of the text line. The length of the text line is the length of the extension line along the character arrangement direction in the text line, and the height of the text line is the height perpendicular to the extension line;

[0024] A text line segmentation unit, configured to segment a text line with an aspect ratio greater than a preset value into multiple text lines with an aspect ratio less than or equal to the preset value;

[0025] A quality score prediction unit, configured to respectively perform feature extraction and quality score prediction on the multiple text lines with an aspect ratio less than or equal to the preset value to obtain the quality scores of the multiple text lines with an aspect ratio less than or equal to the preset value.

[0026] In some embodiments of the present application, based on the above technical solution, the text line segmentation unit includes:

[0027] A text line projection sub-unit, configured to project the text line with an aspect ratio greater than a preset value onto the length direction of the text line to form a one-dimensional set of projection points. Among them, the pixel points at the positions of the characters form real point projections in the length direction of the text line, and the pixel points at the positions other than the positions of the characters form virtual point projections in the length direction of the text line. The real point projections and the virtual point projections form the set of projection points;

[0028] A text line segmentation subunit, configured to obtain segmentation points on the line segments where the virtual dot projections are aggregated, and segment the text line according to the segmentation points, so as to segment a text line with an aspect ratio greater than a preset value into multiple text lines with an aspect ratio less than or equal to the preset value.

[0029] In some embodiments of the present application, based on the above technical solution, the quality score prediction module may include:

[0030] An input unit, configured to input the text region in the text image into a second neural network;

[0031] A feature extraction unit, configured to extract features from the text region through the convolutional layer of the second neural network to obtain planar features;

[0032] A dimensionality reduction processing unit, configured to perform dimensionality reduction processing on the planar features through the pooling layer of the second neural network to obtain feature vectors;

[0033] A fully connected calculation unit, configured to perform a fully connected calculation on the feature vectors through the fully connected layer of the second neural network to obtain the predicted quality score of the text region.

[0034] In some embodiments of the present application, based on the above technical solution, the quality score prediction module may further include:

[0035] A dataset acquisition unit, configured to acquire a text recognition dataset, where the text recognition dataset includes text images and the recognition accuracy rate of the text images. Among them, the recognition accuracy rate of the text image is the ratio of the number of recognizable characters in the text image to the actual number of characters, the number of recognizable characters is the number of characters that can be correctly recognized by a character recognition model in the text image, and the actual number of characters is the number of characters actually included in the text image;

[0036] A dataset annotation unit, configured to perform a proportional conversion on the recognition accuracy rate of the text image to obtain the quality score of the text image, and annotate the quality score of the text image to the text recognition dataset;

[0037] A neural network training unit, configured to input the text recognition dataset into the second neural network and train the second neural network.

[0038] In some embodiments of the present application, based on the above technical solution, the text region is a single-character region, and the single-character region is a continuous image region in the text image that contains one character.

[0039] According to one aspect of the embodiments of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for detecting the quality of a text image in the above technical solution.

[0040] According to one aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes: a processor; and a memory for storing executable instructions of the processor. Wherein, the processor is configured to execute the method for detecting the quality of a text image in the above technical solution by executing the executable instructions.

[0041] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for detecting the quality of a text image in the above technical solution.

[0042] In the technical solution provided by the embodiments of the present application, by performing magnification processing or reduction processing on a text image whose character scale exceeds a preset scale range, so that the character scale of the text image is within the preset scale range, and the receptive field of the first neural network is also within the preset scale range. Thereby, the matching degree between the receptive field of the first neural network and the character size of the text image is relatively high, which can improve the accuracy of detecting the text area in the text image, and further can improve the accuracy of predicting the quality score of the text area and even the text image, and at the same time improve the robustness of the text image quality detection solution of the present embodiment.

[0043] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 Schematically shown is an exemplary system architecture block diagram applying the technical solution of the present application.

[0046] Figure 2 Schematically shown is a flowchart of the steps of the quality detection method in some embodiments of the present application.

[0047] Figure 3 Schematically shows a comparison diagram of the detection effects of an embodiment of the present application and the related art.

[0048] Figure 4 Schematically shows a flowchart of steps for predicting the quality score of a text region in a text image in an embodiment of the present application to obtain the quality score of the text region.

[0049] Figure 5 Schematically shows a process diagram of a second neural network in an embodiment of the present application for extracting features, performing dimensionality reduction processing, and performing fully connected calculations on a text region, and obtaining the quality score of the text region.

[0050] Figure 6 Schematically shows a partial flowchart of steps before inputting a text region in a text image into a second neural network in a quality detection method of an embodiment of the present application.

[0051] Figure 7 Schematically shows a flowchart of steps for predicting the quality score of a text region in a text image in an embodiment of the present application to obtain the quality score of the text region.

[0052] Figure 8 Schematically shows a flowchart of steps for splitting a text line with an aspect ratio greater than a preset value into multiple text lines with an aspect ratio less than or equal to the preset value in an embodiment of the present application.

[0053] Figure 9 Schematically shows a flowchart of steps for obtaining a preset weight corresponding to each text region in an embodiment of the present application.

[0054] Figure 10 Schematically shows a scenario diagram for detecting text lines, the quality scores of text lines, and the number of characters in a text image according to an embodiment of the present application.

[0055] Figure 11 Schematically shows a scenario diagram for detecting text lines, the quality scores of text lines, and the aspect ratios of text lines in a text image according to an embodiment of the present application.

[0056] Figure 12 Schematically shows a structural block diagram of a quality detection device for a text image provided by an embodiment of the present application.

[0057] Figure 13 Schematically shows a computer system structural block diagram of an electronic device for implementing an embodiment of the present application. Detailed implementation manners

[0058] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0059] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this application.

[0060] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0061] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all the content and operations / steps, nor do they necessarily have to be executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0062] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a manner similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0063] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace the human eye for tasks such as object recognition, tracking, and measurement, and further performs image processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision research related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0064] Figure 1 Schematically shows an exemplary system architecture block diagram to which the technical solution of the present application is applied.

[0065] As Figure 1 shown, the system architecture 100 may include a terminal device 110, a network 120, and a server 130. The terminal device 110 may include various electronic devices such as smart phones, tablet computers, laptop computers, and desktop computers. The server 130 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The network 120 may be a communication medium of various connection types that can provide a communication link between the terminal device 110 and the server 130. For example, it may be a wired communication link or a wireless communication link.

[0066] According to implementation requirements, the system architecture in the embodiments of the present application may have any number of terminal devices, networks, and servers. For example, the server 130 may be a server group composed of multiple server devices. In addition, the technical solution provided by the embodiments of the present application may be applied to the terminal device 110, or may be applied to the server 130, or may be jointly implemented by the terminal device 110 and the server 130. The present application makes no special limitation thereto.

[0067] For example, the quality detection method for text images according to the embodiments of the present application can be implemented on the server 130. After the user uploads a text image through the terminal device 110, the server can implement the quality detection method according to the embodiments of the present application to detect the quality of the text image. Thereby, the matching degree between the receptive window of the first neural network and the character size of the text image is relatively high, which can improve the accuracy of detecting the text area in the text image, and further improve the accuracy of predicting the quality score of the text area and even the text image. At the same time, the robustness of the text image quality detection solution according to the embodiments of the present application is improved.

[0068] In addition, in the technical solution provided by the embodiments of the present application, first, the quality score of the text area in the text image is predicted to obtain the quality score of the text area, and then the preset weights corresponding to each text area are obtained, and the quality scores of the text areas are weighted and summed based on the preset weights to obtain the quality score of the text image. Thus, the problem of predicting the quality score of the text image is transformed into the problem of predicting the quality score of the text area and obtaining the weights corresponding to each text area. Therefore, compared with directly predicting the quality score of the text image, predicting the quality score of the text area first can exclude the influence of the area in the text image that does not contain characters on the quality score prediction result, and the accuracy of the quality score prediction is higher. At the same time, different text areas can have different preset weights, and weighted summing the quality scores of the text areas based on the preset weights is beneficial to further improving the robustness of the text image quality detection solution according to the embodiments of the present application.

[0069] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted device, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions in this regard.

[0070] With the widespread use of intelligent devices in our daily lives, it is often necessary to submit text images captured by mobile devices in the company's business processes, resulting in a rapid increase in the number of text images. Therefore, intelligent document recognition has become increasingly important for business process automation. However, intelligent document recognition is very sensitive to the quality of text images. Since inevitable distortions during the image capture process may result in low-quality text images, the recognition accuracy of the captured text images usually decreases, which may seriously hinder subsequent business processes. For example, in the online insurance underwriting business, if a low-quality document image submitted for a claim is not detected as soon as possible, it needs to be recaptured immediately. Because once the underwriting document image cannot be obtained due to the loss of the paper copy or the user's non-cooperation in providing it, key information may be lost in the business process. Since the quality of text images uploaded by users varies, it is necessary to evaluate the quality of such text images in advance to reject low-quality text images.

[0071] The following will make a detailed description of the quality detection method provided by this application in combination with specific implementation manners.

[0072] Figure 2 The step flow chart of the quality detection method of some embodiments of this application is schematically shown. The execution subject of this quality detection method can be a terminal device, a server, etc., and this application does not limit this. As Figure 2 shown, this quality detection method mainly can include the following steps S210 to step S250.

[0073] S210. Detect the character scale of the text image, where the character scale of the text image is the average scale of the characters in the text image;

[0074] S220. When the character scale of the text image is less than the first preset scale, perform magnification processing on the text image so that the character scale of the text image is within the preset scale range. When the character scale of the text image is greater than the second preset scale, perform reduction processing on the text image so that the character scale of the text image is within the preset scale range. The preset scale range is the scale range greater than the first preset scale and less than the second preset scale;

[0075] S230. Input the text image into the first neural network, and perform feature extraction and mapping processing on the text image through the first neural network to detect one or more text regions in the text image; wherein, a receptive window is configured in the first neural network, the receptive window is within the preset scale range, and the receptive window is used to move on the text image to perform feature extraction on the text image; the text region is a continuous image region in the text image composed of some or all of the characters;

[0076] S240. Perform quality score prediction on the text regions in the text image to obtain the quality scores of the text regions;

[0077] S250. Obtain preset weights corresponding to each text region, and perform weighted summation processing on the quality scores of the text regions based on the preset weights to obtain the quality score of the text image.

[0078] Among them, the text image is an image containing text, that is, an image containing characters. The character scale of the text image is the average scale of the characters in the text image. Specifically, the scale can be height or area, etc. In some embodiments, the average scale of the characters in the text image can be the average value of the character heights of each character in the text image; in other embodiments, the average scale of the characters in the text image can be the average value of the character areas of each character in the text image. In a specific embodiment, the MSER (Maximally Stable Extremal Regions) detector can be used as a single-character detector to detect individual characters in the text image. Then, find the average height of all characters in the text image, and according to the calculated average height, perform adaptive scaling processing on the original image as described in step S220, so that the character scale of the text image is within a preset scale range to adapt to the scale size of the receptive window of the first neural network, which can improve the accuracy of the first neural network in detecting text regions. Among them, the receptive window is the matrix for the convolutional layer in the convolutional neural network to perform convolutional processing on data, that is, the receptive field of the convolutional neural network.

[0079] In some embodiments, after inputting the text image into the first neural network, first perform feature extraction on the text image through the feature extraction unit of the first neural network, and then perform mapping processing through the mapping processing unit of the first neural network, so as to detect one or more text regions in the text image. In a specific embodiment, the first neural network can be a PSE (Progressive Scale Expansion) segmentation network, that is, PSENet. Thus, multi-directional, multi-angle, and multi-scale detection of the text image can be realized, and it can detect and distinguish large-scale characters and small-scale characters in a text image with characters having a scale difference of several times, and can also detect and distinguish text images with characters arranged obliquely, bent, or inverted. Please refer to Figure 3 , Figure 3 which schematically shows a comparison diagram of the detection effects of an embodiment of the present application and related technologies. As Figure 3 shown, the text regions of the same text image are detected using related technologies and some embodiments of the present application. In this embodiment, the text region is a text line. The text image 310 shows that related technologies can only detect a single direction (such as Figure 3As shown in [the figure], for the text lines in the horizontal direction, the vertical text lines cannot be detected, and for the characters with a relatively large scale, such as "respect teachers and value education", they cannot be completely detected, resulting in a certain edge loss. However, the text image 320 shows that the technical solution of the present application can achieve the detection of text images in multiple directions (horizontal and vertical), multiple angles (horizontal angle and vertical angle), and multiple scales (large scale and small scale), and can achieve the complete detection of text lines, which can greatly reduce the edge loss of text lines and characters during the detection process. In addition to the PSE segmentation network, the first neural network can also be other neural networks based on CNN (Convolutional Neural Network) for detecting text regions in text images, and the present application does not limit this.

[0080] It can be understood that different from scene images, text images essentially focus more on text. Therefore, in the embodiments of the present application, by performing weighted summation processing on the quality scores of text regions based on preset weights to obtain the quality score of the text image, it can more accurately reflect the quality of the text image.

[0081] Figure 4 Schematically shows the flowchart of steps for predicting the quality score of a text region in a text image in an embodiment of the present application to obtain the quality score of the text region. As Figure 4 shown, based on the above embodiments, the step of predicting the quality score of the text region in the text image in step S240 to obtain the quality score of the text region can further include the following steps S410 to S440.

[0082] S410. Input the text region in the text image into the second neural network;

[0083] S420. Extract features from the text region through the convolutional layer of the second neural network to obtain planar features;

[0084] S430. Perform dimensionality reduction processing on the planar features through the pooling layer of the second neural network to obtain feature vectors;

[0085] S440. Perform full connection calculation on the feature vectors through the fully connected layer of the second neural network to obtain the predicted quality score of the text region.

[0086] In some embodiments, the structure of the second neural network may include a convolutional layer, a pooling layer, and a fully connected layer.

[0087] Specifically, Figure 5 Schematically shows the schematic diagram of the process of the second neural network in an embodiment of the present application for feature extraction, dimensionality reduction processing, and full connection calculation of a text region, and obtaining the predicted quality score of the text region. Please refer to Figure 5, the convolutional layer of the second neural network extracts features from the text region through convolution, max pooling, and convolution in steps to obtain the first planar feature. Then, the pooling layer of the second neural network obtains the corresponding feature vector through max pooling and min pooling of the first planar feature. Next, the fully connected layer of the second neural network calculates the quality score corresponding to the text region through fully connected calculation of the feature vector.

[0088] In some embodiments, the text region can be a text line. The second neural network can be a DIQA (Deep CNN - Based Blind Image Quality Predictor) framework based on text lines. In a specific embodiment, the second neural network can be constructed based on ResNet (Residual Network). Since ResNet has excellent feature representation ability, it can extract features from text images more accurately, thereby improving the accuracy of predicting the quality score. The second neural network can also be other CNN networks such as VGG (Visual Geometry Group) network.

[0089] In some embodiments, a text line with an aspect ratio greater than a preset value can be split into multiple text lines with an aspect ratio less than or equal to the preset value. It should be noted that the size of text images is usually much larger than the image size accepted by deep convolutional neural networks. To meet the situation where the image size accepted by deep convolutional neural networks is relatively fixed, if the text image size is adjusted by undersampling before detection, characters with smaller sizes may become blurred or even unrecognizable after undersampling, ultimately resulting in inaccurate detection of text lines, and thus relatively inaccurate prediction of the quality score of the text image.

[0090] Therefore, in order to avoid the deterioration caused by undersampling of text lines, in the embodiments of the present application, the first neural network performs feature extraction and mapping processing on the text image. After detecting the text area in the text image, the text area is input into the second neural network, that is, the deep convolutional neural network, for quality score prediction. Thus, by detecting the text area in the text image and then inputting the text area into the second neural network for quality score prediction, the text image is decomposed into one or more text areas and then input into the second neural network, which can avoid the problem of excessive image size caused by directly inputting the text image into the neural network, and at the same time can avoid the image deterioration caused by resizing the text image in an undersampling manner before detection. Moreover, the input images are all text areas, and most of the information in the images of the text areas is text information, which is beneficial for the second neural network to extract and analyze features, and can improve the accuracy and robustness of the second neural network's prediction of image quality. Based on the above effects, it can be understood that the text image quality detection method of the embodiments of the present application not only has a relatively accurate effect in the quality detection of business document images such as insurance underwriting and contracts, but also can have a relatively accurate quality detection effect in natural text image quality assessment scenarios with large interference and high difficulty in recognizing the text itself, such as bank card capture images, medical record form capture images, invoice capture images, etc.

[0091] In the second neural network, the L2 loss can be used as the estimation loss to describe the difference between the predicted and the true quality. Specifically, the estimation loss L2 is defined as:

[0092]

[0093] where q is the predicted quality score, and q gt is the quality score annotated for the text image in the text recognition dataset. When the conversion ratio of the equivalent ratio conversion is 1, q gt is also the accuracy of the text recognition of this text image.

[0094] Before training the second neural network, the parameters of the fully connected layer can be randomly initialized under a uniform normal distribution within the range of (-0.1, 0.1).

[0095] Figure 6 Schematically shows a partial step flowchart of the quality detection method of an embodiment of the present application before inputting the text area in the text image into the second neural network. As Figure 6 shown, on the basis of the above embodiment, before inputting the text area in the text image in step S410 into the second neural network, the following steps S610 to step S630 can be further included.

[0096] S610. Obtain a text recognition dataset, which includes text images and the recognition accuracy rates of the text images. Herein, the recognition accuracy rate of a text image is the ratio of the number of recognizable characters in the text image to the actual number of characters. The number of recognizable characters is the number of characters that can be correctly recognized by a character recognition model from the text image, and the actual number of characters is the number of characters actually included in the text image;

[0097] S620. Perform a geometric conversion on the recognition accuracy rate of the text image to obtain the quality score of the text image, and label the quality score of the text image in the text recognition dataset;

[0098] S630. Input the text recognition dataset into a second neural network and train the second neural network.

[0099] Specifically, performing a geometric conversion on the recognition accuracy rate of the text image to obtain the quality score of the text image can be to geometrically convert the recognition accuracy rate of the text image into a quality score on a percentage scale, and label the quality score of the text image in the text recognition dataset. The text recognition dataset is used to train the second neural network to improve the accuracy of the second neural network in predicting the quality score.

[0100] Among them, the conversion ratio for the geometric conversion can be 0.1, 1, 10, 100, etc., and this application does not limit this.

[0101] In some embodiments, the text images in the text recognition dataset can be text line images, that is, images containing one or more serially arranged characters. Herein, the serially arranged characters are also characters arranged in series one by one. Specifically, the serially arranged characters can be single-line characters or single-column characters. Thus, it can be understood that training the second neural network based on text line images can reduce the background clutter and noise of the images, which is conducive to reducing the influence of the background clutter and noise of the images on the quality score, and thus can improve the accuracy of quality score prediction. And, since in some embodiments of this application, the text area can be a text line, at this time, first detect the text line in the text image, and then predict the quality score of the text line in the text image according to the second neural network; in this case, if training the second neural network based on single-line text images, it can make the attributes of the input images in the training process of the second neural network close to the attributes of the input images when using the second neural network to predict the quality score, both being images containing one or more serially arranged characters, which is conducive to improving the prediction accuracy of using the second neural network to predict the quality score of the text area.

[0102] In some embodiments, the text images in the text recognition dataset can come from artificial images synthesized by algorithms such as fuzzy noise, and the quality scores of the text images in the text recognition dataset can be automatically generated during the fuzzy noise synthesis process.

[0103] In some other embodiments, the text images in the text recognition dataset can be from real text images, and the quality scores of the text images in the text recognition dataset can be from manual scoring.

[0104] In still some other embodiments, the text images in the text recognition dataset can be from the recognition data of real text images by an OCR (Optical Character Recognition) model in a text recognition task. In this embodiment, the recognition accuracy rate of the text image is converted by a ratio to obtain the quality score of the text image, and the quality score of the text image is labeled in the text recognition dataset. Compared with the quality scores manually given, it has the advantage of being more objective. Moreover, since the accuracy rate is a continuous real number instead of the discrete quality scores of manual scoring, it can optimize the training of the parameter values of the second neural network and help the training task converge better. Furthermore, it is difficult to subjectively score text images, and the dataset of the quality scores of images is relatively lacking, while the dataset of the recognition accuracy rate of text images in the text recognition task is already very common. This solution uses the dataset of the recognition accuracy rate of text images in the text recognition task and converts the quality score of the text image according to the recognition accuracy rate of the text image. At the same time, real text images are more in line with the actual scenario than artificial images synthesized by algorithms such as fuzzy noise, which is beneficial to improving the training effect of the parameters of the second neural network, and thus beneficial to improving the prediction accuracy rate of the quality score of the text area using the second neural network.

[0105] Figure 7 Schematically shown is a flowchart of the steps for predicting the quality score of a text area in a text image in an embodiment of the present application to obtain the quality score of the text area. As Figure 7 shown, on the basis of the above embodiments, the text area is a text line, and the text line is a continuous image area in the text image that contains one or more serially arranged characters. The step of predicting the quality score of the text area in the text image in step S240 to obtain the quality score of the text area can further include the following steps S710 to S730.

[0106] S710. Detect the text lines in the text image and the aspect ratio of the text lines. The aspect ratio is the ratio of the length to the height of the text line. The length of the text line is the length of the extension line along the direction in which the characters in the text line are arranged, and the height of the text line is the height perpendicular to the extension line.

[0107] S720. Divide the text lines with an aspect ratio greater than a preset value into multiple text lines with an aspect ratio less than or equal to the preset value.

[0108] S730. Feature extraction and quality score prediction are respectively performed on multiple text lines with an aspect ratio less than or equal to a preset value, and the quality scores of the multiple text lines with an aspect ratio less than or equal to the preset value are obtained.

[0109] In some embodiments of the present application, a text line with an aspect ratio greater than a preset value is segmented into multiple text lines with an aspect ratio less than or equal to the preset value, which can segment a text line with a longer length and more text characters into text lines with a shorter length and an aspect ratio less than or equal to the preset value. Thus, the long text is adaptively segmented, avoiding the input of overly long text lines into the second neural network, and greatly improving the prediction accuracy of the second neural network.

[0110] In some embodiments, the text image in the text recognition dataset can be a text line image, that is, an image containing one or more serially arranged characters. More specifically, the aspect ratio of the text line image can be less than or equal to the preset value. Thus, during the training stage of the second neural network using the text recognition dataset and the stage of predicting the quality score of the text line, the aspect ratio of the processed images is less than or equal to the preset value, which is beneficial to the optimization of the second neural network, can reduce the computational amount of the second neural network, and improve the prediction efficiency of predicting the quality score of the text line.

[0111] Figure 8 Schematically shows the step flowchart of segmenting a text line with an aspect ratio greater than a preset value into multiple text lines with an aspect ratio less than or equal to the preset value in an embodiment of the present application. As Figure 8 shown, based on the above embodiments, the step of segmenting a text line with an aspect ratio greater than a preset value into multiple text lines with an aspect ratio less than or equal to the preset value in step S720 can further include the following steps S810 to step S820.

[0112] S810. Project a text line with an aspect ratio greater than a preset value onto the length direction of the text line to form a one-dimensional set of projection points. Among them, the pixel points at the positions of the characters form real point projections in the length direction of the text line, and the pixel points at the positions other than the positions of the characters form virtual point projections in the length direction of the text line. The real point projections and the virtual point projections form the set of projection points;

[0113] S820. Take segmentation points on the line segments where the virtual point projections gather, and segment the text line according to the segmentation points to segment a text line with an aspect ratio greater than a preset value into multiple text lines with an aspect ratio less than or equal to the preset value.

[0114] Specifically, the real point projection can be a black dot pixel, the virtual point projection can be a white dot pixel, and the one-dimensional set of projection points can be a line segment composed of black dot pixels, a line segment composed of white dot pixels, or a line segment composed of black dot pixels and white dot pixels. By taking a segmentation point on the line segment where the virtual point projections gather and segmenting the text line according to the segmentation point, it can be understood that the line segment where the virtual point projections gather is the projection line segment of some pixel points at positions other than the positions where the characters are located in the length direction of the text line. By taking a segmentation point on the line segment where the virtual point projections gather, it is possible to segment the text line at positions other than the positions where the characters are located, thereby avoiding splitting a single character into two different text lines, which may cause difficulties in character recognition and reduce the accuracy of quality score prediction.

[0115] Figure 9 Schematically shows a flowchart of steps for obtaining preset weights corresponding to each text region in an embodiment of the present application. As Figure 9 shown, based on the above embodiments, obtaining the preset weights corresponding to each text region in step S250 may further include the following steps S910 to step S920.

[0116] S910. Obtain the number of characters in each text region, and perform a summation operation on the number of characters in each text region to obtain the total number of characters in the text image;

[0117] S920. Respectively obtain the ratio of the number of characters in each text region to the total number of characters in the text image, and use the ratio as the preset weight corresponding to each text region.

[0118] It can be understood that the longer text lines in the text image occupy more content, and the quality of the longer text lines has a greater impact on the quality of the text image. Therefore, by respectively obtaining the ratio of the number of characters in each text region to the total number of characters in the text image and using the ratio as the preset weight corresponding to each text region, it is possible to make the preset weight of the longer text lines larger, so that the text image quality detection method of the embodiment of the present application conforms to the determination logic of actual quality determination, can improve the accuracy of text image quality detection, and improve the detection efficiency of the quality detection method.

[0119] In a specific embodiment, the quality score of the text image can be:

[0120]

[0121] where q j is the predicted quality score of the j-th text line, and w j is the preset weight of the j-th text line in the text image. In some embodiments, can be equal to the sum value after multiplying the predicted quality scores of all text lines by their corresponding preset weights. wj The definition can be as follows:

[0122]

[0123] Wherein, R (j) is the number of characters in the j-th text line, and ∑ k R (k) means that there are k text lines in the text image, and the sum value of the number of characters in the k text lines. In some embodiments, the number of characters in a text line can be detected by a single-character detector, such as an MSER detector. In some other embodiments, the number of characters R (k) in a text line is approximately obtained from the aspect ratio of the length to the height of the text line. For example:

[0124]

[0125] Wherein, line_w is the length of the text line, and line_h is the height of the text line. The length of the text line is the length of the extension line along the arrangement direction of the characters in the text line, and the height of the text line is the height perpendicular to the extension line. Thus, by approximately calculating the number of characters in a text line using the aspect ratio of the length to the height of the text line, the detection process of the quality detection method in the embodiments of the present application can be simplified, and the detection amount can be reduced.

[0126] Alternatively, in some embodiments, obtaining the preset weights corresponding to each text region includes: obtaining the aspect ratio of the length to the height of each text region, and performing a summation operation on the aspect ratios of the length to the height of each text region to obtain the total aspect ratio; obtaining the ratio of the aspect ratio of each text region to the total aspect ratio, and using the ratio as the preset weight corresponding to each text region. Thus, by directly calculating the preset weights corresponding to each text region using the aspect ratio of the length to the height of the text line, the calculation amount can be reduced, and the detection efficiency of the quality detection method can be improved.

[0127] Figure 10 Schematically shows a scene diagram for detecting text lines, quality scores of text lines, and the number of characters in a text image according to an embodiment of the present application. Please refer to Figure 10 , perform text line detection on the text image, and then perform quality score prediction and number of characters detection on the text lines in the text image. The results of the quality score prediction and number of characters detection of the text lines are displayed below the corresponding text lines. For example, below the text line "Ms.", it is shown that the number of characters in the text line "Ms." is 2, and the predicted quality score is 0.8. At this time, the predicted quality score of the text image can be calculated according to formula (1) of the quality score of the text image and the results of the quality score prediction and number of characters detection of the text lines.

[0128] Figure 11Schematically shown is a schematic diagram of a scenario for detecting text lines in a text image, as well as the quality scores and aspect ratios of the text lines, according to an embodiment of the present application. Please refer to Figure 11 , perform text line detection on the text image, and then perform quality score prediction and aspect ratio detection on the text lines in the text image. The results of the quality score prediction and aspect ratio detection of the text lines are displayed below the corresponding text lines. For example, the aspect ratio of the text line "Ultrasound description:" is 4.4, and the predicted quality score is 0.848. At this time, the predicted quality score of the text image can be calculated according to formula (1) of the quality score of the text image, the quality scores, and aspect ratios of the text lines.

[0129] In some embodiments, the text region is a single-character region, and the single-character region is a continuous image region in the text image that contains one character. Then, perform quality score prediction on the single-character regions in the text image to obtain the quality scores of each single-character region, and then obtain the preset weights corresponding to each single-character region, and perform weighted summation processing on the quality scores of the single-character regions based on the preset weights to obtain the quality score of the text image. Among them, the preset weights of the single-character regions are all 1. Alternatively, the preset weights of the single-character regions can be determined by the ratio of the area of the single-character region to the total area of the single-character regions. For example:

[0130] The quality score of the text image can be:

[0131]

[0132] where q j is the predicted quality score of the jth text line, and v j is the preset weight of the jth single-character region in the text image. In some embodiments, can be equal to the sum of the products of the predicted quality scores of all single-character regions and the corresponding preset weights. The definition of v j can be as follows:

[0133]

[0134] where S (j) is the area size of the jth single-character region, and ∑ k S (k)There are k single-character regions in the text image, and the sum of the areas of the k single-character regions. Thus, it can be understood that the larger the text in the text image, the higher its importance, and the larger the text, the greater the impact on the quality of the text image. Therefore, obtaining the preset weights corresponding to each single-character region and performing weighted summation processing on the quality scores of the single-character regions based on the preset weights to obtain the quality score of the text image can make the preset weights of the larger single-character regions larger, making the text image quality detection method of the embodiment of the present application conform to the determination logic of actual quality determination and improving the accuracy of text image quality detection.

[0135] In some embodiments, the first neural network and the second neural network can be combined into the same deep neural network, and the eigenvalue obtained by the first neural network through feature extraction of the text image is shared with the second neural network. Specifically, the text image quality detection method may further include:

[0136] Input the eigenvalue obtained by the first neural network through feature extraction of the text image into the second neural network, where the eigenvalue is used to help the second neural network perform feature extraction on the text region through the convolutional layer to obtain one or more of planar features, feature vectors, and predicted quality scores.

[0137] Thus, the computational complexity of the text image quality detection method of the embodiment of the present application can be reduced, and the efficiency of text image quality score detection can be improved.

[0138] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that these steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0139] The following introduces the device embodiments of the present application, which can be used to execute the text image quality detection method in the above embodiments of the present application. Figure 12 Schematically shows the structural block diagram of the text image quality detection device provided by the embodiment of the present application. As Figure 12 shown, the text image quality detection device 1200 includes:

[0140] A character scale detection module 1210, configured to detect the character scale of the text image, where the character scale of the text image is the average scale of the characters in the text image;

[0141] A scaling module 1220, configured to, when the character scale of a text image is less than a first preset scale, perform magnification processing on the text image so that the character scale of the text image is within a preset scale range, and when the character scale of the text image is greater than a second preset scale, perform reduction processing on the text image so that the character scale of the text image is within the preset scale range, where the preset scale range is a scale range greater than the first preset scale and less than the second preset scale;

[0142] A text region detection module 1230, configured to input the text image into a first neural network, perform feature extraction and mapping processing on the text image through the first neural network to detect a text region in the text image; wherein, a receptive window is configured in the first neural network, the receptive window is within the preset scale range, and the receptive window is used to move on the text image to perform feature extraction on the text image; the text region is a continuous image region in the text image composed of some or all characters;

[0143] A quality score prediction module 1240, configured to perform quality score prediction on the text region in the text image to obtain the quality score of the text region;

[0144] A weighted summation module 1250, configured to obtain preset weights corresponding to each text region, and perform weighted summation processing on the quality scores of the text regions based on the preset weights to obtain the quality score of the text image.

[0145] In some embodiments of the present application, based on the above embodiments, the weighted summation module includes:

[0146] A word count acquisition unit, configured to acquire the word counts of each text region, and perform summation operation on the word counts of each text region to obtain the total word count of the text image;

[0147] A preset weight calculation unit, configured to respectively obtain the ratio of the word count of each text region to the total word count of the text image, and use the ratio as the preset weight corresponding to each text region.

[0148] In some embodiments of the present application, based on the above embodiments, the text region is a text line, the text line is a continuous image region in the text image containing one or more serially arranged characters, and the quality score prediction module includes:

[0149] An aspect ratio detection unit, configured to detect the text line and the aspect ratio of the text line in the text image, the aspect ratio is the ratio of the length and height of the text line, the length of the text line is the length of the extension line along the character arrangement direction in the text line, and the height of the text line is the height perpendicular to the extension line of the text line;

[0150] A text line segmentation unit, configured to segment a text line with an aspect ratio greater than a preset value into multiple text lines with an aspect ratio less than or equal to the preset value;

[0151] A quality score prediction unit, configured to perform feature extraction and quality score prediction on multiple text lines with an aspect ratio less than or equal to the preset value respectively, to obtain the quality scores of the multiple text lines with an aspect ratio less than or equal to the preset value.

[0152] In some embodiments of the present application, based on the above embodiments, the text line segmentation unit includes:

[0153] A text line projection sub-unit, configured to project a text line with an aspect ratio greater than a preset value onto the length direction of the text line to form a one-dimensional set of projection points. Among them, the pixel points at the positions where the characters are located form real point projections in the length direction of the text line, and the pixel points at the positions other than the positions where the characters are located form virtual point projections in the length direction of the text line. The real point projections and the virtual point projections form the set of projection points;

[0154] A text line segmentation sub-unit, configured to take segmentation points on the line segments where the virtual point projections gather, and segment the text line according to the segmentation points, so as to segment a text line with an aspect ratio greater than a preset value into multiple text lines with an aspect ratio less than or equal to the preset value.

[0155] In some embodiments of the present application, based on the above embodiments, the quality score prediction module may include:

[0156] An input unit, configured to input the text area in the text image into a second neural network;

[0157] A feature extraction unit, configured to perform feature extraction on the text area through the convolutional layer of the second neural network to obtain planar features;

[0158] A dimensionality reduction processing unit, configured to perform dimensionality reduction processing on the planar features through the pooling layer of the second neural network to obtain feature vectors;

[0159] A fully connected calculation unit, configured to perform a fully connected calculation on the feature vectors through the fully connected layer of the second neural network to obtain the predicted quality score of the text area.

[0160] In some embodiments of the present application, based on the above embodiments, the quality score prediction module may further include:

[0161] A dataset acquisition unit, configured to acquire a text recognition dataset, where the text recognition dataset includes text images and the recognition accuracy rate of the text images. The recognition accuracy rate of the text images is the ratio of the number of recognizable characters in the text images to the actual number of characters. The number of recognizable characters is the number of characters that can be correctly recognized by a character recognition model in the text images, and the actual number of characters is the number of characters actually included in the text images;

[0162] A dataset annotation unit, configured to perform a geometric conversion on the recognition accuracy rate of the text images to obtain the quality scores of the text images, and annotate the quality scores of the text images to the text recognition dataset;

[0163] A neural network training unit, configured to input the text recognition dataset into a second neural network and train the second neural network.

[0164] In some embodiments of the present application, based on the above embodiments, the text area is a single-character area, and the single-character area is a continuous image area in the text image that contains one character.

[0165] The specific details of the text image quality detection device provided in the embodiments of the present application have been described in detail in the corresponding method embodiments, and will not be elaborated here.

[0166] Figure 13 Schematically shows a computer system block diagram of an electronic device for implementing the embodiments of the present application.

[0167] It should be noted that Figure 13 The computer system 1300 of the shown electronic device is only an example, and should not bring any limitations to the functions and usage scopes of the embodiments of the present application.

[0168] As Figure 13 shown, the computer system 1300 includes a central processing unit 1301 (Central Processing Unit, CPU), which can execute various appropriate actions and processes according to the programs stored in the read-only memory 1302 (Read-Only Memory, ROM) or the programs loaded from the storage section 1308 into the random access memory 1303 (Random Access Memory, RAM). In the random access memory 1303, various programs and data required for system operations are also stored. The central processing unit 1301, the read-only memory 1302, and the random access memory 1303 are connected to each other through a bus 1304. The input / output interface 1305 (Input / Output interface, that is, I / O interface) is also connected to the bus 1304.

[0169] The following components are connected to the input / output interface 1305: an input section 1306 including a keyboard, a mouse, etc.; an output section 1307 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a local area network card, a modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the input / output interface 1305 as required. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1310 as required so that a computer program read from the removable medium 1311 is installed into the storage section 1308 as required.

[0170] Specifically, according to an embodiment of the present application, the processes described in each method flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1309, and / or installed from the removable medium 1311. When the computer program is executed by the central processing unit 1301, various functions defined in the system of the present application are executed.

[0171] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0172] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0173] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0174] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0175] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.

[0176] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for quality detection of text images, characterized in that, it includes: detecting the character scale of the text image, where the character scale of the text image is the average scale of the characters in the text image; when the character scale of the text image is less than a first preset scale, performing magnification processing on the text image so that the character scale of the text image is within a preset scale range, and when the character scale of the text image is greater than a second preset scale, performing reduction processing on the text image so that the character scale of the text image is within a preset scale range, where the preset scale range is a scale range greater than the first preset scale and less than the second preset scale; inputting the text image into a first neural network, and performing feature extraction and mapping processing on the text image through the first neural network to detect one or more text regions in the text image; wherein, a receptive window is configured in the first neural network, the receptive window is within the preset scale range, and the receptive window is used to move on the text image to perform feature extraction on the text image; the text region is a continuous image region in the text image composed of some or all of the characters; performing quality score prediction on the text regions in the text image to obtain the quality scores of the text regions; obtaining the number of characters in each of the text regions, and performing a summation operation on the number of characters in each of the text regions to obtain the total number of characters in the text image; respectively obtaining the ratio of the number of characters in each of the text regions to the total number of characters in the text image, using the ratio as the preset weight corresponding to each of the text regions, and performing a weighted summation process on the quality scores of the text regions based on the preset weight to obtain the quality score of the text image.

2. The method for quality detection of text images according to claim 1, characterized in that, the text region is a text line, and the text line is a continuous image region in the text image containing one or more serially arranged characters. The performing quality score prediction on the text regions in the text image to obtain the quality scores of the text regions includes: detecting the text lines in the text image and the aspect ratio of the text lines, where the aspect ratio of the text lines is the ratio of the length and height of the text lines, the length of the text line is the length of the extension line along the character arrangement direction in the text line, and the height of the text line is the height perpendicular to the extension line of the text line; dividing the text lines with an aspect ratio greater than a preset value into multiple text lines with an aspect ratio less than or equal to the preset value; respectively performing feature extraction and quality score prediction on the multiple text lines with an aspect ratio less than or equal to the preset value to obtain the quality scores of the multiple text lines with an aspect ratio less than or equal to the preset value.

3. The method for quality detection of text images according to claim 2, characterized in that, the dividing the text lines with an aspect ratio greater than a preset value into multiple text lines with an aspect ratio less than or equal to the preset value includes: Project the text line with an aspect ratio greater than a preset value onto the length direction of the text line to form a one-dimensional set of projection points. Among them, the pixel points at the positions of the characters form real-point projections in the length direction of the text line, and the pixel points at the positions other than the positions of the characters form virtual-point projections in the length direction of the text line. The real-point projections and the virtual-point projections form the set of projection points; Take segmentation points on the line segments where the virtual-point projections gather, and segment the text line according to the segmentation points to segment the text line with an aspect ratio greater than the preset value into multiple text lines with an aspect ratio less than or equal to the preset value.

4. The method for quality detection of a text image according to claim 1, characterized in that, the predicting the quality score of the text area in the text image to obtain the quality score of the text area includes: inputting the text area in the text image into a second neural network; extracting features of the text area through the convolutional layer of the second neural network to obtain planar features; performing dimensionality reduction processing on the planar features through the pooling layer of the second neural network to obtain feature vectors; performing a fully connected calculation on the feature vectors through the fully connected layer of the second neural network to obtain the predicted quality score of the text area.

5. The method for quality detection of a text image according to claim 4, characterized in that, before inputting the text area in the text image into the second neural network, the method further includes: obtaining a text recognition data set, where the text recognition data set includes text images and the recognition accuracy of the text images. Among them, the recognition accuracy of the text image is the ratio of the number of recognizable characters in the text image to the actual number of characters. The number of recognizable characters is the number of characters that can be correctly recognized by a character recognition model in the text image, and the actual number of characters is the actual number of characters included in the text image; performing a geometric conversion on the recognition accuracy of the text image to obtain the quality score of the text image, and labeling the quality score of the text image in the text recognition data set; inputting the text recognition data set into the second neural network to train the second neural network.

6. The method for quality detection of a text image according to claim 1, characterized in that, the text area is a single-character area, and the single-character area is a continuous image area in the text image that contains one character.

7. A device for quality detection of a text image, characterized in that, comprising: a character scale detection module configured to detect the character scale of the text image, where the character scale of the text image is the average scale of the characters in the text image; A scaling module, configured to, when the character scale of the text image is smaller than a first preset scale, perform magnification processing on the text image so that the character scale of the text image falls within a preset scale range, and when the character scale of the text image is larger than a second preset scale, perform reduction processing on the text image so that the character scale of the text image falls within the preset scale range, where the preset scale range is a scale range greater than the first preset scale and less than the second preset scale; A text region detection module, configured to input the text image into a first neural network, perform feature extraction and mapping processing on the text image through the first neural network to detect a text region in the text image; wherein, a receptive window is configured in the first neural network, the receptive window falls within the preset scale range, and the receptive window is used to move on the text image to perform feature extraction on the text image; the text region is a continuous image region composed of some or all characters in the text image; A quality score prediction module, configured to perform quality score prediction on the text region in the text image to obtain the quality score of the text region; A weighted summation module, configured to obtain a preset weight corresponding to each text region, and perform weighted summation processing on the quality scores of the text regions based on the preset weight to obtain the quality score of the text image; The weighted summation module includes: A word count acquisition unit, configured to acquire the word count of each text region, and perform a summation operation on the word counts of each text region to obtain the total word count of the text image; A preset weight calculation unit, configured to respectively acquire the ratio of the word count of each text region to the total word count of the text image, and use the ratio as the preset weight corresponding to each text region.

8. A computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the quality detection method of the text image according to any one of claims 1 to 6 is implemented.

9. An electronic device, characterized in that, it includes: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the quality detection method of the text image according to any one of claims 1 to 6 by executing the executable instructions.

10. A computer program product, characterized in that, the computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the quality detection method of the text image according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image recognition method and device, computer equipment and storage medium

    CN111914834A

  • OCV detection method for code spraying characters

    CN111967456A

  • Identifying Matching Canonical Documents in Response to a Visual Query

    US20110129153A1