Image text recognition method and device, electronic equipment and program product

By constructing and optimizing the text detection model, the problem of low recognition accuracy caused by image text tilt and distortion is solved, and more efficient text recognition and business operations are achieved.

CN120472468APending Publication Date: 2025-08-12INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510558180.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the prior art, image text is prone to tilt and distortion, resulting in a low accuracy of text recognition, especially in the counter business of financial institutions, which affects business processing efficiency.

Method used

By building an initial text detection model, using the backbone network and feature pyramid module for feature extraction and upsampling, combining data augmentation and model cropping technology, the text detection model is optimized, and text position information is acquired and cropped, corrected and identified.

Benefits of technology

It improves the accuracy of image text recognition, ensures the clarity and orientation of text, and improves the accuracy and efficiency of business operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472468A_ABST
    Figure CN120472468A_ABST
Patent Text Reader

Abstract

The invention discloses an image text recognition method and device, electronic equipment and a program product, and relates to the technical field of artificial intelligence, and the recognition method comprises the steps: obtaining a to-be-detected image, inputting the to-be-detected image to a preset text detection model, and obtaining a text image; processing the text image according to the text position information to obtain a processed text image; and inputting the processed text image into a preset text recognition model to obtain a text in the to-be-detected image. According to the method and the device, the technical problem that the text recognition accuracy is relatively low due to the fact that the image text is easy to incline and distort in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an image text recognition method and device, electronic equipment, and program product. Background Art

[0002] In the counter business of financial institutions, it is necessary to manually enter the text that has been printed on the voucher into the system. For example, the business personnel need to manually enter the customer code, customer name, payment account number and other information in the entrustment agreement into the system. However, the font size of such fields in the entrustment agreement is small, and the business personnel are prone to make mistakes when entering these fields. If the input errors are not discovered in time, remote authorization will result in remote rejection or remote return, which reduces the business personnel's business processing efficiency.

[0003] Furthermore, when recognizing printed text on a voucher, current OCR (Optical Character Recognition) models are unable to accurately recognize the printed text on the voucher due to problems such as uneven paper and low image clarity of the images captured by scanners and high-definition cameras.

[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0005] The embodiments of the present invention provide a method for recognizing image text, an apparatus thereof, an electronic device, and a program product, so as to at least solve the technical problem in the related art that image text is prone to tilt and distortion, resulting in low text recognition accuracy.

[0006] According to one aspect of an embodiment of the present application, a method for recognizing image text is provided, including: obtaining an image to be detected, and inputting the image to be detected into a preset text detection model to obtain a text image, wherein the text image corresponds to text position information; processing the text image based on the text position information to obtain a processed text image; inputting the processed text image into a preset text recognition model to obtain the text in the image to be detected.

[0007] Furthermore, before inputting the image to be detected into the preset text detection model to obtain the text image, it also includes: constructing an initial text detection model; obtaining multiple historical text images, and training the initial text detection model based on all historical text images to obtain a preset text detection model.

[0008] Furthermore, the step of constructing an initial text detection model includes: constructing a backbone network module, wherein the backbone network module is used to extract features of historical text images to obtain feature maps, and the backbone network module includes: a multi-layer neural network, each layer of the neural network is used to output feature maps of different resolutions; constructing a feature pyramid module, wherein the feature pyramid module is used to upsample the feature maps of different resolutions output by each layer of the neural network; adjusting the feature pyramid module using a preset strategy to obtain an adjusted feature pyramid module; and constructing an initial text detection model based on the backbone network module and the adjusted feature pyramid module.

[0009] Furthermore, based on all historical text images, the initial text detection model is trained to obtain a preset text detection model, including: performing image transformation on each historical text image to obtain a transformed historical text image; adding all historical text images and all transformed historical text images to an image training set; based on the image training set, the initial text detection model is trained to obtain a preset text detection model.

[0010] Furthermore, the preset text detection model includes: multiple filters, each filter corresponds to a weight matrix. After the initial text detection model is trained based on all historical text images to obtain the preset text detection model, it also includes: mapping the weight matrix of each filter to obtain a weight vector, and using the weight vector as the spatial point coordinates of the filter; determining the center point coordinates based on all spatial point coordinates; determining the distance between each spatial point coordinate and the center point coordinate, and sorting all distances to obtain a sorted set; using the spatial point coordinates corresponding to the minimum distance in the sorted set as the target spatial point coordinates, and cutting out the filter corresponding to the target spatial point coordinates.

[0011] Furthermore, the step of processing the text image based on the text position information to obtain a processed text image includes: cropping the text image based on the text position information to obtain a text area image; adjusting the resolution of the text area image to obtain an adjusted text area image; and determining the processed text image based on the adjusted text area image.

[0012] Furthermore, the step of determining the processed text image based on the adjusted text area image includes: using a direction classifier to analyze the adjusted text area image to obtain rotation information of the adjusted text area image; based on the rotation information, determining whether the adjusted text area image needs to be corrected; if the adjusted text area image needs to be corrected, correcting the adjusted text image to obtain a processed text image.

[0013] According to another aspect of an embodiment of the present application, a device for recognizing image text is also provided, including: a first input unit, used to obtain an image to be detected, and input the image to be detected into a preset text detection model to obtain a text image, wherein the text image corresponds to text position information; a first processing unit, used to process the text image based on the text position information to obtain a processed text image; a second input screening unit, used to input the processed text image into a preset text recognition model to obtain the text in the image to be detected.

[0014] Furthermore, the image text recognition device includes: a first construction module, used to construct an initial text detection model before inputting the image to be detected into a preset text detection model to obtain a text image; a first training module, used to obtain multiple historical text images, and based on all historical text images, train the initial text detection model to obtain a preset text detection model.

[0015] Furthermore, the first construction module includes: a first construction sub-module, used to construct a backbone network module, wherein the backbone network module is used to extract features of historical text images to obtain feature maps, and the backbone network module includes: a multi-layer neural network, each layer of the neural network is used to output feature maps of different resolutions; a second construction sub-module, used to construct a feature pyramid module, wherein the feature pyramid module is used to upsample the feature maps of different resolutions output by each layer of the neural network; a first adjustment sub-module, used to adjust the feature pyramid module using a preset strategy to obtain an adjusted feature pyramid module; and a third construction sub-module, used to construct an initial text detection model based on the backbone network module and the adjusted feature pyramid module.

[0016] Furthermore, the first training module includes: a first transformation submodule, used to perform image transformation on each historical text image to obtain a transformed historical text image; a first adding submodule, used to add all historical text images and all transformed historical text images to an image training set; and a first training submodule, used to train the initial text detection model based on the image training set to obtain a preset text detection model.

[0017] Furthermore, the preset text detection model includes: multiple filters, each filter corresponds to a weight matrix, and the image text recognition device includes: a first mapping module, which is used to train the initial text detection model based on all historical text images to obtain the preset text detection model, and then map the weight matrix of each filter to obtain a weight vector, and use the weight vector as the spatial point coordinate of the filter; a first determination module, which is used to determine the center point coordinate based on all spatial point coordinates; a second determination module, which is used to determine the distance between each spatial point coordinate and the center point coordinate, and sort all distances to obtain a sorted set; a first clipping module, which is used to use the spatial point coordinate corresponding to the minimum distance in the sorted set as the target spatial point coordinate, and clip the filter corresponding to the target spatial point coordinate.

[0018] Furthermore, the first processing unit includes: a second cropping module, used to crop the text image based on the text position information to obtain a text area image; a first adjustment module, used to adjust the resolution of the text area image to obtain an adjusted text area image; and a third determination module, used to determine the processed text image based on the adjusted text area image.

[0019] Furthermore, the third determination module includes: a first analysis submodule, used to analyze the adjusted text area image using a direction classifier to obtain rotation information of the adjusted text area image; a first judgment submodule, used to judge whether the adjusted text area image needs to be corrected based on the rotation information; and a first correction submodule, used to correct the adjusted text image if the adjusted text area image needs to be corrected to obtain a processed text image.

[0020] According to another aspect of an embodiment of the present application, a computer program product is also provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the above-mentioned image text recognition methods is implemented.

[0021] According to another aspect of an embodiment of the present application, an electronic device is also provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors implement any of the above-mentioned image text recognition methods.

[0022] In the present invention, an image to be detected is obtained and input into a preset text detection model to obtain a text image. The text image is processed based on text position information to obtain a processed text image. The processed text image is input into a preset text recognition model to obtain the text in the image to be detected, thereby solving the technical problem in the related art that the text in the image is prone to tilt and distortion, resulting in low text recognition accuracy.

[0023] In the present invention, a preset text detection model is used to perform text detection on an input image to be detected, and a text image including text position information is obtained. The text image is processed according to the text position information to obtain a processed text image, which can ensure the clarity of the text and ensure that the direction of the text is an upright rectangle. Thereafter, the processed text image is input into a preset text recognition model to obtain the text in the image to be detected, and deformed text can be accurately identified, thereby achieving the technical effect of improving the accuracy of image text recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0025] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing an image text recognition method is shown;

[0026] Figure 2 is a flowchart of the image text recognition method according to Example 1 of the present application;

[0027] Figure 3 is a schematic diagram of an optional DBNet network structure based on a recursive feature pyramid according to an embodiment of the present application;

[0028] Figure 4 is a schematic diagram of an optional image text recognition device according to an embodiment of the present application;

[0029] Figure 5 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) collected and involved in the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse. For example, an interface is set up between this system and the relevant users or institutions. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or institution through the interface, and obtain relevant information after receiving the consent information fed back by the aforementioned user or institution.

[0033] In the present invention, a text detection method is used to obtain polygon information of the text area in the image to be detected. At the same time, a recursive feature pyramid network and a model cropping technology are introduced to improve the text detection efficiency. According to the obtained polygon information, the text polygon area is cropped, perspective, transformed and corrected to obtain a processed text image. Thereafter, text recognition is performed on the processed text image to obtain a final recognition result. By optimizing the text detection and text recognition training models, the accuracy of text recognition is improved.

[0034] The present invention will be described in detail below with reference to various embodiments.

[0035] Example 1

[0036] According to an embodiment of the present application, an embodiment of a method for recognizing image text is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0037] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing an image text recognition method is shown in FIG. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more ( Figure 1 102a, 102b, ..., 102n are used to illustrate) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera, wherein the network interface may be connected to a wired and / or wireless network. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0038] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0039] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image text recognition method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizes the above-mentioned image text recognition method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0040] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0041] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0042] Under the above operating environment, this application provides Figure 2 The image text recognition method shown. Figure 2 is a flowchart of the image text recognition method according to Example 1 of the present application, such as Figure 2 As shown, the method includes the following steps:

[0043] Step S201 : obtaining an image to be detected, and inputting the image to be detected into a preset text detection model to obtain a text image, wherein the text image corresponds to text position information.

[0044] In an embodiment of the present invention, by obtaining an image to be detected (such as a scanned copy or photo of a credential of a financial institution) and inputting the image to be detected into a preset text detection model (such as an optimized DBNet (Differentiable Binarization Network, a text detection model based on deep learning) algorithm model), the text area in the image can be detected and its location information (such as the coordinates of the bounding box, the coordinates of the polygon vertices, etc.) can be output.

[0045] Step S202 : processing the text image according to the text position information to obtain a processed text image.

[0046] In embodiments of the present invention, a text image can be cropped, corrected, and optimized based on the text location information to produce a processed text image. For example, the text region can be separated from the image background to generate an isolated text image, reducing noise interference. For non-rectangular text regions, perspective transformation can be applied to correct them into rectangles for easier recognition.

[0047] Step S203: input the processed text image into a preset text recognition model to obtain the text in the image to be detected.

[0048] In an embodiment of the present invention, the processed text image is input into a preset text recognition model (a text recognition model obtained through training, the specific model is not limited here), and the preset text recognition model uses its learned feature mapping capability to recognize the text in the image and convert it into a text string. In this way, the text in the image to be detected can be obtained.

[0049] Optionally, during the text recognition process, data augmentation (by setting multiple reference points in the image, then moving these reference points, and generating new images through geometric transformations) can be used to improve data diversity and the generalization ability of the model. The accuracy of text recognition can also be improved through learning rate strategies (for example, setting a larger learning rate in the early stage of training to speed up the convergence of the model, and gradually reducing the learning rate in the later stages of training to prevent the model from oscillating during the convergence process) and regularization methods (such as L1 regularization and L2 regularization (regularization technology used to reduce the complexity of the model and prevent overfitting)).

[0050] In summary, the preset text detection model locates the text area from the image to be detected and outputs a text image with position information. Subsequently, based on this position information, the detected text image is cropped, corrected, and enhanced to obtain a processed text image. After that, the processed text image is input into the preset text recognition model to obtain accurate text recognition results, thereby solving the technical problem in related technologies that the image text is prone to tilt and distortion, resulting in low text recognition accuracy.

[0051] In order to accurately obtain the preset text detection model, in the image text recognition method provided in Example 1 of the present application, an initial text detection model is constructed; multiple historical text images are obtained, and the initial text detection model is trained based on all the historical text images to obtain the preset text detection model.

[0052] In an embodiment of the present invention, an initial text detection model (such as a DBNet model) needs to be constructed. The initial text detection model first extracts the features of the input image, then upsamples the feature pyramid to the same size and performs feature cascade to obtain fused features, and then uses the fused features to predict probability maps and threshold maps. Finally, an approximate binary map is calculated using the probability map and threshold map to obtain the final detection result.

[0053] In an embodiment of the present invention, multiple historical text images are obtained as training sets (covering various types of vouchers, including but not limited to entrusted deduction agreements, transfer receipts, checks, etc.), and the initial text detection model is trained based on all historical text images to obtain a preset text detection model.

[0054] In order to improve the accuracy of constructing the initial text detection model, in the image text recognition method provided in Example 1 of the present application, a backbone network module is constructed, wherein the backbone network module is used to extract features of historical text images to obtain feature maps, and the backbone network module includes: a multi-layer neural network, each layer of the neural network is used to output feature maps of different resolutions; a feature pyramid module is constructed, wherein the feature pyramid module is used to upsample the feature maps of different resolutions output by each layer of the neural network; the feature pyramid module is adjusted using a preset strategy to obtain an adjusted feature pyramid module; and an initial text detection model is constructed based on the backbone network module and the adjusted feature pyramid module.

[0055] In an embodiment of the present invention, constructing an initial text detection model includes constructing a backbone network module and a feature pyramid module. The backbone network module comprises a multi-layer neural network responsible for extracting high-level features from the input historical text image. Starting from the initial input layer, the neural network gradually refines the intrinsic features of the image through a series of operations such as convolution and pooling, ultimately outputting feature maps of different resolutions at different layers. The feature pyramid module restores the feature maps output by the backbone network to a higher resolution by upsampling them.

[0056] In an embodiment of the present invention, a preset strategy (i.e., a recursive strategy) is adopted to adjust the feature pyramid module to obtain an adjusted feature pyramid module (i.e., a recursive feature pyramid network). This strategy adds a loop on the basis of the feature pyramid module, and inputs the fused feature map into the backbone network again for secondary feature extraction. This not only increases the depth of the model, but also enables the model to capture text features from different angles and levels, thereby improving the efficiency and accuracy of text detection.

[0057] In the embodiment of the present invention, an initial text detection model is constructed based on the backbone network module and the adjusted feature pyramid module, such as Figure 3 As shown, Figure 3This is a schematic diagram of an optional DBNet network structure based on a recursive feature pyramid according to an embodiment of the present application. The network architecture includes: a backbone network module and a recursive feature pyramid module. Among them, the backbone network module first receives an original image containing text as input. The input image enters the backbone network module, and different layers of the backbone network module output feature maps of different resolutions (such as 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32). The multi-scale feature maps generated by the backbone network enter the recursive feature pyramid module. In this module, different levels will obtain feature maps of the same size as the previous level through upsampling operations (i.e., Up×2). By cascading feature maps at different levels, the DBNet network based on the recursive feature pyramid can integrate local details and global context information. Afterwards, the feature maps of different resolutions (such as 1 / 4, 1 / 8, and 1 / 16) output by different levels of the recursive feature pyramid module are concatenated (i.e., Concat) after convolution (i.e., conv) and upsampling operations (i.e., UP×2, UP×4, and UP×8) to obtain a fused feature map. At the same time, the output of each layer of the recursive feature pyramid module is input into the backbone network again for additional feature extraction, forming a cyclic feature transfer.

[0058] In order to accurately obtain the preset text detection model, in the image text recognition method provided in Example 1 of the present application, each historical text image is transformed to obtain a transformed historical text image; all historical text images and all transformed historical text images are added to an image training set; based on the image training set, the initial text detection model is trained to obtain a preset text detection model.

[0059] Optionally, each collected historical text image is subjected to image transformations (such as hue transformation, transparency transformation, rotation, background blurring, and saturation transformation). This process generates image samples under various deformations and lighting conditions, thereby alleviating the problem of model overfitting. For example, rotation and background blurring can simulate the imaging of text at different angles and in complex backgrounds, while saturation transformation helps the model understand text boundaries under different color contrasts.

[0060] In an embodiment of the present invention, all historical text images that have not been transformed, as well as all historical text images that have undergone the above-mentioned transformation, are integrated into the same image training set (the image training set contains a large number of positive samples (i.e., images containing text) and negative samples (i.e., images without text)), which can help the model learn more common features and make its generalization ability stronger. The initial text detection model is trained based on the image training set, so that the model has a certain anti-interference ability to changes in the input data, noise, etc.

[0061] The preset text detection model includes: multiple filters, and the filters correspond to weight matrices. In order to improve the accuracy of the preset text detection model, in the image text recognition method provided in Example 1 of the present application, the weight matrix of each filter is mapped to obtain a weight vector, and the weight vector is used as the spatial point coordinates of the filter; based on all spatial point coordinates, the center point coordinates are determined; the distance between each spatial point coordinate and the center point coordinate is determined, and all distances are sorted to obtain a sorted set; the spatial point coordinates corresponding to the minimum distance in the sorted set are used as the target spatial point coordinates, and the filter corresponding to the target spatial point coordinates is trimmed.

[0062] Optionally, since the preset text detection model may have parameter redundancy, a lighter network can be obtained while ensuring the accuracy of the model by removing redundant channels, filters, neurons, etc. in the model.

[0063] In an embodiment of the present invention, the preset text detection model includes multiple filters, each filter corresponds to a weight matrix, and the weight matrix of each filter is mapped to a weight vector, and the weight vector is used as the point coordinates of the filter in multidimensional space (i.e., spatial point coordinates), which facilitates subsequent distance calculation and analysis.

[0064] In an embodiment of the present invention, after the weight matrices of all filters are mapped to spatial point coordinates, a center point coordinate can be calculated and determined based on these spatial point coordinates (such as calculating the average value of all spatial point coordinates), and the Euclidean distance between each spatial point coordinate and the center point coordinate is calculated. All calculated distances are sorted to generate a sorted set, and the spatial point coordinates closest to the center point are selected as the target spatial point coordinates. Since it is closest to the center point, it can be considered that the information of the filter overlaps with other filters. In order to reduce the redundancy of the model and improve the efficiency of the model, the filter corresponding to the target spatial point coordinate is trimmed, the model is slimmed down, and it is also ensured that the model does not lose key feature recognition capabilities.

[0065] In order to accurately determine the processed text image, in the image text recognition method provided in Example 1 of the present application, the text image is cropped based on the text position information to obtain a text area image; the resolution of the text area image is adjusted to obtain an adjusted text area image; and the processed text image is determined based on the adjusted text area image.

[0066] In the embodiment of the present invention, the text image is cropped according to the text position information, with the purpose of separating the pure text portion from the original image, which can improve the accuracy of text recognition and reduce the amount of calculation.

[0067] In an embodiment of the present invention, the cropped text area images may have different sizes and resolutions. In order to ensure the consistency and efficiency of the direction classifier and the text recognition model, these text area images can be unified to a higher resolution. The high-resolution image (i.e., the adjusted text area image) provides clearer details, which is conducive to the model capturing the subtle features of the text, thereby improving the recognition accuracy. Therefore, the processed text image can be determined based on the adjusted text area image.

[0068] In order to accurately obtain the processed text image, in the image text recognition method provided in Example 1 of the present application, a direction classifier is used to analyze the adjusted text area image to obtain the rotation information of the adjusted text area image; based on the rotation information, it is determined whether the adjusted text area image needs to be corrected; if the adjusted text area image needs to be corrected, the adjusted text image is corrected to obtain the processed text image.

[0069] In an embodiment of the present invention, a direction classifier can identify the rotation angle of a text area image, that is, whether the text is deflected relative to the positive direction of the image. By sending the adjusted text area image to the direction classifier for analysis, the rotation information of the adjusted text area image can be obtained, and based on the rotation information, it is determined whether the adjusted text area image needs to be corrected. If the adjusted text area image needs to be corrected, the adjusted text image is corrected (such as rotating it 30 degrees) based on the rotation information to obtain a processed text image. Even in the case of poor image quality or a large text tilt, the accuracy of text recognition can be improved through direction correction.

[0070] The image text recognition method provided in the embodiment of the present application can reduce the number of model parameters and improve the accuracy of the model by adopting a recursive feature pyramid in the DBNet model and trimming the filters in the model. The amount of sample data for model learning is increased by the data enhancement method, and the accuracy of the direction classifier in correcting text in the image is improved by increasing the resolution of the input image. At the same time, in the text recognition model, the learning sample data of the model is increased by the data enhancement method, which improves the model's accuracy in recognizing distorted text in the image, thereby improving the business operation accuracy of financial institution personnel.

[0071] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0072] Example 2

[0073] The present application also provides an image text recognition device. It should be noted that the image text recognition device of the present application can be used to execute the image text recognition method provided in the present application. The image text recognition device provided in the present application is introduced below.

[0074] According to an embodiment of the present application, a device for implementing the above-mentioned image text recognition method is also provided. Figure 4 is a schematic diagram of an optional image text recognition device according to an embodiment of the present application, such as Figure 4 As shown, the image text recognition device may include: a first input unit 40 , a first processing unit 41 , and a second input unit 42 .

[0075] The first input unit 40 is used to obtain an image to be detected and input the image to be detected into a preset text detection model to obtain a text image, wherein the text image corresponds to text position information;

[0076] The first processing unit 41 is used to process the text image according to the text position information to obtain a processed text image;

[0077] The second input unit 42 is used to input the processed text image into a preset text recognition model to obtain the text in the image to be detected.

[0078] The image text recognition device provided in the embodiment of the present application can obtain the image to be detected through the first input unit 40, and input the image to be detected into a preset text detection model to obtain a text image. The text image can be processed by the first processing unit 41 according to the text position information to obtain a processed text image. The processed text image can be input into the preset text recognition model through the second input unit 42 to obtain the text in the image to be detected.

[0079] Optionally, the image text recognition device includes: a first construction module, used to construct an initial text detection model before inputting the image to be detected into a preset text detection model to obtain a text image; a first training module, used to obtain multiple historical text images and train the initial text detection model based on all historical text images to obtain a preset text detection model.

[0080] Optionally, the first construction module includes: a first construction sub-module, used to construct a backbone network module, wherein the backbone network module is used to extract features of historical text images to obtain feature maps, and the backbone network module includes: a multi-layer neural network, each layer of the neural network is used to output feature maps of different resolutions; a second construction sub-module, used to construct a feature pyramid module, wherein the feature pyramid module is used to upsample the feature maps of different resolutions output by each layer of the neural network; a first adjustment sub-module, used to adjust the feature pyramid module using a preset strategy to obtain an adjusted feature pyramid module; and a third construction sub-module, used to construct an initial text detection model based on the backbone network module and the adjusted feature pyramid module.

[0081] Optionally, the first training module includes: a first transformation submodule, used to perform image transformation on each historical text image to obtain a transformed historical text image; a first adding submodule, used to add all historical text images and all transformed historical text images to an image training set; a first training submodule, used to train the initial text detection model based on the image training set to obtain a preset text detection model.

[0082] Optionally, the preset text detection model includes: multiple filters, each filter corresponds to a weight matrix, and the image text recognition device includes: a first mapping module, which is used to train the initial text detection model based on all historical text images to obtain the preset text detection model, and then map the weight matrix of each filter to obtain a weight vector, and use the weight vector as the spatial point coordinate of the filter; a first determination module, which is used to determine the center point coordinate based on all spatial point coordinates; a second determination module, which is used to determine the distance between each spatial point coordinate and the center point coordinate, and sort all distances to obtain a sorted set; a first clipping module, which is used to use the spatial point coordinate corresponding to the minimum distance in the sorted set as the target spatial point coordinate, and clip the filter corresponding to the target spatial point coordinate.

[0083] Optionally, the first processing unit includes: a second cropping module, used to crop the text image based on the text position information to obtain a text area image; a first adjustment module, used to adjust the resolution of the text area image to obtain an adjusted text area image; and a third determination module, used to determine the processed text image based on the adjusted text area image.

[0084] Optionally, the third determination module includes: a first analysis submodule, used to analyze the adjusted text area image using a direction classifier to obtain rotation information of the adjusted text area image; a first judgment submodule, used to judge whether the adjusted text area image needs to be corrected based on the rotation information; and a first correction submodule, used to correct the adjusted text image if the adjusted text area image needs to be corrected to obtain a processed text image.

[0085] The above-mentioned image text recognition device may also include a processor and a memory. The above-mentioned first input unit 40, first processing unit 41, second input unit 42, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions.

[0086] The processor includes a kernel that retrieves the corresponding program unit from the memory. One or more kernels can be provided. By adjusting the kernel parameters, the processed text image is input into a preset text recognition model to obtain the text in the image to be detected.

[0087] The above-mentioned memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0088] It should be noted that the first input unit 40, the first processing unit 41, and the second input unit 42 correspond to steps S201 to S203 in Example 1. The examples and application scenarios implemented by the above units and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above units can also be part of the device and can be run in the computer terminal 10 provided in Example 1.

[0089] Example 3

[0090] The embodiment of the present application may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile terminal or an electronic device.

[0091] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.

[0092] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the image text recognition method: obtaining the image to be detected, and inputting the image to be detected into a preset text detection model to obtain a text image, wherein the text image corresponds to text position information; processing the text image based on the text position information to obtain a processed text image; inputting the processed text image into a preset text recognition model to obtain the text in the image to be detected.

[0093] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the image text recognition method: constructing an initial text detection model; obtaining multiple historical text images, and training the initial text detection model based on all historical text images to obtain a preset text detection model.

[0094] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the image text recognition method: constructing a backbone network module, wherein the backbone network module is used to extract features of historical text images to obtain feature maps, and the backbone network module includes: a multi-layer neural network, each layer of the neural network is used to output feature maps of different resolutions; constructing a feature pyramid module, wherein the feature pyramid module is used to upsample the feature maps of different resolutions output by each layer of the neural network; adjusting the feature pyramid module using a preset strategy to obtain an adjusted feature pyramid module; and constructing an initial text detection model based on the backbone network module and the adjusted feature pyramid module.

[0095] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the image text recognition method: performing image transformation on each historical text image to obtain the transformed historical text image; adding all historical text images and all transformed historical text images to an image training set; based on the image training set, training the initial text detection model to obtain a preset text detection model.

[0096] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the image text recognition method: mapping the weight matrix of each filter to obtain a weight vector, and using the weight vector as the spatial point coordinates of the filter; determining the center point coordinates based on all spatial point coordinates; determining the distance between each spatial point coordinate and the center point coordinate, and sorting all distances to obtain a sorted set; using the spatial point coordinates corresponding to the minimum distance in the sorted set as the target spatial point coordinates, and cutting out the filter corresponding to the target spatial point coordinates.

[0097] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the image text recognition method: cropping the text image based on the text position information to obtain a text area image; adjusting the resolution of the text area image to obtain an adjusted text area image; and determining the processed text image based on the adjusted text area image.

[0098] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the image text recognition method: using a direction classifier to analyze the adjusted text area image to obtain rotation information of the adjusted text area image; based on the rotation information, judging whether the adjusted text area image needs to be corrected; if the adjusted text area image needs to be corrected, correcting the adjusted text image to obtain a processed text image.

[0099] Optionally, Figure 5 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 5 As shown, the electronic device may include: one or more ( Figure 5 Only one is shown) processor 502, memory 504, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0100] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image text recognition method and device in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned image text recognition method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, corporate intranet, local area network, mobile communication network and combinations thereof.

[0101] The processor can call the information and application programs stored in the memory through the transmission device to execute the above steps in the above image text recognition method.

[0102] By using the embodiment of the present application, a solution for recognizing image text is provided. Through a recursive feature pyramid network, the accuracy of feature extraction and text area detection can be improved, the generalization ability of the model can be improved through data enhancement, and the learning rate adjustment strategy and regularization method are used to prevent model overfitting, and the network redundancy is reduced through filter clipping. It can realize efficient and accurate detection and recognition of tilted and distorted text on vouchers, and realize high-precision automatic extraction of text information, thereby solving the technical problem in related technologies that image text is prone to tilt and distortion, resulting in low text recognition accuracy.

[0103] It can be understood by those skilled in the art that Figure 5 The structure shown is for illustration only, and the electronic device may also be a terminal device such as a smart phone, a tablet computer, a PDA, or a mobile Internet device (MID). Figure 5 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 5 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 5 Different configurations shown.

[0104] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0105] Example 4

[0106] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the image text recognition method provided in the first embodiment.

[0107] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0108] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing the steps of the image text recognition method.

[0109] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0110] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0111] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0112] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0113] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0114] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0115] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for recognizing image text, characterized in that: include: Acquire an image to be detected, and input the image to be detected into a preset text detection model to obtain a text image, wherein the text image corresponds to text position information; Processing the text image according to the text position information to obtain a processed text image; The processed text image is input into a preset text recognition model to obtain the text in the image to be detected.

2. The image text recognition method according to claim 1, characterized in that: Before inputting the image to be detected into a preset text detection model to obtain a text image, the method further includes: Build an initial text detection model; A plurality of historical text images are obtained, and the initial text detection model is trained based on all the historical text images to obtain the preset text detection model.

3. The image text recognition method according to claim 2, characterized in that: The steps to build the initial text detection model include: Constructing a backbone network module, wherein the backbone network module is used to extract features from the historical text image to obtain a feature map, and the backbone network module includes: a multi-layer neural network, each layer of the neural network is used to output a feature map of different resolutions; Constructing a feature pyramid module, wherein the feature pyramid module is used to upsample the feature maps of different resolutions output by each layer of the neural network; Adjusting the feature pyramid module using a preset strategy to obtain an adjusted feature pyramid module; Based on the backbone network module and the adjusted feature pyramid module, the initial text detection model is constructed.

4. The image text recognition method according to claim 2, characterized in that: The step of training the initial text detection model based on all the historical text images to obtain the preset text detection model includes: Performing image transformation on each of the historical text images to obtain a transformed historical text image; Adding all of the historical text images and all of the transformed historical text images to an image training set; Based on the image training set, the initial text detection model is trained to obtain the preset text detection model.

5. The image text recognition method according to claim 2, characterized in that: The preset text detection model includes: a plurality of filters, each of which has a corresponding weight matrix. After the initial text detection model is trained based on all the historical text images to obtain the preset text detection model, the method further includes: Mapping the weight matrix of each filter to obtain a weight vector, and using the weight vector as the spatial point coordinates of the filter; Based on the coordinates of all the spatial points, determine the coordinates of the center point; Determine the distance between the coordinates of each of the spatial points and the coordinates of the center point, and sort all the distances to obtain a sorted set; The spatial point coordinates corresponding to the minimum distance in the sorted set are used as target spatial point coordinates, and the filter corresponding to the target spatial point coordinates is clipped.

6. The image text recognition method according to claim 1, characterized in that: The step of processing the text image according to the text position information to obtain a processed text image includes: Cropping the text image according to the text position information to obtain a text area image; Adjusting the resolution of the text region image to obtain an adjusted text region image; The processed text image is determined based on the adjusted text region image.

7. The image text recognition method according to claim 6, characterized in that: The step of determining the processed text image based on the adjusted text region image comprises: Analyzing the adjusted text region image using a direction classifier to obtain rotation information of the adjusted text region image; Based on the rotation information, determining whether correction of the adjusted text region image is required; In the case where the adjusted text region image needs to be corrected, the adjusted text image is corrected to obtain the processed text image.

8. An image text recognition device, characterized in that: include: A first input unit is used to obtain an image to be detected and input the image to be detected into a preset text detection model to obtain a text image, wherein the text image corresponds to text position information; a first processing unit, configured to process the text image according to the text position information to obtain a processed text image; The second input unit is used to input the processed text image into a preset text recognition model to obtain the text in the image to be detected.

9. A computer program product, characterized in that The invention comprises a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image text recognition method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: The invention comprises one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the image text recognition method described in any one of claims 1 to 7.