Image text recognition method, device, equipment and readable storage medium

By applying the image text recognition model of feature pyramid network and sequence transformation network in the field of automobile finance loans, the problem of inefficient automatic text recognition of certificate pictures is solved, automatic verification is realized, and cost and risk are reduced.

CN111368709BActive Publication Date: 2025-06-06WEBANK (CHINA)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010134748.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-02-28
Publication Date
2025-06-06
Estimated Expiration
2040-02-28

AI Technical Summary

Technical Problem

In the field of automobile finance loans, automatic text recognition technology for certificate pictures has not been effectively solved, resulting in time-consuming and labor-intensive verification process, inefficient efficiency, and increasing loan risks.

Method used

By obtaining the picture to be identified and entering the preset picture text recognition model, the text area is recognized using the feature pyramid network (FPN), and the text content is extracted through the sequence transformation network, thereby automatically identifying the text in the picture.

Benefits of technology

It realizes automatic identification of text in document pictures, reduces verification costs, improves verification efficiency, and reduces loan risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111368709B_ABST
    Figure CN111368709B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment and readable storage medium for recognizing text in an image, and relates to the field of financial technology. The method comprises the steps of: obtaining an image to be recognized, inputting the image to be recognized into a preset image text recognition model to obtain the region coordinates of a text region corresponding to the image to be recognized, and obtaining the text content corresponding to the text region; determining an associated text region according to the region coordinates corresponding to the text region and the text content corresponding to the region coordinates; and obtaining a text recognition result containing semantics in the image to be recognized according to the associated text region. The present invention realizes automatic recognition of text in an image, and improves the recognition efficiency of text in an image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text recognition technology in financial technology (Fintech), and in particular to a method, device, equipment and readable storage medium for recognizing image text. Background Art

[0002] With the development of computer technology, more and more technologies are applied in the financial field. The traditional financial industry is gradually transforming into financial technology (Fintech), and text recognition technology is no exception. However, due to the security and real-time requirements of the financial industry, higher requirements are also placed on text recognition technology.

[0003] In the field of auto finance loans, the funding party often requires the borrower to take photos and upload pictures of various certificate information (such as ID card, driver's license, vehicle registration certificate, etc.) for authenticity verification. Then these pictures will be reviewed and verified by a dedicated person, and then further risk control will be carried out to finally decide whether to lend. The step of dedicated review and verification is very time-consuming and labor-intensive, especially for the vehicle registration certificate, which has a lot of information, including but not limited to: driving license number, vehicle license number, frame number, engine number, ID number, registration authority, license plate number, unified social credit code, name and mortgagee, etc. And because the format of the vehicle registration certificate is not as regular as the ID card, many fields will be relatively messy, and even seriously blurred and misplaced, which increases the workload and difficulty of the verification personnel; and because the number of certificate pictures is huge, sometimes it is impossible to verify all pictures, so only a part of the pictures can be randomly checked for verification, which significantly increases the risk of the loan.

[0004] It can be seen that if the text in the image can be automatically recognized and then the image is automatically verified based on the recognized text, the cost of verifying the ID image can be reduced and the efficiency of verifying the ID image can be improved. Therefore, how to automatically recognize the text in the image is an urgent problem to be solved. Summary of the invention

[0005] The main purpose of the present invention is to provide a method, device, equipment and readable storage medium for recognizing text in an image, aiming to solve the technical problem of how to automatically recognize text in an image.

[0006] To achieve the above object, the present invention provides a method for recognizing text in an image, the method comprising the steps of:

[0007] Obtaining a picture to be identified, and inputting the picture to be identified into a preset picture text recognition model to obtain the region coordinates of a text region corresponding to the picture to be identified, and obtaining text content corresponding to the text region;

[0008] Determining an associated text area according to area coordinates corresponding to the text area and text content corresponding to the area coordinates;

[0009] A text recognition result containing semantics in the image to be recognized is obtained according to the associated text area.

[0010] Preferably, the step of inputting the image to be recognized into a preset image text recognition model to obtain the region coordinates of the text region corresponding to the image to be recognized includes:

[0011] Inputting the image to be recognized into a preset image text recognition model, and recognizing the text area in the image to be recognized by using a feature pyramid network FPN in the image text recognition model;

[0012] The region coordinates of the text region are determined according to the pixel values ​​of the text region in the image to be recognized.

[0013] Preferably, the step of identifying the text area in the image to be identified by using the FPN in the image text recognition model includes:

[0014] Performing feature extraction and feature fusion on the image to be recognized through the FPN in the image text recognition model, obtaining a first feature map corresponding to the image to be recognized;

[0015] Inputting the first feature map into the convolution layer of the FPN to obtain a second feature map corresponding to the image to be recognized;

[0016] Determine a text area in the image to be recognized based on the text pixels in the second feature map.

[0017] Preferably, the step of determining the text area in the to-be-recognized picture according to the text pixels in the second feature map comprises:

[0018] Determine text pixels in the second feature map, and determine core pixels in the second feature map according to the text pixels;

[0019] Each pixel in the second feature map is classified based on the core pixel, and the text area in the image to be identified is determined according to the classification result obtained by the classification.

[0020] Preferably, the step of obtaining the text content corresponding to the text area includes:

[0021] Inputting the text region into the network structure of the image text recognition model to obtain a third feature map corresponding to the text region;

[0022] Inputting the third feature graph into a sequence transformation network corresponding to the network structure to obtain a serialized fourth feature graph;

[0023] A fully connected network is connected according to the fourth feature graph and each node in the sequence transformation network to obtain text content corresponding to the text area.

[0024] Preferably, before the step of obtaining the image to be identified, the method further includes:

[0025] Obtaining a first sample image for model training, and annotating the first sample image to obtain a training sample set consisting of the annotated first sample images;

[0026] The training sample set is input into the image text recognition model to train the image text recognition model.

[0027] Preferably, the step of labeling the first sample images to obtain a training sample set consisting of the labeled first sample images includes:

[0028] Annotating the first sample images to obtain annotated first sample images, and calculating the number of images of the annotated first sample images;

[0029] If the number of pictures is less than the preset number, performing picture simulation according to the marked first sample pictures to obtain second sample pictures;

[0030] The second sample image and the labeled first sample image are used as a training sample set.

[0031] Preferably, the image to be identified is a certificate image of the lender, and after the step of obtaining a text recognition result containing semantics in the image to be identified according to the associated text area, the step further includes:

[0032] Comparing the text recognition result with the pre-stored certificate information of the lender to obtain a comparison result;

[0033] If it is determined according to the comparison result that the text recognition result is consistent with the certificate information, then it is determined that the certificate corresponding to the image to be recognized is a real certificate.

[0034] In addition, to achieve the above-mentioned purpose, the present invention further provides a picture text recognition device, the picture text recognition device comprising:

[0035] An acquisition module is used to acquire the image to be identified;

[0036] An input module, used for inputting the image to be recognized into a preset image text recognition model to obtain the region coordinates of the text region corresponding to the image to be recognized, and to obtain the text content corresponding to the text region;

[0037] A determination module, configured to determine an associated text region according to region coordinates corresponding to the text region and text content corresponding to the region coordinates;

[0038] A processing module is used to obtain a text recognition result containing semantics in the image to be recognized according to the associated text area.

[0039] In addition, to achieve the above-mentioned purpose, the present invention also provides a picture text recognition device, which includes a memory, a processor, and a picture text recognition program stored in the memory and executable on the processor, and when the picture text recognition program is executed by the processor, the steps of the picture text recognition method corresponding to the federated learning server are implemented.

[0040] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a picture text recognition program is stored, and when the picture text recognition program is executed by a processor, the steps of the picture text recognition method as described above are implemented.

[0041] The present invention obtains a picture to be recognized, inputs the picture to be recognized into a picture text recognition model to obtain the region coordinates of a text region corresponding to the picture to be recognized, and obtains text content corresponding to the text region, determines an associated text region according to the region coordinates corresponding to the text region and the text content corresponding to the region coordinates, and obtains a text recognition result containing semantics in the picture to be recognized according to the associated text region correspondence, thereby realizing automatic recognition of text in the picture and improving recognition efficiency of text in the picture. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a flow chart of the first embodiment of the method for recognizing text in an image according to the present invention;

[0043] Figure 2 It is a flow chart of a second embodiment of a method for recognizing text in an image according to the present invention;

[0044] Figure 3 It is a functional schematic diagram module diagram of a preferred embodiment of the image text recognition device of the present invention;

[0045] Figure 4 It is a structural diagram of the hardware operating environment involved in the embodiment of the present invention.

[0046] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0047] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0048] The present invention provides a method for recognizing image text, referring to Figure 1 , Figure 1 FIG. 1 is a flow chart of a first embodiment of a method for recognizing text in an image according to the present invention.

[0049] The embodiment of the present invention provides an embodiment of a method for recognizing text in an image. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be executed in an order different from that shown here.

[0050] The image text recognition method is applied to an image text recognition device, which may be a server or a terminal. The terminal may include mobile terminals such as mobile phones, tablet computers, laptop computers, cameras, PDAs, and fixed terminals such as digital TVs and desktop computers. In various embodiments of the image text recognition method, for ease of description, the execution subject is omitted to illustrate various embodiments. The image text recognition method includes:

[0051] Step S10, obtaining a picture to be recognized, inputting the picture to be recognized into a preset picture text recognition model to obtain the region coordinates of the text region corresponding to the picture to be recognized, and obtaining the text content corresponding to the text region.

[0052] Obtain the picture to be identified. Specifically, the picture to be identified can be obtained in real time, such as being collected in real time by a picture text recognition device, or being sent to a recognition device by other terminals, or being pre-stored. The picture to be identified can be one picture or multiple pictures. When the picture to be identified is pre-stored, a timed task can be pre-set, and the pre-stored picture to be identified can be obtained through the timed task. In this embodiment, the picture to be identified can be a certificate picture, or a picture of a license plate number. After obtaining the picture to be identified, the picture to be identified is input into a preset picture text recognition model to obtain the area coordinates of the text area corresponding to the picture to be identified, and to obtain the text content corresponding to the text area. It should be noted that each picture to be identified corresponds to at least one text area, and each text area has at least one character, and the text character is the text content corresponding to the text area. In this embodiment, the shape of the text area is not limited. In this embodiment, the picture text recognition model is an OCR (Optical Character Recognition) model. In other embodiments, the picture text recognition model can also be other models that can recognize text in pictures. Specifically, an OCR engine may be installed to implement the calling of the OCR model.

[0053] It should be noted that the text recognition process in multiple images to be recognized is the same as the text recognition process in a single image to be recognized. Therefore, for ease of description, the embodiment of the present invention is described using a single image to be recognized.

[0054] Furthermore, the step of inputting the image to be recognized into a preset image text recognition model to obtain the region coordinates of the text region corresponding to the image to be recognized includes:

[0055] Step a: input the image to be identified into a preset image text recognition model, and identify the text area in the image to be identified through the feature pyramid network FPN in the image text recognition model.

[0056] Step b: determining the region coordinates of the text region according to the pixel values ​​of the text region in the image to be identified.

[0057] Specifically, after obtaining the picture to be identified, the picture to be identified is input into the FPN (Feature Pyramid Networks) in the preset picture text recognition model to identify the text area in the picture to be identified through FPN. It should be noted that FPN is the basic architecture in the picture text recognition model. FPN is mainly used to solve the multi-scale problem in object detection. Through simple network connection changes, the performance of small object detection is greatly improved without basically increasing the calculation amount of the original model. After the text area in the picture to be identified is obtained through the FPN in the picture text recognition model, the regional coordinates of the text area are determined according to the pixel value of the identified text area in the picture to be identified. It can be understood that since each pixel in the picture to be identified has a corresponding pixel value, the position of the corresponding pixel in the picture to be identified can be determined by the pixel value. Therefore, the regional coordinates of the text area can be determined by the pixel value of each pixel in the picture to be identified.

[0058] Furthermore, the step of identifying the text area in the image to be identified by using the FPN in the image text recognition model includes:

[0059] Step a1: extract and fuse features of the image to be recognized through FPN in the image text recognition model to obtain a first feature map corresponding to the image to be recognized.

[0060] Further, it should be noted that FPN is divided into the first half and the second half. The first half is used to extract features from large to small in the image to be identified, and the second half is used to fuse features from small to large in the image to be identified. It can be understood that the larger the feature, the more information the feature contains and the more complex it is. In the first half of FPN, each layer uses a structure similar to ResNet (Residual Network). In FPN, it includes convolutional layers and pooling layers. There are K1 layers in the first and second half of FPN, respectively, where the size of K1 can be pre-set, and this embodiment does not specifically limit the size of K1. It can be understood that when the image to be identified passes through the K1 layer, K1 feature maps corresponding to the image to be identified will be obtained, that is, each time a layer is passed, a feature map will be obtained. Since the characteristics of each layer are different, the size of the resulting feature map is also different. It can be seen that by performing feature extraction and feature fusion on the image to be recognized through the FPN in the image text recognition model, the first feature map corresponding to the image to be recognized can be obtained. At this time, the sizes of the first feature maps corresponding to the image to be recognized are inconsistent.

[0061] Step a2: input the first feature map into the convolution layer of the FPN to obtain a second feature map corresponding to the image to be identified.

[0062] In order to facilitate the subsequent processing of the first feature map and improve the processing efficiency of the first feature map, after obtaining each first feature map corresponding to the image to be identified, the target feature map with the largest area in the first feature map is determined, and then each first feature map is upsampled to adjust the size of each first feature map to the same size as the target feature map. Figure 1 The size of each first feature map is adjusted to match the target feature map. Figure 1 In the same process, each first characteristic graph can be enlarged in the same proportion, that is, when the first characteristic graph is enlarged, the length and width are enlarged in the same proportion.

[0063] After obtaining the adjusted first feature map, the adjusted first feature map is input into the convolution layer of FPN to obtain the second feature map corresponding to the image to be identified. The number of convolution layers can be set to K2, where the size of K2 can be the same as K1 or different from K1. In the K2 convolution layers, the size of each convolution layer can be the same or different, and the user can set the size of each convolution layer according to specific needs.

[0064] Step a3: determine the text area in the image to be identified based on the text pixels in the second feature map.

[0065] After obtaining the second feature map, determine the text pixels in the second feature map, and determine the text area in the image to be identified based on the text pixels in the second feature map. Specifically, in the obtained second feature map, each pixel has a corresponding label value, and the label value can be used to determine whether the pixel is a text pixel. For example, the label value representing a text pixel can be set to "1", and the label value representing a non-text pixel can be set to "0". Therefore, when the label value corresponding to a pixel in the second feature map is "1", it indicates that the pixel is a text pixel; when the label value corresponding to a pixel in the second feature map is "0", it indicates that the pixel is a non-text pixel. This embodiment does not limit the expression form of the label value, such as the label value representing a text pixel can be expressed as "true", and the label value representing a non-text pixel can be expressed as "false".

[0066] Further, step a3 includes:

[0067] Step a31, determining text pixel points in the second feature map, and determining core pixel points in the second feature map according to the text pixel points.

[0068] Step a32, classifying each pixel in the second feature map based on the core pixel, and determining the text area in the image to be identified according to the classification result obtained by the classification.

[0069] Further, after obtaining the second feature map, the text pixels in the second feature map are determined, and the pixels in the second feature map corresponding to the text pixels are classified to obtain the pixel classification result. Specifically, the core pixel in the second feature map is determined, wherein the core pixel is a pixel whose surrounding pixels are all text pixels, that is, the four pixels above, below, left and right of the core pixel are all core pixels. After determining the core pixel, the core pixel is expanded to the surrounding area with the core pixel as the center. The expansion method can be specifically as follows: if the pixel to be expanded is recorded as x, the corresponding value is f(x), f(x) indicates how many proportions of pixels in the adjacent pixels of the pixel to be expanded have been successfully expanded, or are text pixels themselves, and the number of adjacent pixels to the pixel to be expanded can be set according to specific needs. It should be noted that if a certain pixel is successfully expanded, it indicates that the pixel is also confirmed as a text pixel. When the f(x) corresponding to the pixel to be expanded is greater than the preset threshold, it indicates that the pixel x is successfully expanded, wherein the preset threshold can be set according to specific needs, such as 0.6, 0.7 or 0.85. For example, when the preset threshold is set to 0.7, if 8 of the 10 adjacent pixels corresponding to the pixel x to be expanded are text pixels (including pixels that are confirmed as text pixels due to successful expansion), the expanded pixel can be confirmed as a text pixel. It is understandable that the pixel to be expanded is a pixel that was not originally a text pixel.

[0070] After performing the above operation on all pixels in all second feature maps, all text pixels and all non-text pixels in the second feature map can be determined, thereby obtaining the classification results of each pixel, that is, the pixels in the second feature map are divided into text pixels and non-text pixels. After determining the text pixels and non-text pixels in the second feature map, the text area in the image to be identified can be determined based on the text pixels in the second feature map.

[0071] Furthermore, the step of obtaining the text content corresponding to each text area includes:

[0072] Step c: input the text area into the network structure of the image text recognition model to obtain a third feature map corresponding to the text area.

[0073] Furthermore, after determining the text area in the image to be recognized, the text area is input into the network structure (backbone) in the image recognition model to obtain the third feature map corresponding to the text area. Specifically, each text area in the image to be recognized can be cut to obtain the regional image corresponding to each text area, and the regional image can be input into the network structure in the image text recognition model; or each text area in the image to be recognized can be annotated to obtain the annotated image to be recognized, and the annotated image to be recognized can be input into the network structure of the image text recognition model. The network structure can be ResNet or VGG (Visual Geometry Group)16, etc.

[0074] Step d: input the third feature graph into a sequence transformation network corresponding to the network structure to obtain a serialized fourth feature graph.

[0075] After the third feature graph is obtained, the second feature graph is input into the sequence transformation network corresponding to the network structure, and the third feature graph is serialized by the sequence transformation network to obtain a serialized fourth feature graph. In this embodiment, the sequence transformation network includes but is not limited to LSTM (Long Short-Term Memory) and BiLSTM.

[0076] Step e: connecting the fully connected network according to the fourth feature graph and each node in the sequence transformation network to obtain the text content corresponding to the text area.

[0077] It should be noted that, in this embodiment, each node in the sequence transformation network is connected to a fully connected network, and this embodiment does not limit what kind of fully connected network. After the fourth feature graph is obtained, the fourth feature graph is input into the fully connected network connected to each node in the sequence transformation network, and then the text content corresponding to the text area is obtained through the softmax of the fully connected network connection.

[0078] Step S20: determining an associated text region according to region coordinates corresponding to the text region and text contents corresponding to the region coordinates.

[0079] After obtaining the text content corresponding to the text area in the image to be identified, the associated text area in the text area is determined according to the area coordinates corresponding to the text area and the text content corresponding to the area coordinates. It can be understood that in the process of determining the associated text area, the associated area coordinates can be determined first, the text content corresponding to the associated area coordinates can be determined as the text target content, and then the associated text target content can be determined according to the semantics of each text target content, and the text area corresponding to the associated text target content can be determined as the associated text area; or the associated text target content can be determined first, and then the associated area coordinates can be determined.

[0080] It is understandable that the associated text areas can be determined according to the size of the coordinates of the corresponding areas of each text area, and the associated text areas can be one or more of the upper associated text area, the lower associated text area, the left associated text area and the right associated text area. It is understandable that the associated text areas are adjacent text areas, such as determining that the A text area is associated with the B text area, the C text area and the D text area. In the process of determining the associated text content, it is mainly determined based on semantics, such as the text content corresponding to the A text area is "identity card number", and the text content corresponding to the B text area meets the requirements corresponding to the "identity card number", such as a total of 18 characters, including address code, date of birth code, digital code and check code. At this time, it can be finally determined that the A text area and the B text area are associated.

[0081] Step S30, obtaining a text recognition result containing semantics in the to-be-recognized image according to the associated text area.

[0082] After the associated text area is determined, the text recognition result containing semantics in the image to be identified corresponding to the associated text area is obtained. Specifically, corresponding semantics can be assigned to each text area according to the text content. For example, a "name" is assigned to the text area corresponding to "Zhang Xiaoming", an "address" is assigned to the text area corresponding to "Village D, Town C, City B, Guangdong Province", and a "name tag" is assigned to the text area corresponding to "name", such as "Name: Zhang Xiaoming" on the ID card. Then, the semantics corresponding to each text area in the image to be identified are determined to obtain the text recognition result in the image to be identified, that is, a structured text recognition result containing semantics is obtained.

[0083] This embodiment obtains a picture to be recognized, inputs the picture to be recognized into a picture text recognition model to obtain the region coordinates of the text area corresponding to the picture to be recognized, and obtains the text content corresponding to the text area, determines the associated text area according to the region coordinates corresponding to the text area and the text content corresponding to the region coordinates, and obtains the semantic text recognition result in the picture to be recognized according to the associated text area correspondence, thereby realizing automatic recognition of text in pictures and improving the recognition efficiency of text in pictures.

[0084] It should be noted that the present embodiment can perform the recognition of text in the picture online or offline. When the recognition of text in the picture is performed online, the picture to be recognized is obtained in real time; when the recognition of text in the picture is performed offline, the picture to be recognized is pre-stored. If the recognition of text in the picture is performed online, the user can upload the picture to be recognized through the applet corresponding to the picture text recognition model. The text recognition result obtained at this time allows the user to modify it online. After obtaining the text recognition result, the text recognition result will be output to allow the user corresponding to the picture to be recognized to confirm whether the text recognition result is accurate. If the user confirms that the text recognition result is inaccurate, the inaccurate text recognition result can be modified. It should be noted that the user at this time is not the user of the information corresponding to the picture to be recognized, but the user corresponding to the picture text recognition model. For example, when the picture to be recognized is a certificate picture, the user is not the owner of the certificate, but the user who verifies the certificate picture. This is to prevent the user corresponding to the picture to be recognized from falsifying information. By performing the recognition of the picture text online, the load of the equipment used for the picture text recognition is guaranteed to be more balanced, and the correctness of the obtained text recognition result is guaranteed by giving the text recognition result to the user for confirmation. Furthermore, the image to be identified can be annotated using the text recognition results to obtain an annotated image to be identified, and the annotated image to be identified can be used to further train the image text recognition model to improve the accuracy of text recognition performed by the image text recognition model.

[0085] When performing image text recognition offline, there can be multiple images to be recognized each time, and the image text recognition is performed regularly at this time. The reason for offline image text recognition is that sometimes it is inconvenient or unnecessary to perform online recognition of image text. In the process of offline recognition of images to be recognized, you can choose to perform it when the network is better and the load of the corresponding device for image text recognition is lower, so as to improve the recognition efficiency of image text without worrying about network delays. Furthermore, if some images to be recognized contain sensitive information, in order to ensure the security of sensitive information, it can also be limited to offline recognition of images to be recognized containing sensitive information.

[0086] Furthermore, a second embodiment of the method for recognizing text in an image of the present invention is proposed. The difference between the second embodiment of the method for recognizing text in an image and the first embodiment of the method for recognizing text in an image is that, referring to Figure 2 , the image text recognition method also includes:

[0087] Step S40, obtaining a first sample image for model training, and annotating the first sample image to obtain a training sample set consisting of the annotated first sample images.

[0088] Obtain a first sample image for model training, wherein the first sample image can be obtained from other terminals when needed, or can be obtained in real time. It can be understood that the first sample image is also a picture containing text. In this embodiment, the number of first sample images is not limited, and the user can set the number of first sample images according to specific needs. After obtaining the first sample image, annotate the first sample image to obtain the annotated first sample image, and form the annotated first sample image into a training sample set. Specifically, a prompt message can be sent to the annotator to prompt the annotator to annotate the corresponding first sample image according to the prompt message to determine each character in each first sample image. Further, the text area and each character in each first sample image can also be annotated, and the text content can be determined by the annotated characters. In other embodiments, annotation can also be performed automatically.

[0089] Step S50: input the training sample set into the image text recognition model to train the image text recognition model.

[0090] The first sample image in the training sample set is input into the image text recognition model to train the image text recognition model, and the image text recognition model is stored. It should be noted that the processing of inputting the first sample image into the image text recognition model is consistent with the processing of inputting the image to be recognized into the image text recognition model, and this embodiment will not be repeated here.

[0091] In this embodiment, a first sample image is obtained to perform model training to obtain a picture text recognition model. When it is necessary to recognize text in a picture, the trained picture text recognition model can be used.

[0092] Furthermore, the step of labeling the first sample images to obtain a training sample set consisting of the labeled first sample images includes:

[0093] Step h: annotate the first sample images to obtain annotated first sample images, and calculate the number of images of the annotated first sample images.

[0094] Step i: If the number of pictures is less than a preset number, a picture simulation is performed based on the annotated first sample picture to obtain a second sample picture.

[0095] Step j: taking the second sample image and the annotated first sample image as a training sample set.

[0096] After obtaining the first sample picture, the first sample picture is annotated to obtain the annotated first sample picture, and the number of pictures of the annotated first sample picture is calculated to determine whether the picture data is less than the preset number. If it is determined that the picture data is less than the preset number, the picture simulation is performed according to the annotated first sample picture to obtain the second sample picture. Specifically, in the process of picture simulation, in order to improve the authenticity of the obtained second sample picture, it is necessary to ensure that the font in the second sample picture is the same as the first sample picture, and the background similarity between the second sample picture and the background of the first sample picture is greater than the preset similarity, wherein the present embodiment does not limit the size of the preset similarity. The corpus in the second sample picture can be determined according to the characteristics and format of the text in the first sample picture, that is, the text content in the second sample picture is determined. For example, the fictitious ID number in the second sample picture has the same format as the real ID number. It should be noted that the simulated second sample picture is also a well-annotated picture.

[0097] After the second sample images are obtained, the second sample images and the annotated first sample images are used as training sample sets. Further, if it is determined that the number of images is greater than or equal to the preset number, the number of first sample images is sufficient and no simulation is required to obtain the second sample images.

[0098] If so many sample images are not needed after obtaining the second sample image, a portion of sample images can be selected from the first sample image and the second sample image to form a training sample set. The quantity ratio between the first sample image and the second sample image in the training sample set is also set according to specific needs. This embodiment does not impose specific restrictions on the size of the quantity ratio.

[0099] Since the training of the image text recognition model requires a large number of sample images, it may be impossible to take out sufficiently real sample images for annotation due to reasons such as time, labor costs, insufficient real data volume and data sensitivity. Therefore, this embodiment simulates the sample images to increase the sample images for training the image text recognition model, thereby improving the accuracy of the trained image text recognition model in recognizing text.

[0100] Furthermore, a third embodiment of the image text recognition method of the present invention is proposed.

[0101] The third embodiment of the method for recognizing text in an image is different from the first and / or second embodiments of the method for recognizing text in an image in that the method for recognizing text in an image further includes:

[0102] Step k: compare the text recognition result with the pre-stored certificate information of the lender to obtain a comparison result.

[0103] If the image text recognition method is applied to loan business, such as auto finance loan or house mortgage loan, etc. After the text recognition result is obtained, the text recognition result is compared with the borrower's certificate information pre-stored in the database to obtain a comparison result. It can be understood that the borrower's certificate information will be filled in the loan information, and the certificate information is stored in the database.

[0104] Step 1: If it is determined according to the comparison result that the text recognition result is consistent with the certificate information, then it is determined that the certificate corresponding to the image to be recognized is a real certificate.

[0105] After the comparison result is obtained, if the text recognition result is determined to be consistent with the certificate information according to the comparison result, then the certificate corresponding to the image to be identified is determined to be a real certificate. It can be understood that the image to be identified is a certificate image at this time; if the text recognition result is determined to be inconsistent with the certificate information according to the comparison result, then the certificate corresponding to the image to be identified is determined to be a false certificate. It should be noted that when the information in the text recognition result is the same as the certificate information, it can be determined that the text recognition result is consistent with the certificate information, otherwise, it is determined that the text recognition result is inconsistent with the certificate information; it can also be determined that the text recognition result is consistent with the certificate information when the similarity between the information in the text recognition result and the certificate information is greater than the preset information similarity, otherwise, it is determined that the text recognition result is inconsistent with the certificate information. When it is determined that the certificate corresponding to the image to be identified is a real certificate, the lender is loaned; when the certificate corresponding to the image to be identified is a false certificate, the lender is refused to lend.

[0106] This embodiment compares the text recognition result with the lender's certificate information, thereby automatically verifying the lender's certificate information to ensure the authenticity of the lender's certificate information, thereby improving the security of the loan and reducing the risk of the loan.

[0107] In addition, the present invention also provides a device for recognizing image text, referring to Figure 3 , the image text recognition device comprises:

[0108] An acquisition module 10 is used to acquire a picture to be identified;

[0109] An input module 20, used for inputting the image to be recognized into a preset image text recognition model to obtain the region coordinates of the text region corresponding to the image to be recognized, and to obtain the text content corresponding to the text region;

[0110] A determination module 30, configured to determine an associated text region according to region coordinates corresponding to the text region and text content corresponding to the region coordinates;

[0111] The processing module 40 is used to obtain a text recognition result containing semantics in the image to be recognized according to the associated text area.

[0112] Furthermore, the input module 20 includes:

[0113] A first input unit, used for inputting the image to be recognized into a preset image text recognition model;

[0114] A recognition unit, used for recognizing a text area in the image to be recognized by using a feature pyramid network (FPN) in the image text recognition model;

[0115] The first determining unit is used to determine the region coordinates of the text region according to the pixel values ​​of the text region in the image to be identified.

[0116] Furthermore, the identification unit includes:

[0117] A feature processing subunit, used for performing feature extraction and feature fusion on the image to be recognized through the FPN in the image text recognition model to obtain a first feature map corresponding to the image to be recognized;

[0118] An input subunit, used to input the first feature map into the convolution layer of the FPN to obtain a second feature map corresponding to the picture to be recognized;

[0119] A determination subunit is used to determine the text area in the image to be identified based on the text pixels in the second feature map.

[0120] Furthermore, the determination subunit is also used to determine the text pixel points in the second feature map, and determine the core pixel points in the second feature map based on the text pixel points; classify each pixel point in the second feature map based on the core pixel points, and determine the text area in the image to be identified based on the classification results obtained by classification.

[0121] Furthermore, the input module 20 further includes:

[0122] The second input unit is used to determine the text pixel points in the second feature map, determine the core pixel points in the second feature map according to the text pixel points; classify each pixel point in the second feature map based on the core pixel points, and determine the text area in the image to be identified according to the classification result obtained by classification.

[0123] A processing unit is used to connect the fully connected network according to the fourth feature graph and each node in the sequence transformation network to obtain the text content corresponding to the text area.

[0124] Furthermore, the acquisition module 10 is also used to acquire a first sample image for model training;

[0125] The image text recognition device also includes:

[0126] A labeling module, used to label the first sample images to obtain a training sample set consisting of the labeled first sample images;

[0127] The input module 20 is further used to input the training sample set into the image text recognition model to train the image text recognition model.

[0128] Furthermore, the marking module includes:

[0129] a labeling unit, configured to label the first sample image to obtain a labeled first sample image;

[0130] A calculation unit, used to calculate the number of images of the first sample image after annotation;

[0131] A simulation unit, configured to perform picture simulation according to the annotated first sample picture to obtain a second sample picture if the number of pictures is less than a preset number;

[0132] The second determining unit is configured to use the second sample image and the labeled first sample image as a training sample set.

[0133] Furthermore, the image text recognition device also includes:

[0134] A comparison module, used for comparing the text recognition result with the pre-stored certificate information of the lender to obtain a comparison result;

[0135] The determination module 30 is further configured to determine that the certificate corresponding to the image to be identified is a genuine certificate if it is determined according to the comparison result that the text recognition result is consistent with the certificate information.

[0136] The specific implementation of the image text recognition device of the present invention is basically the same as the various embodiments of the above-mentioned image text recognition method, and will not be repeated here.

[0137] In addition, the present invention also provides a device for recognizing image text. Figure 4 As shown, Figure 4 It is a structural diagram of the hardware operating environment involved in the embodiment of the present invention.

[0138] It should be noted that Figure 4 That is, it can be a schematic diagram of the structure of the hardware operating environment of the image text recognition device. The image text recognition device of the embodiment of the present invention can be a terminal device such as a PC, a portable computer, etc.

[0139] like Figure 4As shown, the image text recognition device may include: a processor 1001, such as a CPU, a memory 1005, a user interface 1003, a network interface 1004, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0140] Those skilled in the art will understand that Figure 4 The structure of the image text recognition device shown in the figure does not constitute a limitation on the image text recognition device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0141] like Figure 4 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a picture text recognition program. The operating system is a program that manages and controls the hardware and software resources of the picture text recognition device, and supports the operation of the picture text recognition program and other software or programs.

[0142] exist Figure 4 In the picture text recognition device shown, the user interface 1003 is mainly used to connect to other terminals and communicate data with other terminals, and the picture to be recognized and / or the first sample picture can be obtained through other terminals; the network interface 1004 is mainly used for the background server and communicates data with the background server; the processor 1001 can be used to call the picture text recognition program stored in the memory 1005, and execute the steps of the picture text recognition method as described above.

[0143] The specific implementation of the image text recognition device of the present invention is basically the same as the various embodiments of the above-mentioned image text recognition method, and will not be repeated here.

[0144] In addition, an embodiment of the present invention further proposes a computer-readable storage medium, on which a picture text recognition program is stored. When the picture text recognition program is executed by a processor, the steps of the picture text recognition method as described above are implemented.

[0145] The specific implementation of the computer-readable storage medium of the present invention is basically the same as the various embodiments of the above-mentioned image text recognition method, and will not be repeated here.

[0146] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0147] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0148] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0149] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for recognizing text in an image. It is characterized in that The image text recognition method comprises the following steps: Obtaining a picture to be identified, and inputting the picture to be identified into a preset picture text recognition model to obtain the region coordinates of a text region corresponding to the picture to be identified, and obtaining text content corresponding to the text region; Determining an associated text area according to area coordinates corresponding to the text area and text content corresponding to the area coordinates; Obtaining a text recognition result containing semantics in the image to be recognized according to the semantics corresponding to the associated text area; The step of determining the associated text area according to the area coordinates corresponding to the text area and the text content corresponding to the area coordinates comprises: Determining region coordinates associated with the text region, and determining text content corresponding to the associated region coordinates as text target content; Determine the associated text target content according to the semantics of each of the text target contents, and determine the text area corresponding to the associated text target content as the associated text area, wherein the associated text area is a text area adjacent to the text area.

2. The method for recognizing text in an image according to claim 1, It is characterized in that The step of inputting the image to be recognized into a preset image text recognition model to obtain the region coordinates of the text region corresponding to the image to be recognized includes: Inputting the image to be recognized into a preset image text recognition model, and recognizing the text area in the image to be recognized by using a feature pyramid network FPN in the image text recognition model; The region coordinates of the text region are determined according to the pixel values ​​of the text region in the image to be recognized.

3. The method for recognizing text in an image according to claim 2, It is characterized in that The step of identifying the text area in the image to be identified by using the FPN in the image text recognition model comprises: Performing feature extraction and feature fusion on the image to be recognized through the FPN in the image text recognition model, obtaining a first feature map corresponding to the image to be recognized; Inputting the first feature map into the convolution layer of the FPN to obtain a second feature map corresponding to the image to be recognized; Determine a text area in the image to be recognized based on the text pixels in the second feature map.

4. The method for recognizing text in an image according to claim 3, It is characterized in that The step of determining the text area in the to-be-recognized picture according to the text pixels in the second feature map comprises: Determine text pixels in the second feature map, and determine core pixels in the second feature map according to the text pixels; Each pixel in the second feature map is classified based on the core pixel, and the text area in the image to be identified is determined according to the classification result obtained by the classification.

5. The method for recognizing text in an image according to claim 1, It is characterized in that The step of obtaining the text content corresponding to the text area comprises: Inputting the text region into the network structure of the image text recognition model to obtain a third feature map corresponding to the text region; Inputting the third feature graph into a sequence transformation network corresponding to the network structure to obtain a serialized fourth feature graph; A fully connected network is connected according to the fourth feature graph and each node in the sequence transformation network to obtain text content corresponding to the text area.

6. The method for recognizing text in an image according to claim 1, It is characterized in that Before the step of obtaining the image to be identified, the method further includes: Obtaining a first sample image for model training, and annotating the first sample image to obtain a training sample set consisting of the annotated first sample images; The training sample set is input into the image text recognition model to train the image text recognition model.

7. The method for recognizing text in an image according to claim 6, It is characterized in that The step of labeling the first sample images to obtain a training sample set consisting of the labeled first sample images comprises: Annotating the first sample images to obtain annotated first sample images, and calculating the number of images of the annotated first sample images; If the number of pictures is less than the preset number, performing picture simulation according to the marked first sample pictures to obtain second sample pictures; The second sample image and the labeled first sample image are used as a training sample set.

8. The method for recognizing text in an image according to any one of claims 1 to 7, It is characterized in that The image to be identified is a certificate image of the lender. After the step of obtaining a text recognition result containing semantics in the image to be identified according to the semantics corresponding to the associated text area, the method further includes: Comparing the text recognition result with the pre-stored certificate information of the lender to obtain a comparison result; If it is determined according to the comparison result that the text recognition result is consistent with the certificate information, then it is determined that the certificate corresponding to the image to be recognized is a real certificate.

9. A device for recognizing text in an image, It is characterized in that The image text recognition device comprises: An acquisition module is used to acquire the image to be identified; An input module, used for inputting the image to be recognized into a preset image text recognition model to obtain the region coordinates of the text region corresponding to the image to be recognized, and to obtain the text content corresponding to the text region; A determination module, configured to determine an associated text region according to region coordinates corresponding to the text region and text content corresponding to the region coordinates; A processing module, used for obtaining a text recognition result containing semantics in the to-be-recognized picture according to the semantics corresponding to the associated text area; Among them, the determination module is also used to: determine the area coordinates associated with the text area, and determine the text content corresponding to the associated area coordinates as the text target content; determine the associated text target content according to the semantics of each of the text target contents, and determine the text area corresponding to the associated text target content as the associated text area, wherein the associated text area is a text area adjacent to the text area.

10. A device for recognizing text in pictures, It is characterized in that The image text recognition device includes a memory, a processor, and a image text recognition program stored in the memory and executable on the processor. When the image text recognition program is executed by the processor, the steps of the image text recognition method as described in any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a picture text recognition program, and when the picture text recognition program is executed by a processor, the steps of the picture text recognition method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Text detection method and device, computer readable storage medium and computer device

    CN109886330A

  • An OCR identification method and electronic equipment thereof

    CN109919014A

  • End-to-end optical character detection and recognition method and system

    CN110598690A