Image and text recognition information processing method and device
Through the text labeling confidence and character proportion of the image text recognition model, targeted labeling corrections are made, which solves the problem of inefficient labeling of image training data and realizes efficient model training data preparation.
Patent Information
- Application Number
- CN202210150494.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-18
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2042-02-18
AI Technical Summary
In the prior art, the labeling efficiency of image text recognition training data is inefficient, and the labor cost is huge, which affects the efficiency of model training.
Through the confidence of the text labeling results of the image text recognition model, targeted text recognition labeling corrections are carried out, and text positioning labeling corrections are carried out according to the proportion of characters to filter out the correct labels that are really needed, and improve the labeling efficiency of training data.
It effectively improves the labeling efficiency of image training data, reduces the number of manual labeling, and improves the efficiency of model training.
Smart Images

Figure CN114529907B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence and can also be used in the financial field. Specifically, it relates to a method and device for image and text recognition information processing. Background Art
[0002] The development of artificial intelligence technology requires the training of many network models, and the training of network models requires a large amount of manually labeled training data.
[0003] In the field of image text recognition, the existing technology mostly labels training data manually using annotation tools on images, which has huge labor costs and low efficiency in preparing training data, which is not conducive to faster and better model training. Summary of the Invention
[0004] In response to the problems in the prior art, the present application provides an image text recognition information processing method and device, which can effectively improve the labeling efficiency of image training data used for model training.
[0005] In order to solve at least one of the above problems, the present application provides the following technical solutions:
[0006] In a first aspect, the present application provides a method for processing image text recognition information, comprising:
[0007] Perform image recognition on the image to be annotated according to a preset image and text recognition model to obtain text recognition information and text positioning information;
[0008] Correcting the character recognition annotation in the character recognition information according to the confidence level of the character recognition information, and determining the proportion of the corrected characters in the character recognition annotation when correcting the character recognition annotation;
[0009] The text positioning annotation in the text positioning information is corrected according to the character proportion to obtain an image annotated with the corrected text recognition information and text positioning information, and the image is used as image training data for the neural network model.
[0010] Furthermore, the step of correcting the text positioning annotation in the text positioning information according to the character proportion to obtain an image annotated with the corrected text recognition information and text positioning information includes:
[0011] Monitoring the recognition and annotation correction operation, obtaining an image after the recognition and annotation correction, and determining the proportion of the corrected characters in the text recognition and annotation;
[0012] If the character ratio of the corrected characters is higher than the character threshold, the text positioning annotation in the text positioning information of the image is corrected to obtain an image after the text positioning annotation is corrected.
[0013] Furthermore, it also includes:
[0014] If the confidence level is within a preset confidence level range, a marking review operation is performed on the image after text recognition and annotation correction; if the character ratio is within a preset character ratio range, a marking review operation is performed on the image after text positioning and annotation correction.
[0015] Furthermore, after performing the annotation review operation on the image after the text recognition annotation correction and / or the text positioning annotation correction, the method further includes:
[0016] The text recognition annotation and / or text positioning annotation of the image is updated according to the correction content of the annotation review operation.
[0017] Furthermore, it also includes:
[0018] If the confidence level is higher than the confidence threshold and the character percentage is lower than the character threshold, the image corrected by the text recognition annotation and the image corrected by the text positioning annotation are used as image training data for the neural network model.
[0019] Furthermore, the method of using the image as image training data for the neural network model further includes:
[0020] An image marked with the corrected text recognition information and text positioning information is used as image training data for a neural network model to perform image text recognition training on the neural network model to obtain a trained neural network model. The neural network model is used to perform image recognition on financial transaction credential images or identity credential images to obtain image recognition information for transaction processing.
[0021] In a second aspect, the present application provides an image and text recognition information processing device, comprising:
[0022] An initial recognition module is used to perform image recognition on the image to be annotated according to a preset image and text recognition model to obtain text recognition information and text positioning information;
[0023] a recognition annotation correction module, configured to correct the character recognition annotation in the character recognition information according to the confidence level of the character recognition information, and determine the proportion of the corrected characters in the character recognition annotation during the correction;
[0024] The positioning annotation correction module is used to correct the text positioning annotation in the text positioning information according to the character ratio, obtain an image annotated with the corrected text recognition information and text positioning information, and use the image as image training data for the neural network model.
[0025] Furthermore, the positioning annotation correction module includes:
[0026] a correction character ratio determination unit, configured to monitor the recognition and annotation correction operation, obtain an image after the recognition and annotation correction, and determine the character ratio of the correction character in the text recognition and annotation;
[0027] The high correction rate positioning correction unit is used to correct the text positioning annotation in the text positioning information of the image if the character ratio of the corrected character is higher than the character threshold, so as to obtain an image after the text positioning annotation correction.
[0028] Furthermore, it also includes:
[0029] a text review unit, configured to perform a label review operation on the image after the text recognition and labeling correction if the confidence level is within a preset confidence level range;
[0030] The positioning review unit is used to perform a marking review operation on the image after the text positioning marking correction if the character ratio is within a preset character ratio value range.
[0031] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the image and text recognition information processing method when executing the program.
[0032] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the image text recognition information processing method when executed by a processor.
[0033] In a fifth aspect, the present application provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the image text recognition information processing method.
[0034] It can be seen from the above technical solution that the present application provides an image text recognition information processing method and device, which performs targeted corrections to the text recognition annotations based on the confidence level of the text annotation results of the image text recognition model, and then performs targeted corrections to the text positioning annotations based on the proportion of characters that need to be corrected during the text recognition annotation correction, thereby accurately screening out the annotations that really need to be corrected and performing correction operations, thereby effectively improving the annotation efficiency of image training data used for model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0036] Figure 1 This is one of the flowcharts of the image text recognition information processing method in the embodiment of the present application;
[0037] Figure 2 This is a second flow chart of the image text recognition information processing method in an embodiment of the present application;
[0038] Figure 3 This is the third flow chart of the image text recognition information processing method in the embodiment of the present application;
[0039] Figure 4 This is one of the structural diagrams of the image and text recognition information processing device in the embodiment of the present application;
[0040] Figure 5 This is the second structural diagram of the image and text recognition information processing device in the embodiment of the present application;
[0041] Figure 6 This is the third structural diagram of the image and text recognition information processing device in the embodiment of the present application;
[0042] Figure 7 This is an overall flow chart of the image and text recognition information processing method in a specific embodiment of the present application;
[0043] Figure 8 Schematic diagram of the structure of the electronic device in the embodiment of the present application. DETAILED DESCRIPTION
[0044] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0045] Taking into account the problem that most of the annotated training data in the prior art are manually annotated on images using annotation tools, which has huge labor costs and low efficiency in preparing training data, and is not conducive to faster and better model training, the present application provides an image text recognition information processing method and device, which performs targeted corrections to the text recognition annotations by adjusting the confidence level of the text annotation results of the image text recognition model, and then performs targeted corrections to the text positioning annotations based on the proportion of characters that need to be corrected when correcting the text recognition annotations, and then accurately screens out the annotations that really need to be corrected for correction operations, thereby effectively improving the annotation efficiency of image training data used for model training.
[0046] In order to effectively improve the efficiency of labeling image training data for model training, this application provides an embodiment of an image text recognition information processing method, see Figure 1 The image character recognition information processing method specifically includes the following contents:
[0047] Step S101: performing image recognition on the image to be annotated according to a preset image and text recognition model to obtain text recognition information and text positioning information.
[0048] Optionally, the present application may first use a preset (which may be an existing) image text recognition model to recognize the annotated image, output text location information Loc, text recognition information Rec and confidence Con of the text recognition information, and then correct the text recognition annotation of the low-confidence image based on the numerical comparison result of the confidence and the confidence threshold.
[0049] Step S102: correcting the character recognition annotation in the character recognition information according to the confidence level of the character recognition information, and determining the character ratio of the corrected character in the character recognition annotation.
[0050] Optionally, the present application may adopt existing annotation correction methods to realize the correction of text recognition annotations, or may adopt manual correction methods to realize it, thereby obtaining an image after text recognition annotation correction.
[0051] Optionally, when performing correction on the text recognition annotation, the present application can monitor the correction process to determine how many characters are corrected, and further determine the correction ratio.
[0052] For example, the text recognition annotation output by the image text recognition model is "Zhang Mousan", and its confidence level is low. Therefore, after verification, it can be manually corrected to "Zhang Mouer". At this time, the system monitors the correction process to determine that one character is corrected out of three characters, so the correction ratio is 33%.
[0053] Step S103: Perform positioning annotation correction on the text positioning annotation in the text positioning information according to the character ratio, obtain an image marked with the corrected text recognition information and text positioning information, and use this image as the image training data of the neural network model.
[0054] Optionally, if the character ratio of the corrected characters is higher than the character threshold, it indicates that the image text recognition model is very inaccurate in recognizing the text of this image, and there is a strong correlation between the inaccuracy of the text recognition information and the inaccuracy of the text positioning information. That is, a high character ratio of the corrected characters indicates inaccurate text positioning.
[0055] Therefore, the present application performs positioning annotation correction on the text positioning information of the image after text recognition annotation correction, and obtains an image after positioning annotation correction.
[0056] Optionally, the present application can use existing annotation correction methods to implement the correction of text positioning annotation, or can also use manual correction methods to implement it, thereby obtaining an image after positioning annotation correction.
[0057] At this time, the present application has successively screened out the images that need to be corrected for text recognition and text positioning from the initial large number of images, completed the correction, and finally obtained the image training data that meets the model training requirements and can be used for model training.
[0058] As can be seen from the above description, the image text recognition information processing method provided by the embodiments of the present application can, through the confidence level of the text annotation result of the image text recognition model, perform targeted correction of the text recognition annotation, and then perform targeted correction of the text positioning annotation according to the character ratio that needs to be corrected during the text recognition annotation correction, so as to accurately screen out the annotations that really need to be corrected for correction operations, thereby effectively improving the annotation efficiency of the image training data used for model training.
[0059] In order to improve the correction efficiency of text recognition annotation, in an embodiment of the image text recognition information processing method of the present application, the above step S101 may further specifically include the following content:
[0060] Image text recognition is performed on the image to be annotated according to a preset image text recognition model to obtain text positioning information, text recognition information and confidence of the text recognition information. Text recognition and annotation correction is performed on the text recognition information of the image whose confidence is lower than the confidence threshold to obtain an image after text recognition and annotation correction.
[0061] Optionally, the present application may first use a preset (which may be an existing) image text recognition model to recognize the annotated image, output text location information Loc, text recognition information Rec and confidence Con of the text recognition information, and then correct the text recognition annotation of the low-confidence image based on the numerical comparison result of the confidence and the confidence threshold.
[0062] Optionally, the present application may adopt existing annotation correction methods to realize the correction of text recognition annotations, or may adopt manual correction methods to realize it, thereby obtaining an image after text recognition annotation correction.
[0063] In order to improve the efficiency of correcting text positioning annotation, in one embodiment of the image text recognition information processing method of the present application, see Figure 2 , the above step S102 may further specifically include the following contents:
[0064] Step S201: monitoring the recognition and annotation correction operation, obtaining an image after the recognition and annotation correction, and determining the character ratio of the corrected characters in the text recognition and annotation.
[0065] Step S202: If the character ratio of the corrected characters is higher than the character threshold, the text positioning annotation in the text positioning information of the image is corrected to obtain an image after the text positioning annotation is corrected.
[0066] Optionally, when performing text recognition and annotation correction, the present application can monitor the correction process to determine how many characters are corrected, and then determine the correction ratio.
[0067] Optionally, if the character ratio of the corrected characters is higher than the character threshold, it indicates that the image text recognition model is very inaccurate in recognizing the text of the image, and the inaccuracy of the text recognition information is strongly correlated with the inaccurate recognition of the text positioning information, that is, the character ratio of the corrected characters is high, indicating that its text positioning is also inaccurate.
[0068] Therefore, the present application performs text positioning and annotation correction on the text positioning information of the image that has undergone text recognition and annotation correction to obtain an image that has undergone text positioning and annotation correction.
[0069] In order to be able to make targeted corrections, in one embodiment of the image text recognition information processing method of the present application, see Figure 3 , and can also include the following:
[0070] Step S301: If the confidence level is within a preset confidence level range, a marking review operation is performed on the image after the text recognition and marking correction.
[0071] Step S302: If the character ratio is within a preset character ratio value range, a marking review operation is performed on the image after the text positioning marking correction.
[0072] Optionally, for images whose confidence or character ratio does not reach the threshold but whose representation is not very accurate, this application can perform a labeling review operation on them.
[0073] Optionally, the annotation review operation may be implemented by using an existing annotation review method or by manual review, thereby obtaining a reviewed image.
[0074] Optionally, after the annotation review is completed, the present application may update the text recognition annotation and / or text positioning annotation of the image according to the correction content of the annotation review operation.
[0075] Optionally, after the confidence level and character ratio are obtained by the above calculation, if the confidence level is higher than the confidence threshold and the character ratio is lower than the character threshold, it indicates that the image recognition is accurate, and the image corrected by text recognition and annotation and the image corrected by text positioning and annotation can be used as image training data for the neural network model.
[0076] In order to be able to effectively apply the neural network model, in one embodiment of the present application, after using the image marked with the corrected text recognition information and text positioning information as the image training data of the neural network model, the present application can further perform image text recognition training on the neural network model to obtain a trained neural network model.
[0077] In a specific example of a financial scenario, the neural network model can be used to perform image recognition on images of financial transaction credentials (such as bills) or identity credentials (such as bank cards, ID cards) to obtain image recognition information for transaction processing.
[0078] In order to effectively improve the labeling efficiency of image training data used for model training, this application provides an embodiment of an image text recognition information processing device for implementing all or part of the content of the image text recognition information processing method, see Figure 4 The image and text recognition information processing device specifically includes the following contents:
[0079] The initial recognition module 10 is used to perform image recognition on the image to be annotated according to a preset image and text recognition model to obtain text recognition information and text positioning information.
[0080] The recognition annotation correction module 20 is used to correct the character recognition annotation in the character recognition information according to the confidence level of the character recognition information, and determine the character ratio of the corrected characters in the character recognition annotation.
[0081] The positioning annotation correction module 30 is used to correct the text positioning annotation in the text positioning information according to the character proportion, obtain an image annotated with the corrected text recognition information and text positioning information, and use the image as image training data for the neural network model.
[0082] From the above description, it can be seen that the image text recognition information processing device provided in the embodiment of the present application can perform targeted corrections to the text recognition annotations based on the confidence level of the text annotation results of the image text recognition model, and then perform targeted corrections to the text positioning annotations based on the proportion of characters that need to be corrected when correcting the text recognition annotations, thereby accurately screening out the annotations that really need to be corrected and performing correction operations, thereby effectively improving the annotation efficiency of the image training data used for model training.
[0083] In order to improve the efficiency of correcting text recognition annotations, in one embodiment of the image text recognition information processing device of the present application, see Figure 5 , the positioning annotation correction module 30 includes:
[0084] The correction character ratio determination unit 31 is used to monitor the recognition and annotation correction operation, obtain the image after the recognition and annotation correction, and determine the character ratio of the correction character in the text recognition and annotation.
[0085] The high correction rate positioning correction unit 32 is configured to correct the text positioning annotation in the text positioning information of the image if the character ratio of the corrected character is higher than the character threshold, so as to obtain an image after the text positioning annotation correction.
[0086] In order to be able to make targeted corrections, in one embodiment of the image and text recognition information processing device of the present application, see Figure 6 , also includes:
[0087] The text review unit 41 is configured to perform a label review operation on the image after the text recognition and labeling correction if the confidence level is within a preset confidence level range.
[0088] The positioning review unit 42 is configured to perform a marking review operation on the image after the text positioning marking correction if the character ratio is within a preset character ratio value range.
[0089] In order to further illustrate this solution, this application also provides a specific application example of using the above-mentioned image and text recognition information processing device to implement the image and text recognition information processing method, see Figure 7 , specifically including the following contents:
[0090] (1) First, use the original image text recognition model to recognize the annotated image and output the text location information Loc, text recognition information Rec and the confidence level Con of the text recognition information.
[0091] (2) Determine whether the confidence level Con of the text recognition information is greater than 99%.
[0092] (3) If the above judgment result is True, the image and text recognition information are judged as not requiring attention; if the above judgment result is False, the subsequent processing is continued.
[0093] (4) Determine whether the confidence level Con of the text recognition information is greater than 95%.
[0094] (5) If the above judgment result is True, the image and text recognition information are judged to be in need of review; if the above judgment result is False, the image and text recognition information are judged to be in need of correction.
[0095] (6) Manual correction is used to correct the text recognition annotations. The system monitors the correction process to determine how many characters have been corrected, and then determines the correction ratio.
[0096] (7) Determine whether the correction ratio of text recognition annotation is less than 1%.
[0097] (8) If the above judgment result is True, the image and text positioning information are judged as not requiring attention; if the above judgment result is False, the subsequent processing is continued.
[0098] (9) Determine whether the correction ratio is less than 20%.
[0099] (10) If the above judgment result is True, the image and text positioning information is determined to be in need of review; if the above judgment result is False, the image and text positioning information is determined to be in need of correction.
[0100] (11) Manual correction is used to correct the text positioning annotation.
[0101] (12) After the annotation content is completed, an image containing the annotated text positioning information and text recognition information is obtained.
[0102] From the above content, it can be seen that this application does not require manual annotation of all images one by one, which greatly reduces the number of annotations, and does not require manual verification and review of the images recognized by the model one by one. Through a hierarchical approach, manual annotation / review of key images can be made, making more effective use of manual energy.
[0103] From a hardware perspective, in order to effectively improve the efficiency of labeling image training data used for model training, this application provides an embodiment of an electronic device for implementing all or part of the image text recognition information processing method. The electronic device specifically includes the following:
[0104] A processor, memory, a communications interface, and a bus; wherein the processor, memory, and communications interface communicate with each other via the bus; the communications interface is used to transmit information between the image and text recognition information processing device and related devices such as core business systems, user terminals, and related databases; the logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., but this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiments of the image and text recognition information processing method and the embodiments of the image and text recognition information processing device in the embodiments, and their contents are incorporated herein, and repeated parts are not repeated.
[0105] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.
[0106] In practical applications, portions of the image text recognition information processing method may be executed on the electronic device as described above, or all operations may be performed on the client device. The specific selection may be based on the processing capabilities of the client device and the limitations of the user's usage scenario. This application does not impose any restrictions on this. If all operations are performed on the client device, the client device may also include a processor.
[0107] The client device may include a communication module (i.e., a communication unit) that can establish a communication connection with a remote server to implement data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a server structure of a distributed device.
[0108] Figure 8 Schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Figure 8 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that the Figure 8 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.
[0109] In one embodiment, the image text recognition information processing method function may be integrated into the central processing unit 9100. The central processing unit 9100 may be configured to perform the following control:
[0110] Step S101: performing image recognition on the image to be annotated according to a preset image and text recognition model to obtain text recognition information and text positioning information.
[0111] Step S102: correcting the character recognition annotation in the character recognition information according to the confidence level of the character recognition information, and determining the character ratio of the corrected character in the character recognition annotation.
[0112] Step S103: Correct the text positioning annotation in the text positioning information according to the character proportion, obtain an image annotated with the corrected text recognition information and text positioning information, and use the image as image training data for the neural network model.
[0113] From the above description, it can be seen that the electronic device provided in the embodiment of the present application performs targeted corrections to the text recognition annotations by adjusting the confidence level of the text annotation results of the image text recognition model, and then performs targeted corrections to the text positioning annotations based on the proportion of characters that need to be corrected during the text recognition annotation correction, thereby accurately screening out the annotations that really need to be corrected and performing correction operations, thereby effectively improving the annotation efficiency of the image training data used for model training.
[0114] In another embodiment, the image and text recognition information processing device can be configured separately from the central processing unit 9100. For example, the image and text recognition information processing device can be configured as a chip connected to the central processing unit 9100, and the image and text recognition information processing method function can be implemented under the control of the central processing unit.
[0115] like Figure 8 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Figure 8 In addition, the electronic device 9600 may also include all components shown in Figure 8 For components not shown, reference may be made to the prior art.
[0116] like Figure 8 As shown, the central processing unit 9100 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives input and controls the operation of various components of the electronic device 9600.
[0117] Memory 9140 can be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It can store the aforementioned failure-related information and also store programs that execute the relevant information. The CPU 9100 can execute the programs stored in memory 9140 to implement information storage or processing.
[0118] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 may be, for example, a keypad or touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display objects such as images and text. The display may be, for example, an LCD display, but is not limited thereto.
[0119] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), or a SIM card. Alternatively, it may be a memory that retains information even when power is off, can be selectively erased, and is provided with more data. Examples of such memory are sometimes referred to as EPROMs. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs or processes for executing the operation of the electronic device 9600 by the central processing unit 9100.
[0120] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various driver programs for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).
[0121] The communication module 9110 is a transmitter / receiver 9110 that transmits and receives signals via an antenna 9111. The communication module (transmitter / receiver) 9110 is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as in a conventional mobile communication terminal.
[0122] Based on different communication technologies, multiple communication modules 9110 can be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module. The communication module (transmitter / receiver) 9110 is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide audio output via the speaker 9131 and receive audio input from the microphone 9132, thereby implementing common telecommunication functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Furthermore, the audio processor 9130 is also coupled to the central processing unit 9100, enabling local recording via the microphone 9132 and playback of stored audio via the speaker 9131.
[0123] The embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the image text recognition information processing method in the above-mentioned embodiments, where the execution subject is a server or a client. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the computer program implements all steps of the image text recognition information processing method in the above-mentioned embodiments, where the execution subject is a server or a client. For example, when the processor executes the computer program, the following steps are implemented:
[0124] Step S101: performing image recognition on the image to be annotated according to a preset image and text recognition model to obtain text recognition information and text positioning information.
[0125] Step S102: correcting the character recognition annotation in the character recognition information according to the confidence level of the character recognition information, and determining the character ratio of the corrected character in the character recognition annotation.
[0126] Step S103: Correct the text positioning annotation in the text positioning information according to the character proportion, obtain an image annotated with the corrected text recognition information and text positioning information, and use the image as image training data for the neural network model.
[0127] From the above description, it can be seen that the computer-readable storage medium provided in the embodiment of the present application performs targeted corrections to the text recognition annotations by adjusting the confidence level of the text annotation results of the image text recognition model, and then performs targeted corrections to the text positioning annotations based on the proportion of characters that need to be corrected when correcting the text recognition annotations, thereby accurately screening out the annotations that really need to be corrected and performing correction operations, thereby effectively improving the annotation efficiency of the image training data used for model training.
[0128] The embodiments of the present application also provide a computer program product capable of implementing all steps of the image text recognition information processing method in the above-mentioned embodiments, where the execution subject is a server or a client. When the computer program / instructions are executed by a processor, the steps of the image text recognition information processing method are implemented. For example, the computer program / instructions implement the following steps:
[0129] Step S101: performing image recognition on the image to be annotated according to a preset image and text recognition model to obtain text recognition information and text positioning information.
[0130] Step S102: correcting the character recognition annotation in the character recognition information according to the confidence level of the character recognition information, and determining the character ratio of the corrected character in the character recognition annotation.
[0131] Step S103: Correct the text positioning annotation in the text positioning information according to the character proportion, obtain an image annotated with the corrected text recognition information and text positioning information, and use the image as image training data for the neural network model.
[0132] From the above description, it can be seen that the computer program product provided in the embodiment of the present application performs targeted corrections to the text recognition annotations by adjusting the confidence level of the text annotation results of the image text recognition model, and then performs targeted corrections to the text positioning annotations based on the proportion of characters that need to be corrected when correcting the text recognition annotations, thereby accurately screening out the annotations that really need to be corrected and performing correction operations, thereby effectively improving the annotation efficiency of the image training data used for model training.
[0133] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0134] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as a combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0135] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0136] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0137] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
Claims
1. A method for processing image character recognition information, characterized in that: The method comprises: Perform image recognition on the image to be annotated according to a preset image and text recognition model to obtain text recognition information and text positioning information; Correcting the character recognition annotation in the character recognition information according to the confidence level of the character recognition information, and determining the proportion of the corrected characters in the character recognition annotation when correcting the character recognition annotation; Correcting the character positioning annotation in the character positioning information according to the character proportions to obtain an image annotated with the corrected character recognition information and character positioning information, and using the image as image training data for a neural network model; Among them, the inaccuracy of the text recognition information is correlated with the inaccurate recognition of the text positioning information. When the character ratio of the corrected characters is higher than the character threshold, the corresponding text positioning is inaccurate.
2. The image character recognition information processing method according to claim 1, characterized in that: The step of correcting the text positioning mark in the text positioning information according to the character proportion to obtain an image marked with the corrected text recognition information and text positioning information includes: Monitoring the recognition and annotation correction operation, obtaining an image after the recognition and annotation correction, and determining the proportion of the corrected characters in the text recognition and annotation; If the character ratio of the corrected characters is higher than the character threshold, the text positioning annotation in the text positioning information of the image is corrected to obtain an image after the text positioning annotation is corrected.
3. The image character recognition information processing method according to claim 2, characterized in that: Also includes: If the confidence level is within a preset confidence level range, performing a marking review operation on the image after identification and marking correction; If the character ratio is within a preset character ratio value range, a marking review operation is performed on the image after the text positioning marking correction.
4. The image character recognition information processing method according to claim 3, characterized in that: After the image that has undergone the identification mark correction and / or the image that has undergone the text positioning mark correction is marked and reviewed, the method further includes: The text recognition annotation and / or text positioning annotation is updated according to the correction content of the annotation review operation.
5. The image character recognition information processing method according to claim 2, characterized in that: Also includes: If the confidence level is higher than the confidence threshold and the character ratio is lower than the character threshold, the image corrected by the recognition annotation and the image corrected by the text positioning annotation are used as image training data for the neural network model.
6. The image character recognition information processing method according to claim 1, characterized in that: The method of using the image as image training data for the neural network model further includes: An image marked with the corrected text recognition information and text positioning information is used as image training data for a neural network model to perform image text recognition training on the neural network model to obtain a trained neural network model. The neural network model is used to perform image recognition on financial transaction credential images or identity credential images to obtain image recognition information for transaction processing.
7. An image character recognition information processing device, characterized in that: include: An initial recognition module is used to perform image recognition on the image to be annotated according to a preset image and text recognition model to obtain text recognition information and text positioning information; a recognition annotation correction module, configured to correct the character recognition annotation in the character recognition information according to the confidence level of the character recognition information, and determine the proportion of the corrected characters in the character recognition annotation during the correction; a positioning annotation correction module, configured to correct the text positioning annotation in the text positioning information according to the character proportion, obtain an image annotated with the corrected text recognition information and text positioning information, and use the image as image training data for the neural network model; Among them, the inaccuracy of the text recognition information is correlated with the inaccurate recognition of the text positioning information. When the character ratio of the corrected characters is higher than the character threshold, the corresponding text positioning is inaccurate.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the image character recognition information processing method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image character recognition information processing method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the image character recognition information processing method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
OCR model training method and device, computer equipment and storage medium
CN113537184A