Fuzzy character recognition method and device, equipment and storage medium
By improving the fast regional convolutional neural network model and the character recognition model, the problem of difficult recognition of fuzzy characters in financial institutions has been solved, achieving efficient fuzzy character recognition and improving business processing efficiency.
Patent Information
- Application Number
- CN202511010659.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-31
AI Technical Summary
In financial institutions, blurry text captured in images is difficult to recognize, leading to longer business processes and reduced processing efficiency.
An improved fast region convolutional neural network model is used to train a fuzzy text region recognition model. By combining a densely connected convolutional network and a feature pyramid network, and using soft nonmaximum suppression and region of interest alignment algorithms, the accuracy of fuzzy text detection is improved. A text recognition model consisting of a densely connected convolutional network, a bidirectional long short-term memory network, and a decoder is used to process fuzzy text region images and achieve target text recognition.
It improves the accuracy of recognizing fuzzy text, reduces the number of manual reviews, and enhances the processing efficiency of financial institutions' business.
Smart Images

Figure CN120877296A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for fuzzy text recognition. Background Technology
[0002] The widespread use of image acquisition equipment has made it easier for people to obtain the text information they need. However, due to the complexity of the natural scenes in which the text is located, the text in the acquired images is often blurry and difficult to recognize, requiring manual verification. Furthermore, repeated image submissions by customers prolong the business process and affect the processing efficiency of financial institutions, primarily banks.
[0003] Therefore, accurately identifying blurry text in images involved in the business operations of financial institutions is crucial for banks and urgently needs to be addressed. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and storage medium for recognizing fuzzy text, thereby improving the accuracy of recognizing fuzzy text in images and thus improving the processing efficiency of financial institutions' business.
[0005] According to one aspect of the present invention, a method for recognizing fuzzy characters is provided, the method comprising:
[0006] Obtain the image to be detected that contains blurred text;
[0007] The image to be detected is input into the fuzzy text region recognition model to obtain the fuzzy text region image in the image to be detected; wherein, the fuzzy text region recognition model is obtained by training an improved fast region convolutional neural network model;
[0008] The blurred text region image is input into the trained text recognition model to obtain the target text in the blurred text region image.
[0009] According to another aspect of the present invention, a fuzzy character recognition device is provided, the device comprising:
[0010] The image acquisition module is used to acquire images containing blurred text.
[0011] The text region image determination module is used to input the image to be detected into the fuzzy text region recognition model to obtain the fuzzy text region image in the image to be detected; wherein, the fuzzy text region recognition model is obtained by training an improved fast region convolutional neural network model;
[0012] The target text determination module is used to input the blurred text region image into the trained text recognition model to obtain the target text in the blurred text region image.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] At least one processor; and
[0015] A memory that is communicatively connected to at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the fuzzy character recognition method of any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the fuzzy character recognition method of any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the fuzzy character recognition method of any embodiment of the present invention.
[0019] The technical solution of this invention involves acquiring an image to be detected containing fuzzy text; inputting the image to be detected into a fuzzy text region recognition model to obtain an image of the fuzzy text region in the image to be detected; wherein, the fuzzy text region recognition model is obtained by training an improved fast region convolutional neural network model; and the fuzzy text region image is input into a trained text recognition model to obtain the target text in the fuzzy text region image. This technical solution, combining the fuzzy text region recognition model and the text recognition model, achieves accurate recognition of fuzzy text in the image to be detected, improves the recognition accuracy of fuzzy text in the image to be detected, thereby reducing the number of times manual review of fuzzy text in images involved in financial institution business, reducing the number of times customers submit images in financial institution business, and thus improving the processing efficiency of financial institution business.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a fuzzy character recognition method provided in Embodiment 1 of the present invention;
[0023] Figure 2A This is a flowchart of a fuzzy character recognition method provided in Embodiment 2 of the present invention;
[0024] Figure 2B This is a schematic diagram of the structure of a fuzzy text region recognition model provided in Embodiment 2 of the present invention;
[0025] Figure 3 This is a schematic diagram of the structure of a fuzzy character recognition device according to Embodiment 3 of the present invention;
[0026] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the fuzzy character recognition method of this invention. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "target," "first," and "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] Example 1
[0030] Figure 1This is a flowchart of a fuzzy text recognition method provided in Embodiment 1 of the present invention. This embodiment can be applied to the recognition of fuzzy text in images involved in financial institution business (such as identity verification, transaction voucher submission, and contract signing). The method can be executed by a fuzzy text recognition device, which can be implemented in hardware and / or software and can be configured in an electronic device.
[0031] like Figure 1 As shown, the method includes:
[0032] S101. Obtain the image to be detected containing blurred text.
[0033] The image to be detected refers to the image on which blurred text needs to be recognized. For example, the image to be detected can be a blurred ID photo, a blurred contract signing image, a blurred transaction voucher image, or a blurred income certificate image, etc.
[0034] Specifically, images containing blurred text can be obtained from the materials submitted by the client.
[0035] S102. Input the image to be detected into the fuzzy text region recognition model to obtain the fuzzy text region image in the image to be detected.
[0036] The blurred text region image refers to the image in the image to be detected that consists of regions containing blurred text. It should be noted that the blurred text region image is a part of the image to be detected; it is a sub-image of the image to be detected.
[0037] The fuzzy text region recognition model was trained using an improved Fast Region-based Convolutional Neural Network (Faster R-CNN) model. The improvements to the Faster R-CNN model include: replacing the Residual Neural Network (ResNet) with Densely Connected Convolutional Networks (DenseNet) and Feature Pyramid Network (FPN) to improve the model's ability to detect small targets; and replacing the Non-maximum Suppression (NMS) algorithm with Soft Non-maximum Suppression (SoftNMS). The Non-Maximum Suppression (Soft-NMS) algorithm addresses the issue of missed detections of dense targets and improves the model's ability to detect small targets. The Region of Interest (ROI) pooling algorithm in the fast region convolutional neural network model is replaced with the RolPool algorithm to resolve the spatial precision loss caused by quantization in RIO pooling, preserving more spatial precision and improving target detection accuracy. A channel feature enhancement layer is added to the fast region convolutional neural network model to further enhance target detection accuracy and improve the model's ability to detect small targets. The Gaussian weighting function of the Soft-NMS algorithm is as follows:
[0038]
[0039] Among them, s i This represents the confidence score of the current detection box; σ represents a control parameter used to determine the steepness of the decay curve, with a value of 0.5, i.e., σ = 0.5; b i b is the current detection bounding box. j Represents the bounding box with the highest confidence score; IoU(b i ,b j ) represents the intersection-union ratio between the current detection box and the detection box with the highest confidence score.
[0040] Optionally, to further enhance target coverage and improve the model's ability to detect small targets, six aspect ratios can be set for the anchor boxes in the improved fast region convolutional neural network model, namely {1:5, 1:3, 1:2, 2:1, 3:1, 5:1}. A single-scale anchor box can be set for each layer of the feature pyramid network in the improved fast region convolutional neural network model, that is, the anchor box scale corresponding to P2 is 8px×8px, the anchor box scale corresponding to P3 is 16px×16px, the anchor box scale corresponding to P4 is 32px×32px, and the anchor box scale corresponding to P5 is 64px×64px. In addition, three scaling ratios are added to each layer of the feature pyramid network, namely 0.5, 1.0, and 2.0.
[0041] Optionally, to balance model accuracy and speed, the 3×3 convolutions after layers P2, P3, and P4 of the feature pyramid network in the improved fast region convolutional neural network model can be replaced with deformable convolutions.
[0042] Specifically, the fuzzy text region recognition model is obtained by training the improved fast region convolutional neural network model. This can be achieved by training the improved fast region convolutional neural network model with a large number of sample images containing fuzzy text until the model training loss decreases to a preset range, at which point training is stopped, thus obtaining the fuzzy text region recognition model. The preset range can be pre-set based on actual business needs or the experience of those skilled in the art; this embodiment of the invention does not specifically limit it.
[0043] Specifically, the image to be detected is input into the fuzzy text region recognition model. After processing by the fuzzy text region recognition model, the fuzzy text region image in the image to be detected is obtained.
[0044] Optionally, before inputting the image to be detected into the fuzzy text region recognition model, image preprocessing can be performed on the image to obtain a preprocessed image to improve its image quality. Specifically, image preprocessing can involve: using an image enhancement algorithm, such as CLAHE (Contrast Limited Adaptive Histogram Equalization), to enhance the local contrast of the image to obtain a first optimized image; using a Wiener filtering algorithm to filter the first optimized image to remove noise and enhance its clarity, thus obtaining a second optimized image; and using a blind deconvolution algorithm, such as the Lucy-Richardson algorithm, to iteratively refine the second optimized image to obtain the preprocessed image to be detected.
[0045] Optionally, to improve the deblurring effect of the image to be detected and to retain more details during image preprocessing, a trained deep learning model can be used to preprocess the image to be detected, resulting in a preprocessed image. The trained deep learning model can be pre-configured according to actual business needs; this embodiment of the invention does not impose specific limitations on it.
[0046] S103. Input the blurred text region image into the trained text recognition model to obtain the target text in the blurred text region image.
[0047] The character recognition model can consist of a densely connected convolutional network, a bidirectional long short-term memory (BiLSTM) network, and a decoder. The BiLSTM network is an improved recurrent neural network specifically designed for processing sequential data. It should be noted that the decoder refers to a Transformer. Optionally, the character recognition model can also be based on the RobustScanner framework, with the corresponding loss function being: L total =L CTC +λL Attention ; among which, L total L represents the overall training loss of the character recognition model. CTC This represents the CTC (Connectionist Temporal Classification) loss, i.e., the sequence modeling loss; L Attentionλ represents the attention mechanism loss; λ represents a hyperparameter used to balance the contributions of CTC loss and attention mechanism loss, and its value is set to 0.3, i.e., λ = 0.3. Target text refers to the text in the blurred text region image output by the text recognition model.
[0048] The training process for the character recognition model is as follows: A clear dataset consisting of a large number of clear text images is generated using a training dataset synthesis tool. Motion blurring and Gaussian blurring are applied to the clear text images in the clear dataset, along with salt-and-pepper noise and image brightness or contrast adjustments, resulting in a blurred dataset with blurred features, noise, and illumination variations. A progressive model training approach is adopted. First, the character recognition model is trained using the clear dataset containing only clear text images. Then, 10% of the clear text images in the clear dataset are replaced with corresponding blurred images from the blurred dataset, and the character recognition model is trained again, until 50% of the clear text images in the clear dataset are replaced with corresponding blurred images from the blurred dataset, at which point no further processing is performed on the clear dataset. Finally, the character recognition model is trained using the clear dataset containing 50% blurred images until the training loss is reduced to a preset range, at which point training of the character recognition model is stopped, resulting in a trained character recognition model.
[0049] Specifically, the blurred text region image is input into the trained text recognition model. After processing by the trained text recognition model, the target text in the blurred text region image is obtained.
[0050] The technical solution of this invention involves acquiring an image to be detected containing fuzzy text; inputting the image to be detected into a fuzzy text region recognition model to obtain an image of the fuzzy text region in the image to be detected; wherein, the fuzzy text region recognition model is obtained by training an improved fast region convolutional neural network model; and inputting the fuzzy text region image into a character recognition model to obtain the target text in the fuzzy text region image. This technical solution, combining the fuzzy text region recognition model and the character recognition model, achieves accurate recognition of fuzzy text in the image to be detected, improves the recognition accuracy of fuzzy text in the image to be detected, thereby reducing the number of times manual review of fuzzy text in images involved in financial institution operations, reducing the number of times customers submit images in financial institution operations, and thus improving the processing efficiency of financial institution operations.
[0051] Example 2
[0052] Figure 2AThis is a flowchart of a fuzzy text recognition method provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment further optimizes the step of "inputting the image to be detected into the fuzzy text region recognition model to obtain the fuzzy text region image in the image to be detected", and provides an optional implementation scheme.
[0053] It should be noted that the fuzzy text region recognition model includes a feature extraction layer, a candidate region generation layer, a candidate region selection layer, a candidate region pooling layer, a channel feature enhancement layer, and an object detection layer. See also... Figure 2B The connection relationships between the layers in the fuzzy text region recognition model are as follows: the feature extraction layer is connected to the candidate region generation layer; the candidate region generation layer is connected to the candidate region filtering layer; the candidate region pooling layer is connected to both the feature extraction layer and the candidate region filtering layer; the channel feature enhancement layer is connected to the candidate region pooling layer; and the target detection layer is connected to the channel feature enhancement layer.
[0054] It should also be noted that for parts not described in detail in the embodiments of the present invention, reference can be made to the relevant descriptions in other embodiments. For example... Figure 2A As shown, the method includes:
[0055] S201. Obtain the image to be detected containing blurred text.
[0056] S202. The feature extraction layer extracts features from the image to be detected to obtain a multi-scale feature map.
[0057] The feature extraction layer consists of a densely connected convolutional network and a feature pyramid network. The densely connected convolutional network is DenseNet-121, meaning it comprises 121 convolutional layers. The multi-scale feature map refers to the feature map output by the feature extraction layer of the fuzzy text region recognition model.
[0058] Specifically, a densely connected convolutional network can be used to extract features from the image to be detected, obtaining an initial feature map; a feature pyramid network can then be used to perform multi-scale fusion processing on the initial feature map, resulting in a multi-scale feature map. Here, the initial feature map refers to the feature map output by the densely connected convolutional network.
[0059] More specifically, the image to be detected is input into the feature extraction layer of the fuzzy text region recognition model. First, the densely connected convolutional network in the feature extraction layer extracts features from the image to be detected to obtain an initial feature map. Then, the initial feature map is used as the input to the feature pyramid network in the feature extraction layer. The feature pyramid network performs multi-scale fusion processing on the initial feature map to obtain a multi-scale feature map.
[0060] Understandably, by using densely connected convolutional networks and feature pyramid networks to extract features from the image to be detected, the fuzzy text region recognition model can extract feature maps containing rich semantic information and fusing features at different scales from the image to be detected, thereby improving the fuzzy text region recognition model's ability to detect regions containing fuzzy text.
[0061] S203. The multi-scale feature map is processed through the candidate region generation layer to obtain multiple candidate regions.
[0062] The candidate region generation layer consists of a Region Proposal Network (RPN). A candidate region is a region in the image to be detected that may contain text.
[0063] Specifically, the multi-scale feature map output by the feature extraction layer is used as the input to the candidate region generation layer. The region generation network in the candidate region generation layer processes the multi-scale feature map to obtain multiple candidate regions.
[0064] S204. The multiple candidate regions obtained are filtered through the candidate region filtering layer to obtain the filtered candidate regions.
[0065] Specifically, the soft nonmaximum suppression algorithm in the candidate region filtering layer is used to filter the multiple candidate regions to obtain the filtered candidate regions. More specifically, the multiple candidate regions output by the candidate region generation layer are used as input to the candidate region filtering layer, and the soft nonmaximum suppression algorithm in the candidate region filtering layer is used to filter the multiple candidate regions to obtain the filtered candidate regions.
[0066] Understandably, by using the soft nonmaximum suppression algorithm in the candidate region filtering layer to filter multiple candidate regions output by the candidate region generation layer, the problem of missing detection in densely texted areas of the image to be detected is solved, while preserving the diversity of candidate regions.
[0067] S205. The candidate regions and multi-scale feature maps after screening are processed by the candidate region pooling layer to obtain multiple pooled feature maps.
[0068] In this context, a pooling feature map refers to a sub-feature map segmented from a multi-scale feature map by the filtered candidate regions. It should be noted that one filtered candidate region corresponds to one pooling feature map.
[0069] Specifically, the candidate regions and multi-scale feature maps are processed by the region of interest alignment algorithm in the candidate region pooling layer to obtain multiple pooled feature maps.
[0070] More specifically, the filtered candidate regions and multi-scale feature maps are input into the candidate region pooling layer. Through the region of interest alignment algorithm in the candidate region pooling layer, the coordinates of the filtered candidate regions are mapped from the image to be detected onto the multi-scale feature map. And through block pooling operation, a pooling feature map of a fixed size is generated for each filtered candidate region.
[0071] S206. The channel feature enhancement layer is used to enhance the channel features of the obtained pooling feature maps to obtain the enhanced feature maps corresponding to each pooling feature map.
[0072] The channel feature enhancement layer consists of an SE module. The SE module (Squeeze-and-Excitation module) is a lightweight channel attention mechanism whose core idea is to dynamically adjust channel weights, enhance the feature responses of important channels, and suppress irrelevant channels. It should be noted that the SE module includes a compression unit, an activation unit, and a recalibration unit. The enhanced feature map refers to the feature map obtained after channel feature enhancement processing of the pooled feature map.
[0073] Specifically, for each pooling feature map, the compression unit in the channel feature enhancement layer performs global average pooling on each channel of the pooling feature map to obtain the channel vector corresponding to each channel in the pooling feature map; then, the activation unit in the channel feature enhancement layer learns the nonlinear relationship between the obtained channel vectors to obtain the channel weight corresponding to each channel in the pooling feature map; finally, the recalibration unit in the channel feature enhancement layer multiplies the channel vector corresponding to each channel in the pooling feature map with the channel weight to obtain the enhanced feature map corresponding to the pooling feature map.
[0074] Understandably, by dynamically adjusting the channel feature weights of each pooling feature map through the channel feature enhancement layer, the fuzzy text region recognition model enhances its ability to perceive fuzzy text features, thereby improving its ability to detect fuzzy text.
[0075] S207. Target detection is performed on the enhanced feature maps corresponding to each pooling feature map through the target detection layer to obtain the blurred text region image in the image to be detected.
[0076] Specifically, the enhanced feature maps corresponding to each pooling feature map are input into the target detection layer. The target detection layer identifies the enhanced feature maps corresponding to each pooling feature map and outputs feature maps containing blurred text, thereby obtaining the blurred text region image in the image to be detected.
[0077] S208. Input the blurred text region image into the trained text recognition model to obtain the target text in the blurred text region image.
[0078] The technical solution of this invention involves: acquiring a detection image containing blurred text; extracting features from the detection image using a feature extraction layer to obtain multi-scale feature maps; processing the multi-scale feature maps using a candidate region generation layer to obtain multiple candidate regions; filtering the multiple candidate regions using a candidate region filtering layer to obtain filtered candidate regions; processing the filtered candidate regions and the multi-scale feature maps using a candidate region pooling layer to obtain multiple pooled feature maps; enhancing the channel features of the multiple pooled feature maps using a channel feature enhancement layer to obtain enhanced feature maps corresponding to each pooled feature map; performing target detection on the enhanced feature maps corresponding to each pooled feature map using a target detection layer to obtain a blurred text region image in the detection image; and inputting the blurred text region image into a trained text recognition model to obtain the target text in the blurred text region image. The above technical solution further refines the processing of the image to be detected by the fuzzy text region recognition model. Through the cooperation between the layers in the fuzzy text region recognition model, the accurate recognition of fuzzy text in the image to be detected is achieved, improving the recognition accuracy of fuzzy text in the image to be detected. This reduces the number of times fuzzy text in images involved in the business of financial institutions, reduces the number of times customers submit images in the business of financial institutions, and thus improves the processing efficiency of financial institutions.
[0079] Example 3
[0080] Figure 3 This is a schematic diagram of a fuzzy text recognition device provided in Embodiment 3 of the present invention. This embodiment is applicable to the recognition of fuzzy text in images involved in financial institution business (such as identity verification, transaction document submission, and contract signing). The device can be implemented in hardware and / or software and can be configured in an electronic device. Figure 3 As shown, the device includes:
[0081] The image acquisition module 301 is used to acquire an image to be detected containing blurred text;
[0082] The text region image determination module 302 is used to input the image to be detected into the fuzzy text region recognition model to obtain the fuzzy text region image in the image to be detected; wherein, the fuzzy text region recognition model is obtained by training an improved fast region convolutional neural network model;
[0083] The target text determination module 303 is used to input the blurred text region image into the trained text recognition model to obtain the target text in the blurred text region image.
[0084] The technical solution of this invention involves acquiring an image to be detected containing fuzzy text; inputting the image to be detected into a fuzzy text region recognition model to obtain an image of the fuzzy text region in the image to be detected; wherein, the fuzzy text region recognition model is obtained by training an improved fast region convolutional neural network model; and the fuzzy text region image is input into a trained text recognition model to obtain the target text in the fuzzy text region image. This technical solution, combining the fuzzy text region recognition model and the text recognition model, achieves accurate recognition of fuzzy text in the image to be detected, improves the recognition accuracy of fuzzy text in the image to be detected, thereby reducing the number of times manual review of fuzzy text in images involved in financial institution business, reducing the number of times customers submit images in financial institution business, and thus improving the processing efficiency of financial institution business.
[0085] Optionally, the fuzzy text region recognition model includes a feature extraction layer, a candidate region generation layer, a candidate region filtering layer, a candidate region pooling layer, a channel feature enhancement layer, and a target detection layer;
[0086] The text region image determination module 302 includes:
[0087] The multi-scale feature map determination unit is used to extract features from the image to be detected through the feature extraction layer to obtain a multi-scale feature map.
[0088] The candidate region determination unit is used to process the multi-scale feature map through the candidate region generation layer to obtain multiple candidate regions;
[0089] The candidate region filtering unit is used to filter multiple candidate regions obtained through the candidate region filtering layer to obtain the filtered candidate regions.
[0090] The pooling feature map determination unit is used to process the filtered candidate regions and multi-scale feature maps through the candidate region pooling layer to obtain multiple pooling feature maps.
[0091] The enhanced feature map determination unit is used to perform channel feature enhancement processing on the multiple pooled feature maps obtained through the channel feature enhancement layer to obtain the enhanced feature map corresponding to each pooled feature map;
[0092] The text region image determination unit is used to perform target detection on the enhanced feature maps corresponding to each pooling feature map through the target detection layer, so as to obtain the blurred text region image in the image to be detected.
[0093] Optionally, the feature extraction layer consists of a densely connected convolutional network and a feature pyramid network;
[0094] The multi-scale feature map determination unit is specifically used for:
[0095] An initial feature map is obtained by extracting features from the image to be detected using a densely connected convolutional network.
[0096] The initial feature map is fused at multiple scales using a feature pyramid network to obtain a multi-scale feature map.
[0097] Optional, candidate region filtering unit, specifically used for:
[0098] The soft nonmaximum suppression algorithm in the candidate region filtering layer is used to filter the multiple candidate regions to obtain the filtered candidate regions.
[0099] Optionally, the pooling feature map determination unit is specifically used for:
[0100] By using the region of interest alignment algorithm in the candidate region pooling layer, the filtered candidate regions and multi-scale feature maps are processed to obtain multiple pooled feature maps.
[0101] Optionally, the device may also include:
[0102] The image preprocessing module is used to preprocess the image to be detected before inputting it into the fuzzy text region recognition model, so as to obtain the preprocessed image to be detected.
[0103] The fuzzy character recognition device provided in the embodiments of the present invention can execute the fuzzy character recognition method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing each fuzzy character recognition method.
[0104] According to embodiments of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product.
[0105] Example 4
[0106] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0107] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0108] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0109] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as fuzzy character recognition methods.
[0110] In some embodiments, the fuzzy character recognition method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the fuzzy character recognition method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the fuzzy character recognition method by any other suitable means (e.g., by means of firmware).
[0111] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0112] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0113] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0114] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0115] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0116] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0117] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0118] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for recognizing fuzzy characters, characterized in that, include: Obtain the image to be detected that contains blurred text; The image to be detected is input into the fuzzy text region recognition model to obtain the fuzzy text region image in the image to be detected; wherein, the fuzzy text region recognition model is obtained by training an improved fast region convolutional neural network model; The blurred text region image is input into the trained text recognition model to obtain the target text in the blurred text region image.
2. The method according to claim 1, characterized in that, The fuzzy text region recognition model includes a feature extraction layer, a candidate region generation layer, a candidate region filtering layer, a candidate region pooling layer, a channel feature enhancement layer, and a target detection layer; The feature extraction layer is used to extract features from the image to be detected to obtain a multi-scale feature map. The multi-scale feature map is processed by the candidate region generation layer to obtain multiple candidate regions; The candidate region filtering layer filters the multiple candidate regions to obtain the filtered candidate regions. The candidate region and the multi-scale feature map are processed by the candidate region pooling layer to obtain multiple pooled feature maps; The channel feature enhancement layer is used to enhance the channel features of the multiple pooling feature maps to obtain the enhanced feature map corresponding to each pooling feature map. The target detection layer performs target detection on the enhanced feature maps corresponding to each pooling feature map to obtain the blurred text region image in the image to be detected.
3. The method according to claim 2, characterized in that, The feature extraction layer consists of a densely connected convolutional network and a feature pyramid network; The step of extracting features from the image to be detected through the feature extraction layer to obtain a multi-scale feature map includes: The image to be detected is used to extract features through the densely connected convolutional network to obtain an initial feature map; The initial feature map is fused using the feature pyramid network to obtain a multi-scale feature map.
4. The method according to claim 2, characterized in that, The step of filtering multiple candidate regions through the candidate region filtering layer to obtain filtered candidate regions includes: The soft nonmaximum suppression algorithm in the candidate region filtering layer is used to filter the multiple candidate regions to obtain the filtered candidate regions.
5. The method according to claim 2, characterized in that, The process of processing the filtered candidate regions and the multi-scale feature maps through the candidate region pooling layer yields multiple pooled feature maps, including: The candidate regions and the multi-scale feature maps are processed by the region of interest alignment algorithm in the candidate region pooling layer to obtain multiple pooled feature maps.
6. The method according to claim 1, characterized in that, Before inputting the image to be detected into the fuzzy text region recognition model, the method further includes: The image to be detected is preprocessed to obtain the preprocessed image to be detected.
7. A fuzzy character recognition device, characterized in that, include: The image acquisition module is used to acquire images containing blurred text. The text region image determination module is used to input the image to be detected into the fuzzy text region recognition model to obtain the fuzzy text region image in the image to be detected; wherein, the fuzzy text region recognition model is obtained by training an improved fast region convolutional neural network model; The target text determination module is used to input the blurred text region image into the trained text recognition model to obtain the target text in the blurred text region image.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the fuzzy character recognition method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the fuzzy character recognition method according to any one of claims 1-6.
10. A computer program product comprising a computer program that, when executed by a processor, implements the fuzzy character recognition method according to any one of claims 1-6.