Texture Background Character Recognition Method, System and Device Based on Image Inpainting
By applying image-based repair methods in industrial scenarios, restoring the texture background and denoising the character diagram, the problem that traditional algorithms are difficult to recognize characters under texture background is solved, and efficient and accurate character recognition and image repair effects are achieved.
Patent Information
- Application Number
- CN202510486895.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-18
AI Technical Summary
In industrial scenarios, traditional algorithms find it difficult to effectively identify product surface characters, especially in texture backgrounds, with high misidentification rates, which seriously affects production efficiency.
Through an image-based repair method, convolution processing, long-distance feature interaction processing and channel attention-based residual connection processing are used to restore texture background, denoise foreground character diagram, enhance character outline details, and add target background style, and finally convert the image into vector form for character recognition.
It improves image repair quality and character recognition accuracy and efficiency, effectively reduces the misrecognition rate under texture background, and improves visual perception quality.
Smart Images

Figure CN120014653B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image restoration, and specifically to a method, system and device for texture background character recognition based on image restoration. Background Art
[0002] Character recognition based on image restoration is a rigid requirement for industrial quality inspection and product traceability. In intelligent manufacturing, product surface characters, such as serial numbers, batch numbers, QR codes, etc., are the core carriers of quality traceability. Taking new energy vehicles as an example, the characters laser-engraved on the surface of battery cells have a very high manual re-inspection rate due to problems such as metal reflection and oxidation and fouling, seriously restricting production efficiency. At the same time, in the manufacturing of electronic products, the silk-screened characters of micro-components are tiny in size and the background is often a dense circuit texture. Traditional algorithms are difficult to distinguish characters from background patterns, and the misrecognition rate is much higher than that of products in a pure background, resulting in poor character recognition effects in industrial scenarios. Summary of the Invention
[0003] The purpose of the present invention is to provide a method, system and device for texture background character recognition based on image restoration.
[0004] The technical solution of the present invention is as follows:
[0005] A method for texture background character recognition based on image restoration includes the following operations:
[0006] S1. The image to be recognized undergoes convolutional processing with several times of gradually decreasing scales, long-distance feature interaction processing, and convolutional processing with several times of gradually increasing scales to obtain an initial feature map to be recognized; the initial feature map to be recognized undergoes convolutional processing with several times of gradually decreasing scales, several times of residual connection processing based on channel attention, and convolutional processing with several times of gradually increasing scales to obtain a texture background restoration map;
[0007] The operation of long-distance feature interaction processing is as follows: the convolutional image to be recognized undergoes several times of mask context information fusion processing based on a fixed window and mask updated context information fusion processing based on a moving window to obtain a long-distance feature interaction map, which is used for the operation of convolutional processing with several times of gradually increasing scales;
[0008] The operation of mask context information fusion processing based on a fixed window is specifically as follows: after the input undergoes mask multi-head attention processing based on a fixed window, it is cascaded with the input in terms of features, and then undergoes full connection processing and multi-layer perceptron processing to obtain an output; in the mask multi-head attention processing based on a fixed window, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the total sum of pixel mask values;
[0009] S2. Obtain the difference map between the image to be recognized and the texture background repair map to get the foreground character map; according to the texture direction in the texture background map, perform texture direction denoising processing on the foreground character map to obtain the denoised character map; after the denoised character map is subjected to contour enhancement processing, add the target background style to obtain the character map to be recognized.
[0010] The operation of texture direction denoising processing is as follows: based on the normalization value of each pixel point in the foreground character map and the corresponding color filter adaptive value, obtain their respective color filter mask values; based on the color filter mask value of each pixel point and the pixel value at the corresponding position in the texture direction variant template map, perform color filter processing on the texture background map to obtain the denoised character map.
[0011] S3. After the character map to be recognized is subjected to feature extraction and encoding processing, it is converted into a vector form to obtain the character feature sequence to be recognized; after the character feature sequence to be recognized is processed by a bidirectional long short-term memory network, merge the repeated characters in the prediction sequence and delete the blank markers to obtain the character recognition result.
[0012] The specific operation of the residual connection processing based on channel attention in S1 is as follows: the initial convolutional feature map to be recognized or the (N - 1)th residual connection feature map, after convolution, batch normalization, Relu activation function, convolution, batch normalization, and channel attention processing, are respectively fused with the initial convolutional feature map to be recognized or the (N - 1)th residual connection feature map to obtain the Nth residual connection feature map; the last residual connection feature map is used as the residual connection feature map to be recognized for performing the operation of several times of convolution processing with gradually increasing scales; the initial convolutional feature map to be recognized is obtained by performing several times of convolution processing with gradually decreasing scales on the initial feature map to be recognized.
[0013] In the channel attention processing in S1, the channel attention weight is obtained through the following formula:
[0014] ,
[0015] ,
[0016] ,
[0017] is the channel attention weight corresponding to the i th pixel position point, is the sigmoid function, is the initial channel attention weight of the i th pixel point within the neighborhood range of the j th pixel position point, is the i th pixel point within the neighborhood range of the j th pixel position point,J is the total number of pixels within the neighborhood of the i th pixel, is the one-dimensional convolution processing with a convolution kernel of K , is the channel dimension, is the convolution kernel coefficient, is the convolution kernel bias, is the nearest odd function.
[0018] The current head attention probability distribution coefficient in S1 is obtained through the following formula:
[0019] ,
[0020] is the i th head attention probability distribution coefficient, is the softmax function, is the i th head's query vector, is the transpose of the i th head's index vector, is the total number of pixels of the input image, is the sum of the pixel mask values of the input image. If there are invalid positions at the current window of the input image, is the mask map corresponding to the input image, is the i th head's content vector, is the dimension of the index vector.
[0021] The color filter mask value in S2 is obtained through the following formula:
[0022] When , the color filter mask value of the i th pixel = 1;
[0023] When , the color filter mask value of the i th pixel = 0;
[0024] ,
[0025] ,
[0026] is the normalized value of the i th pixel in the foreground character map, is the color filter adaptive threshold of the i th pixel in the foreground character map, is thei The pixel value of a pixel point and are respectively the maximum pixel value and the minimum pixel value in the foreground character map, is the gradient function.
[0027] In S2, the operation of performing color filtering on the texture background map is specifically as follows:
[0028] If the color filter mask value i of the th pixel point in the foreground character map = 1, then the color filtering is implemented through the following formula: ,
[0029] If the color filter mask value i of the th pixel point in the foreground character map = 0, then the color filtering is implemented through the following formula: ,
[0030] is the pixel value of the i th pixel point in the denoised character map, is the pixel value of the i th pixel point in the foreground character map, is the pixel value at the corresponding position of the i th pixel point in the texture direction variant template map.
[0031] In S3, the operation of feature extraction and encoding processing can be implemented through several times of convolution, batch normalization, Relu activation function, and max pooling processing.
[0032] A texture background character recognition system based on image inpainting, used to implement the above-mentioned texture background character recognition method based on image inpainting, includes:
[0033] The texture background repair map generation module is used to obtain the initial feature map to be recognized by performing convolutional processing with gradually decreasing scales several times, long-distance feature interaction processing, and convolutional processing with gradually increasing scales several times on the image to be recognized; the initial feature map to be recognized is subjected to convolutional processing with gradually decreasing scales several times, several residual connection processes based on channel attention, and convolutional processing with gradually increasing scales several times to obtain the texture background repair map; the operation of the long-distance feature interaction processing is: the convolutional image to be recognized is subjected to several mask context information fusion processes based on a fixed window and mask updated context information fusion processes based on a moving window to obtain a long-distance feature interaction map, which is used for the operation of performing convolutional processing with gradually increasing scales several times; the operation of the mask context information fusion process based on a fixed window is specifically: after the input is processed by mask multi-head attention based on a fixed window, it is feature-cascaded with the input, and then subjected to fully connected processing and multi-layer perceptron processing to obtain an output; in the mask multi-head attention processing based on a fixed window, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the total sum of pixel mask values;
[0034] The character map to be recognized generation module is used to obtain the foreground character map by obtaining the difference map between the image to be recognized and the texture background repair map; perform texture direction denoising processing on the foreground character map according to the texture direction in the texture background map to obtain the denoised character map; after the denoised character map is subjected to contour enhancement processing, add the target background style to obtain the character map to be recognized; the operation of the texture direction denoising processing is: based on the normalized value of each pixel point in the foreground character map and the corresponding color filter adaptive value, obtain their respective color filter mask values; based on the color filter mask value of each pixel point and the pixel value at the corresponding position in the texture direction variant template map, perform color filter processing on the texture background map to obtain the denoised character map;
[0035] The character recognition result generation module is used to convert the character map to be recognized into a vector form after feature extraction and encoding processing to obtain the character feature sequence to be recognized; the character feature sequence to be recognized is processed by a bidirectional long short-term memory network, and duplicate characters in the prediction sequence are merged and blank tokens are deleted to obtain the character recognition result.
[0036] A texture background character recognition device based on image repair includes a processor and a memory. Among them, when the processor executes the computer program saved in the memory, it implements the above-mentioned texture background character recognition method based on image repair.
[0037] A computer-readable storage medium is used to store a computer program. Among them, when the computer program is executed by a processor, it implements the above-mentioned texture background character recognition method based on image repair.
[0038] The beneficial effects of the present invention are as follows:
[0039] A method for character recognition of texture background based on image inpainting provided by the present invention first performs convolution processing, long-distance feature interaction processing, and residual connection processing based on channel attention on the image to be recognized, borrows visible information to fill the missing areas in the image, reduces color differences and blurriness, and maintains the structural rationality and texture consistency of the filled areas, so as to restore the texture background of the character-covered part and obtain a texture background inpainted image; then, obtains a foreground character image by subtracting the image to be recognized from the texture background inpainted image, performs denoising processing on the foreground character image to eliminate texture noise in all directions and obtain a denoised character image; enhances the character contour details in the denoised character image and then adds the target background style to eliminate the microscopic pixel-scale residual errors existing in the image inpainting process, effectively improves the visual perception quality while retaining the character topological structure, and obtains a character image to be recognized; finally, after performing feature extraction and encoding processing on the character image to be recognized, converts it into a vector form to obtain a character feature sequence to be recognized, and after processing the character feature sequence to be recognized by a bidirectional long short-term memory network, captures the context information of the character sequence based on the image from the front and back directions, merges the repeated characters in the prediction sequence, deletes the blank markers, and obtains a character recognition result; when applied to industrial scene image inpainting and character recognition, it can improve the image inpainting quality and the accuracy and efficiency of character recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] By reading the detailed description of the preferred embodiments below, the solutions and advantages of the present application will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention.
[0041] In the drawings:
[0042] Figure 1 FIG. is a comparison diagram of the image inpainting effects of the method of this embodiment and the existing method in the embodiment;
[0043] Figure 2 FIG. is the effect diagram of the character enhancement process in the embodiment; in Figure 2 , (a) is the image to be recognized, (b) is the texture background inpainted image, (c) is the foreground character image, (d) is the denoised character image obtained after denoising processing in the horizontal texture direction, (e) is the denoised character image obtained after denoising processing in the horizontal and vertical texture directions, (f) is the character image with enhanced contour, and (g) is the character image to be recognized;
[0044] Figure 3 FIG. is the character recognition effect diagram before and after removing the texture background in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings.
[0046] This embodiment provides a method for recognizing characters with a texture background based on image inpainting, including the following operations:
[0047] S1. The image to be recognized undergoes several times of convolutional processing with gradually decreasing scales, long-distance feature interaction processing, and several times of convolutional processing with gradually increasing scales to obtain an initial feature map to be recognized; the initial feature map to be recognized undergoes several times of convolutional processing with gradually decreasing scales, several times of residual connection processing based on channel attention, and several times of convolutional processing with gradually increasing scales to obtain a texture background inpainting map;
[0048] The operation of long-distance feature interaction processing is as follows: the convolutional image to be recognized undergoes several times of mask context information fusion processing based on a fixed window and mask updated context information fusion processing based on a moving window to obtain a long-distance feature interaction map, which is used to perform the operation of several times of convolutional processing with gradually increasing scales;
[0049] The operation of mask context information fusion processing based on a fixed window is specifically as follows: after the input undergoes mask multi-head attention processing based on a fixed window, it is feature-cascaded with the input, and then undergoes full connection processing and multi-layer perceptron processing to obtain an output; in the mask multi-head attention processing based on a fixed window, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the total sum of pixel mask values;
[0050] S2. Obtain a differential map of the image to be recognized and the texture background inpainting map to obtain a foreground character map; according to the texture direction in the texture background map, perform texture direction denoising processing on the foreground character map to obtain a denoised character map; after the denoised character map undergoes contour enhancement processing, add the target background style to obtain a character map to be recognized;
[0051] The operation of texture direction denoising processing is as follows: based on the normalization value of each pixel point in the foreground character map and the corresponding color filter adaptive value, obtain their respective color filter mask values; based on the color filter mask values of each pixel point and the pixel values at the corresponding positions in the texture direction variant template map, perform color filter processing on the texture background map to obtain a denoised character map;
[0052] S3. After the character map to be recognized undergoes feature extraction and encoding processing, it is converted into a vector form to obtain a character feature sequence to be recognized; after the character feature sequence to be recognized undergoes bidirectional long short-term memory network processing, merge the repeated characters in the prediction sequence and delete the blank markers to obtain the character recognition result.
[0053] S1. The image to be recognized undergoes several times of convolutional processing with gradually decreasing scales, long-distance feature interaction processing, and several times of convolutional processing with gradually increasing scales to obtain an initial feature map to be recognized; the initial feature map to be recognized undergoes several times of convolutional processing with gradually decreasing scales, several times of residual connection processing based on channel attention, and several times of convolutional processing with gradually increasing scales to obtain a texture background inpainting map.
[0054] The image to be recognized is processed through convolution, long-distance feature interaction, and residual connection based on channel attention, borrowing visible information to fill in the missing areas in the image, reducing color differences and blurriness, and maintaining the structural rationality and texture consistency of the filled areas, so as to restore the texture background of the character-covered part and obtain a texture background restoration map.
[0055] The above operation of obtaining the texture background restoration map is achieved by putting the image to be recognized into a training image restoration network for processing. Among them, the training image restoration network is obtained by training the training image restoration network with a training set formed by a number of original character images and character defect images.
[0056] The details of the processing process of the image to be recognized in the training image restoration network are as follows.
[0057] First, the image to be recognized undergoes several times of convolution processing with gradually decreasing scales, long-distance feature interaction processing, and several times of convolution processing with gradually increasing scales. After upsampling the spatial resolution to the input size, long-distance feature interaction of the image is performed to extract the detailed features of the texture background, obtaining an initial feature map to be recognized.
[0058] Among them, the operation of long-distance feature interaction processing is: the convolutional image to be recognized undergoes several times of mask context information fusion processing based on a fixed window and several times of mask updated context information fusion processing based on a moving window to obtain a long-distance feature interaction map, which is used to perform the operation of several times of convolution processing with gradually increasing scales.
[0059] Specifically, the convolutional image to be recognized undergoes mask context information fusion processing based on a fixed window, mask updated context information fusion processing based on a moving window, mask context information fusion processing based on a fixed window, mask updated context information fusion processing based on a moving window, and mask context information fusion processing based on a fixed window in sequence to obtain a long-distance feature interaction map, so as to obtain accurate and rich detailed features of the texture background.
[0060] The convolutional image to be recognized is obtained after the image to be recognized undergoes several times of convolution processing with gradually decreasing scales. In the mask context information fusion processing based on a fixed window, the shape and size of the fixed window that can slide through the input image remain unchanged; in the mask updated context information fusion processing based on a moving window, the shape and size of the moving window that can slide through the input image change, and the change form of the shape and size of the moving window can be set according to actual needs, or preferably refer to the moving window change form of the SW-MSA network (Window&Shifted Window based Self-Attention).
[0061] The operation of mask context information fusion processing based on a fixed window is as follows: The input (including the convolutional image to be recognized) is processed by mask multi-head attention based on a fixed window, and then feature concatenation is performed with the input (which can be achieved through a stacking operation). After fully connected processing and multi-layer perceptron processing, an output is obtained for performing mask updated context information fusion processing based on a moving window.
[0062] In the mask multi-head attention processing based on a fixed window, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the sum of pixel mask values, that is, based on the sum of the pixel mask values traversed by the fixed window during the sliding process and the total number of pixels of the input.
[0063] The current head attention probability distribution coefficient is obtained through the following formula:
[0064] ,
[0065] is the i th head attention probability distribution coefficient, is the softmax function, is the i th head's query vector, is the i transpose of the index vector of the th head, is the total number of pixel points of the input image, is the sum of pixel mask values of the input image, is the mask map corresponding to the input image, i is the content vector of the th head,
[0066] The purpose of mask updated context information fusion processing based on a moving window is to only identify the defective area of the input, that is, to identify the area to be filled, perform mask update, and achieve texture background pixel filling. The operation of mask updated context information fusion processing based on a moving window is as follows: The input is processed by mask multi-head attention based on a moving window (preferably the SW-MSA network), and then feature concatenation is performed with the input (which can be achieved through a stacking operation). After fully connected processing and multi-layer perceptron processing, an output is obtained for performing mask context information fusion processing based on a fixed window. During the mask updated context information fusion processing based on a moving window, if the current window token in the input image is valid, that is, there is a filling position, the mask values of all pixel points at the current window are 255. If the token at the current window in the input image is invalid, that is, there is no filling position, the mask values of all pixel points at the current window are 0.
[0067] Then, to perform a refinement filling process on the local area of the above-mentioned initial feature map to be recognized (coarse filling result), the initial feature map to be recognized is subjected to several convolutional processes with gradually decreasing scales, several residual connection processes based on channel attention, and several convolutional processes with gradually increasing scales, and the surrounding local information is used to appropriately repair the missing area to obtain a texture background repair map. For the repair effect diagram, please refer to Figure 1 。
[0068] Specifically, the operation of the residual connection process based on channel attention is as follows: The initial convolutional feature map to be recognized or the (N - 1)-th residual connection feature map is subjected to convolution, batch normalization, Relu activation function, convolution, batch normalization, and channel attention processing, and then respectively fused with the initial convolutional feature map to be recognized or the (N - 1)-th residual connection feature map to obtain the N-th residual connection feature map; The last residual connection feature map is used as the residual connection feature map to be recognized for performing the operation of several convolutional processes with gradually increasing scales; The initial convolutional feature map to be recognized is obtained by subjecting the initial feature map to be recognized to several convolutional processes with gradually decreasing scales.
[0069] In this embodiment, to improve the existing channel attention effect and capture local cross-channel interaction information by considering each channel and its neighboring channels, in the channel attention processing, the channel attention weight is obtained through the following formula:
[0070] ,
[0071] ,
[0072] ,
[0073] is the channel attention weight corresponding to the i -th pixel position point in the initial normalized feature map to be recognized or the (N - 1)-th residual connection normalized feature map, is the sigmoid function, is the i -th pixel position point, j is the initial channel attention weight of the -th pixel point within the neighborhood range of the i -th pixel position point, j is the pixel value of the J -th pixel point within the neighborhood range of the i -th pixel point, is the one-dimensional convolution process with a convolution kernel of K , is the channel dimension, which is a preset value, is the convolution kernel coefficient, is the convolution kernel bias, , is the nearest odd function. The initial normalized feature map to be recognized or the (N-1)-th residual connection normalized feature map is obtained by performing convolution, batch normalization, Relu activation function, convolution, and batch normalization on the initial convolutional feature map to be recognized or the (N-1)-th residual connection feature map respectively.
[0074] S2. Obtain the difference map between the image to be recognized and the texture background repair map to get the foreground character map; perform texture direction denoising on the foreground character map according to the texture direction in the texture background map to get the denoised character map; after the denoised character map is subjected to contour enhancement processing, add the target background style to get the character map to be recognized.
[0075] By subtracting the image to be recognized from the texture background repair map to obtain the foreground character map, perform denoising on the foreground character map to eliminate texture noise in all directions to get the denoised character map; enhance the character contour details in the denoised character map and then add the target background style to eliminate the microscopic pixel-scale residual error existing in the image repair process, effectively improving the visual perception quality while retaining the character topology structure to obtain the character map to be recognized, providing high-quality character data for subsequent character recognition.
[0076] First, obtain the difference map between the image to be recognized and the texture background repair map to get the foreground character map. Specifically, obtain the absolute difference matrix between the image to be recognized and the texture background repair map as the foreground character map. The effect diagram can be seen in Figure 2 (c) in.
[0077] Then, perform texture direction denoising on the foreground character map according to the texture direction in the texture background map to get the denoised character map. The effect diagram can be seen in Figure 2 (d) and (e) in.
[0078] The operation of texture direction denoising is as follows: based on the normalized value of each pixel point in the foreground character map and the corresponding color filter adaptive value, obtain their respective color filter mask values; based on the color filter mask value of each pixel point and the pixel value at the corresponding position in the texture direction variant template map, perform color filtering on the texture background map to get the denoised character map. The effect diagram can be seen in Figure 2 (f) in.
[0079] The color filter mask value is obtained through the following formula:
[0080] When , the color filter mask value i of the -th pixel point
[0081] When , the color filter mask value i of the = 0;
[0082] ,
[0083] ,
[0084] is the normalized value of the i th pixel point in the foreground character map, is the color filter adaptive threshold of the i th pixel point in the foreground character map, is the pixel value of the i th pixel point in the foreground character map, , are the maximum pixel value and the minimum pixel value in the foreground character map respectively, is the gradient function.
[0085] Meanwhile, in any direction, the operation of performing color filter processing on the texture background map is as follows: If the color filter mask value of the i th pixel point in the foreground character map = 1, then the color filter processing is achieved through the following formula: , if the color filter mask value of the i th pixel point in the foreground character map = 0, then the color filter processing is achieved through the following formula: , is the pixel value of the i th pixel point in the denoised character map, is the pixel value of the i th pixel point in the foreground character map, is the pixel value at the corresponding position of the i th pixel point in the texture direction variant template map.
[0086] The above-mentioned texture direction variant template map is obtained by performing an affine transformation on the texture background repair map containing only texture background information. The affine transformation angles are 0°, 45°, 90°, or 135°, corresponding to the horizontal direction, the positive 45° direction, the vertical direction, and the positive 135° direction respectively. The texture direction in the texture background map includes but is not limited to being obtained through the gray-level co-occurrence matrix in the texture background map. When performing texture direction denoising processing on the foreground character map, it includes but is not limited to performing denoising processing in one texture direction, and can also be the superposition of denoising processing in multiple directions.
[0087] Finally, after the denoised character map is subjected to contour enhancement processing, the target background style is added to obtain the character map to be recognized. The effect diagram can be seen in Figure 2 (g) in.
[0088] The operations of the contour enhancement process are specifically as follows: Create a dither matrix of size n×n, where the element values in the dither matrix are between 0 and n 2 −1, and each element value is unique; normalize the dither matrix to obtain a normalized dither matrix; divide the grayscale image of the denoised character image into blocks of size n × n , add the average grayscale value of each block to the element at the corresponding position in the normalized dither matrix to obtain an initial contour enhancement image; perform binarization on the initial contour enhancement image to obtain a contour enhancement character image.
[0089] The above addition of the target background style can be achieved by superimposing the target style standard image on the contour enhancement character image.
[0090] S3. After the character image to be recognized undergoes feature extraction and encoding processing, it is converted into a vector form to obtain a character feature sequence to be recognized; after the character feature sequence to be recognized undergoes processing by a bidirectional long short-term memory network, duplicate characters in the prediction sequence are merged, and blank markers are deleted to obtain a character recognition result.
[0091] After the character image to be recognized undergoes feature extraction and encoding processing, it is converted into a vector form to obtain a character feature sequence to be recognized; after the character feature sequence to be recognized undergoes processing by a bidirectional long short-term memory network, the context information of the character sequence based on the image is captured from the front and back directions, and then duplicate characters in the prediction sequence are merged, and blank markers are deleted to obtain a character recognition result.
[0092] First, after the character image to be recognized undergoes feature extraction and encoding processing, it is converted into a vector form to obtain a character feature sequence to be recognized.
[0093] Among them, the operations of the feature extraction and encoding processing can be implemented through several times of convolution, batch normalization, Relu activation function, and max pooling processing. Specifically, the operations of the feature extraction and encoding processing can sequentially pass through 6 times of convolution, batch normalization, Relu activation function, and max pooling processing, and 1 time of convolution, batch normalization, and Relu activation function. In the feature extraction and encoding processing, each convolution consists of a different number of kernels, and the number of kernels increases with the increase of the depth of the neural network. The more the number of kernels, the more deep features are extracted.
[0094] Then, after the character feature sequence to be recognized undergoes processing by a bidirectional long short-term memory network (which can be implemented by a BiLSTM network or a BiGRU), duplicate characters in the prediction sequence are merged, and blank markers are deleted to obtain a character recognition result. The effect diagram can be seen in Figure 3 .
[0095] This embodiment also provides a texture background character recognition system based on image restoration, which is used to implement the above-mentioned texture background character recognition method based on image restoration, and includes:
[0096] A texture background restoration map generation module, which is used to perform convolution processing with gradually decreasing scales, long-distance feature interaction processing, and convolution processing with gradually increasing scales on the image to be recognized to obtain an initial feature map to be recognized; the initial feature map to be recognized is subjected to convolution processing with gradually decreasing scales, several residual connection processes based on channel attention, and convolution processing with gradually increasing scales to obtain a texture background restoration map; the operation of long-distance feature interaction processing is: the convolution image to be recognized is subjected to several mask context information fusion processes based on a fixed window and mask updated context information fusion processes based on a moving window to obtain a long-distance feature interaction map, which is used to perform the operation of convolution processing with gradually increasing scales; the operation of the mask context information fusion process based on a fixed window is specifically: after the input is subjected to mask multi-head attention processing based on a fixed window, it is cascaded with the input at the feature level, and after full connection processing and multi-layer perceptron processing, an output is obtained; in the mask multi-head attention processing based on a fixed window, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the total sum of pixel mask values;
[0097] A character map to be recognized generation module, which is used to obtain a foreground character map by obtaining the difference map between the image to be recognized and the texture background restoration map; perform texture direction denoising processing on the foreground character map according to the texture direction in the texture background map to obtain a denoised character map; after the denoised character map is subjected to contour enhancement processing, add the target background style to obtain the character map to be recognized; the operation of texture direction denoising processing is: based on the normalized value of each pixel point of the foreground character map and the corresponding color filter adaptive value, obtain their respective color filter mask values; based on the color filter mask value of each pixel point and the pixel value at the corresponding position in the texture direction variant template map, perform color filter processing on the texture background map to obtain a denoised character map;
[0098] A character recognition result generation module, which is used to convert the character map to be recognized into a vector form after feature extraction and encoding processing to obtain a character feature sequence to be recognized; after the character feature sequence to be recognized is processed by a bidirectional long short-term memory network, merge the repeated characters in the prediction sequence and delete the blank markers to obtain the character recognition result.
[0099] This embodiment also provides a texture background character recognition device based on image restoration, which includes a processor and a memory. Among them, when the processor executes the computer program stored in the memory, it implements the above-mentioned texture background character recognition method based on image restoration.
[0100] This embodiment also provides a computer-readable storage medium for storing a computer program, wherein when the computer program is executed by a processor, the above-mentioned texture background character recognition method based on image restoration is implemented.
[0101] A texture background character recognition method based on image restoration provided in this embodiment first performs convolution processing, long-distance feature interaction processing, and residual connection processing based on channel attention on the image to be recognized, borrows visible information to fill the missing areas in the image, reduces color differences and blurriness, and maintains the structural rationality and texture consistency of the filled areas, thereby restoring the texture background of the character-covered part to obtain a texture background restoration map; then, obtains a foreground character map by subtracting the image to be recognized from the texture background restoration map, performs denoising processing on the foreground character map to eliminate texture noise in all directions to obtain a denoised character map; enhances the character contour details in the denoised character map and then adds the target background style to eliminate the microscopic pixel-scale residual error existing in the image restoration process, effectively improving the visual perception quality while retaining the character topology structure to obtain the character map to be recognized; finally, after the character map to be recognized is processed by feature extraction and encoding, it is converted into a vector form to obtain a character feature sequence to be recognized. The character feature sequence to be recognized is processed by a bidirectional long short-term memory network to capture the context information of the character sequence based on the image in the front and back directions, merge the repeated characters in the prediction sequence, and delete the blank markers to obtain the character recognition result; when applied to industrial scene image restoration and character recognition, it can improve the image restoration quality and the accuracy and efficiency of character recognition.
[0102] A texture background character recognition method based on image restoration provided in this embodiment adopts fusion learning in long-distance feature interaction processing, solves the problem of unstable training caused by too large a proportion of invalid tokens, and encourages learning low-frequency basic features to better learn high-frequency details subsequently and reduce the difficulty of optimization; the attention processing adopts a strategy of fixed window and mask update, and uses effective tokens to borrow visible information to fill the holes, reducing color differences and blurriness.
[0103] A texture background character recognition method based on image restoration provided in this embodiment uses the surrounding local information to appropriately repair some missing areas in the residual connection processing based on channel attention, and at the same time uses residual blocks to accelerate the speed of the image restoration model and improve the image restoration performance.
[0104] A texture background character recognition method based on image restoration provided in this embodiment realizes the connection between the two stages of image restoration and character recognition, and eliminates the non-negligible error existing at the microscopic pixel scale of the image restoration network; by adding the target background style after the contour enhancement processing, it effectively improves the visual perception quality while retaining the character topology structure, providing high-quality character data for subsequent character recognition.
Claims
1. A texture background character recognition method based on image restoration, characterized in that: The following operations are included: S1. The image to be identified is subjected to several convolution processes with successively reduced scales, long-distance feature interaction processes, and several convolution processes with successively increased scales to obtain an initial feature map to be identified; the initial feature map to be identified is subjected to several convolution processes with successively reduced scales, several residual connection processes based on channel attention, and several convolution processes with successively increased scales to obtain a texture background restoration map; The operation of long-distance feature interaction processing is as follows: the convolution image to be identified is subjected to several fixed-window-based mask context information fusion processing and moving-window-based mask update context information fusion processing to obtain a long-distance feature interaction map, which is used to perform several convolution processing operations with successively increasing scales; The operation of the fixed window-based mask context information fusion processing is as follows: after the input is processed by the fixed window-based mask multi-head attention, it is feature concatenated with the input, and processed by the full connection processing and the multi-layer perceptron to obtain the output; in the fixed window-based mask multi-head attention processing, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the sum of the pixel mask values; S2, obtaining a difference image between the image to be recognized and the texture background restoration image to obtain a foreground character image; performing texture direction denoising on the foreground character image according to the texture direction in the texture background image to obtain a denoised character image; After the de-noised character image is processed by contour enhancement, the target background style is added to obtain the character image to be recognized; The operation of texture direction denoising is as follows: based on the normalized value of each pixel of the foreground character image and the corresponding filter adaptive value, obtain the respective filter mask value; Based on the color filter mask value of each pixel and the pixel value of the corresponding position in the texture direction variant template image, the texture background image is subjected to color filter processing to obtain a denoised character image; S3. After feature extraction and encoding, the character image to be recognized is converted into a vector form to obtain a feature sequence of the character to be recognized; after the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network, repeated characters in the prediction sequence are merged, and blank marks are deleted to obtain the character recognition result.
2. The texture background character recognition method based on image restoration according to claim 1 is characterized in that: In S1, the operation of residual connection processing based on channel attention is specifically as follows: The initial convolution feature map to be identified or the N-1th residual connection feature map is processed by convolution, batch normalization, Relu activation function, convolution, batch normalization and channel attention, and then fused with the initial convolution feature map to be identified or the N-1th residual connection feature map to obtain the Nth residual connection feature map; the last residual connection feature map is used as the residual connection feature map to be identified, and is used to perform several convolution operations with successively increasing scales; The initial convolution feature map to be identified is obtained by performing convolution processing on the initial feature map to be identified with the scales successively reduced for several times.
3. The texture background character recognition method based on image restoration according to claim 1, characterized in that: In the channel attention processing in S1, the channel attention weight is obtained by the following formula: , , , For the i The channel attention weight corresponding to the pixel position point, is the sigmoid function, For the i The pixel location point in the neighborhood j The initial channel attention weight of pixels, For the i The pixel location point in the neighborhood j The pixel value of a pixel, J For the i The total number of pixels in the neighborhood of a pixel. The convolution kernel is K One-dimensional convolution processing, is the channel dimension, is the convolution kernel coefficient, is the convolution kernel bias, is the nearest odd function.
4. The texture background character recognition method based on image restoration according to claim 1, characterized in that: In S1, the current head attention probability distribution coefficient is obtained by the following formula: , For the i The probability distribution coefficient of individual attention, is the softmax function, For the i The query vector of the head, For the i The transpose of the index vector of the head, is the total number of pixels of the input image, is the sum of the pixel mask values of the input image. If there is an invalid position in the current window in the input image, is the mask map corresponding to the input image, For the i The content vector of the header, is the dimension of the index vector.
5. The texture background character recognition method based on image restoration according to claim 1, characterized in that: In S2, the color filter mask value is obtained by the following formula: when At that time, i The filter mask value of each pixel =1; when At that time, i The filter mask value of each pixel =0; , , Foreground character map i The normalized value of pixels, Foreground character map i The color filter adaptive threshold of pixels, Foreground character map i The pixel value of a pixel, , are the maximum pixel value and the minimum pixel value in the foreground character image, respectively. is the gradient function.
6. The texture background character recognition method based on image restoration according to claim 1, characterized in that: In S2, the operation of performing color filtering on the texture background image is specifically as follows: If the foreground character image i The filter mask value of each pixel =1, the color filtering is achieved by the following formula: , If the foreground character image i The filter mask value of each pixel =0, the color filtering process is implemented by the following formula: , The denoised character image i The pixel value of a pixel, Foreground character map i The pixel value of a pixel, For the i The pixel value of the corresponding position of a pixel in the texture direction variant template image.
7. The texture background character recognition method based on image restoration according to claim 1 is characterized in that: In S3, the feature extraction and encoding processing operations can be implemented through several times of convolution, batch normalization, ReLU activation function and maximum pooling processing.
8. A texture background character recognition system based on image restoration, used to implement the texture background character recognition method based on image restoration according to claim 1, characterized in that: include: The texture background restoration image generation module is used to obtain the initial feature map to be recognized by subjecting the image to be recognized to several convolution processes with successively reduced scales, long-distance feature interaction processes, and several convolution processes with successively increased scales; the initial feature map to be recognized is subjected to several convolution processes with successively reduced scales, several residual connection processes based on channel attention, and several convolution processes with successively increased scales to obtain the texture background restoration image; The operation of long-distance feature interaction processing is as follows: the convolution image to be identified is subjected to several fixed-window-based mask context information fusion processing and moving-window-based mask update context information fusion processing to obtain a long-distance feature interaction map, which is used to perform several convolution processing operations with successively increasing scales; The operation of the fixed window-based mask context information fusion processing is as follows: after the input is processed by the fixed window-based mask multi-head attention, it is feature concatenated with the input, and processed by the full connection processing and the multi-layer perceptron to obtain the output; in the fixed window-based mask multi-head attention processing, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the sum of the pixel mask values; The module for generating a character image to be recognized is used to obtain a difference image between the image to be recognized and the texture background restoration image to obtain a foreground character image; and to perform texture direction denoising on the foreground character image according to the texture direction in the texture background image to obtain a denoised character image. After the denoised character image is processed by contour enhancement, the target background style is added to obtain the character image to be recognized; the operation of texture direction denoising is: based on the normalized value of each pixel of the foreground character image and the corresponding filter color adaptation value, the respective filter color mask values are obtained; Based on the color filter mask value of each pixel and the pixel value of the corresponding position in the texture direction variant template image, the texture background image is subjected to color filter processing to obtain a denoised character image; The character recognition result generation module is used to convert the character image to be recognized into a vector form after feature extraction and encoding processing, so as to obtain the feature sequence of the character to be recognized; after the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network, the repeated characters in the prediction sequence are merged, and the blank marks are deleted to obtain the character recognition result.
9. A texture background character recognition device based on image restoration, characterized in that: The method comprises a processor and a memory, wherein the processor implements the texture background character recognition method based on image restoration as described in any one of claims 1 to 7 when executing the computer program stored in the memory.
10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the texture background character recognition method based on image restoration as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Character recognition system based on Gabor convolution and linear sparse attention
CN113221874A
Inclined character recognition method based on complex background image
CN115439857A