Texture background character recognition method, system and equipment based on image restoration

Through image repair methods, the texture background is restored and character diagrams are denoised, which solves the problem that traditional algorithms are difficult to recognize characters under texture backgrounds, and achieves efficient and accurate character recognition and image repair effects.

CN120014653AActive Publication Date: 2025-05-16SHANDONG UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510486895.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In industrial scenarios, traditional algorithms find it difficult to effectively identify product surface characters, especially in texture backgrounds, with a high misidentification rate, which seriously affects production efficiency.

Method used

Through an image-based repair method, convolution processing, long-distance feature interaction processing and channel attention-based residual connection processing are used to restore texture background, denoise foreground character diagram, enhance character outline details, and add target background style, and finally convert the image into vector form for character recognition.

Benefits of technology

It improves image repair quality and character recognition accuracy and efficiency, effectively reduces the misrecognition rate under texture background, and improves visual perception quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014653A_ABST
    Figure CN120014653A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image restoration, in particular to a texture background character recognition method, system and equipment based on image restoration. According to the method, a to-be-recognized image is subjected to convolution processing, long-distance feature interaction processing and residual connection processing based on channel attention, so that a texture background of a character coverage part is recovered, and a texture background repair image is obtained; secondly, subtracting a to-be-recognized image from the texture background restoration image, enhancing a character contour after eliminating texture noise in each direction, and adding a target background style to obtain a to-be-recognized character image; and finally, performing feature extraction coding processing on the character graph to be recognized, converting the character graph to be recognized into a sequence, performing bidirectional long-short-term memory network processing, merging repeated characters in the prediction sequence, deleting blank marks, and obtaining a character recognition result. The method is applied to industrial scene image restoration and character recognition, and the image restoration quality and the accuracy and efficiency of character recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image restoration, and in particular to a texture background character recognition method, system and device based on image restoration. Background Art

[0002] Character recognition based on image restoration is a rigid requirement for industrial quality inspection and product traceability. In intelligent manufacturing, characters on the surface of products, such as serial numbers, batch numbers, QR codes, etc., are the core carriers of quality traceability. Taking new energy vehicles as an example, the characters laser engraved on the surface of battery cells have a very high manual re-inspection rate due to problems such as metal reflection and oxidation contamination, which seriously restricts production efficiency. At the same time, the silk-screen characters of micro-components in the manufacturing of electronic products are small in size and the background is often dense circuit textures. Traditional algorithms have difficulty distinguishing between characters and background patterns, and the misrecognition rate is far higher than that of products in pure backgrounds, making character recognition effects poor in industrial scenarios. Summary of the invention

[0003] The purpose of the present invention is to provide a texture background character recognition method, system and device based on image restoration.

[0004] The technical solution of the present invention is as follows: A texture background character recognition method based on image restoration includes the following operations: S1. The image to be identified is subjected to several convolution processes with successively reduced scales, long-distance feature interaction processes, and several convolution processes with successively increased scales to obtain an initial feature map to be identified; the initial feature map to be identified is subjected to several convolution processes with successively reduced scales, several residual connection processes based on channel attention, and several convolution processes with successively increased scales to obtain a texture background restoration map; The operation of long-distance feature interaction processing is as follows: the convolution image to be identified is subjected to several fixed-window-based mask context information fusion processing and moving-window-based mask update context information fusion processing to obtain a long-distance feature interaction map, which is used to perform several convolution processing operations with successively increasing scales; The operation of the fixed window-based mask context information fusion processing is as follows: after the input is processed by the fixed window-based mask multi-head attention, it is feature concatenated with the input, and processed by the full connection processing and the multi-layer perceptron to obtain the output; in the fixed window-based mask multi-head attention processing, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the sum of the pixel mask values; S2, obtaining a difference image between the image to be recognized and the texture background restoration image to obtain a foreground character image; performing texture direction denoising on the foreground character image according to the texture direction in the texture background image to obtain a denoised character image; after contour enhancement processing on the denoised character image, adding the target background style to obtain a character image to be recognized; The operation of texture direction denoising is as follows: based on the normalized value of each pixel of the foreground character image and the corresponding color filter adaptive value, the respective color filter mask values ​​are obtained; based on the color filter mask value of each pixel and the pixel value of the corresponding position in the texture direction variant template image, the texture background image is subjected to color filter processing to obtain a denoised character image; S3. After feature extraction and encoding, the character image to be recognized is converted into a vector form to obtain a feature sequence of the character to be recognized; after the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network, repeated characters in the prediction sequence are merged, and blank marks are deleted to obtain the character recognition result.

[0005] The operation of residual connection processing based on channel attention in S1 is specifically as follows: the initial convolution feature map to be identified or the N-1th residual connection feature map is processed by convolution, batch normalization, Relu activation function, convolution, batch normalization and channel attention, and then fused with the initial convolution feature map to be identified or the N-1th residual connection feature map to obtain the Nth residual connection feature map; the last residual connection feature map is used as the residual connection feature map to be identified, and is used to perform several convolution operations with successively increasing scales; the initial convolution feature map to be identified is obtained by performing several convolution operations with successively decreasing scales on the initial feature map to be identified.

[0006] In the channel attention processing in S1, the channel attention weight is obtained by the following formula: , , , For the i The channel attention weight corresponding to the pixel position point, is the sigmoid function, For the i The pixel location point in the neighborhood j The initial channel attention weight of pixels, For the i The pixel location point in the neighborhood j The pixel value of a pixel, J For the i The total number of pixels in the neighborhood of a pixel. The convolution kernel is K One-dimensional convolution processing, is the channel dimension, is the convolution kernel coefficient, is the convolution kernel bias, is the nearest odd function.

[0007] The current head attention probability distribution coefficient in S1 is obtained by the following formula: , For the i The probability distribution coefficient of individual attention, is the softmax function, For the i The query vector of the head, For the i The transpose of the index vector of the head, is the total number of pixels of the input image, is the sum of the pixel mask values ​​of the input image. If there is an invalid position in the current window in the input image, is the mask map corresponding to the input image, For the i The content vector of the header, is the dimension of the index vector.

[0008] The color filter mask value in S2 is obtained by the following formula: when At that time, i The filter mask value of each pixel =1; when At that time, i The filter mask value of each pixel =0; , , Foreground character map i The normalized value of pixels, Foreground character map i The color filter adaptive threshold of pixels, Foreground character map i The pixel value of a pixel, , are the maximum pixel value and the minimum pixel value in the foreground character image, respectively. is the gradient function.

[0009] In S2, the operation of color filtering the texture background image is specifically as follows: If the foreground character image i The filter mask value of each pixel =1, the color filtering process is implemented by the following formula: , If the foreground character image i The filter mask value of each pixel =0, the color filtering process is implemented by the following formula: , The denoised character image i The pixel value of a pixel, Foreground character map i The pixel value of a pixel, For the i The pixel value of the corresponding position of a pixel in the texture direction variant template image.

[0010] In S3, the feature extraction and encoding processing operations can be achieved through several convolutions, batch normalization, ReLU activation function and maximum pooling processing.

[0011] A texture background character recognition system based on image restoration, used to implement the above-mentioned texture background character recognition method based on image restoration, comprising: The texture background repair map generation module is used for the image to be identified to obtain the initial feature map to be identified after several convolution processes with successively reduced scales, long-distance feature interaction processes, and several convolution processes with successively increased scales; the initial feature map to be identified is subjected to several convolution processes with successively reduced scales, several residual connection processes based on channel attention, and several convolution processes with successively increased scales to obtain the texture background repair map; the operation of the long-distance feature interaction process is: the convolution image to be identified is subjected to several mask context information fusion processes based on fixed windows and mask update context information fusion processes based on moving windows to obtain a long-distance feature interaction map, which is used to perform several operations of successively increased scale convolution processes; the operation of the mask context information fusion process based on fixed windows is specifically: after the input is processed by the mask multi-head attention based on fixed windows, it is feature concatenated with the input, and is processed by full connection and multi-layer perceptron to obtain the output; in the mask multi-head attention process based on fixed windows, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the sum of the pixel mask values; The module for generating a character image to be recognized is used to obtain a difference image between the image to be recognized and the texture background restoration image to obtain a foreground character image; according to the texture direction in the texture background image, the foreground character image is subjected to texture direction denoising to obtain a denoised character image; after the denoised character image is subjected to contour enhancement processing, the target background style is added to obtain a character image to be recognized; the operation of texture direction denoising is as follows: based on the normalized value of each pixel point of the foreground character image and the corresponding color filter adaptive value, the respective color filter mask values ​​are obtained; based on the color filter mask value of each pixel point and the pixel value of the corresponding position in the texture direction variant template image, the texture background image is subjected to color filter processing to obtain a denoised character image; The character recognition result generation module is used to convert the character image to be recognized into a vector form after feature extraction and encoding processing, so as to obtain the feature sequence of the character to be recognized; after the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network, the repeated characters in the prediction sequence are merged, and the blank marks are deleted to obtain the character recognition result.

[0012] A texture background character recognition device based on image restoration comprises a processor and a memory, wherein the processor implements the above-mentioned texture background character recognition method based on image restoration when executing a computer program stored in the memory.

[0013] A computer-readable storage medium is used to store a computer program, wherein the computer program implements the above-mentioned texture background character recognition method based on image restoration when executed by a processor.

[0014] The beneficial effects of the present invention are: The invention provides a texture background character recognition method based on image restoration. First, the image to be recognized is subjected to convolution processing, long-distance feature interaction processing, and residual connection processing based on channel attention, and the missing area in the image is filled with visible information to reduce color difference and blur, and maintain the structural rationality and texture consistency of the filled area, so as to restore the texture background of the character covering part and obtain a texture background restoration image; then, a foreground character image is obtained by subtracting the image to be recognized from the texture background restoration image, and the foreground character image is denoised to eliminate texture noise in all directions to obtain a denoised character image; after enhancing the character contour details in the denoised character image, The target background style is added to eliminate the micro-pixel scale residual error in the image restoration process, effectively improve the visual perception quality while retaining the character topology, and obtain the character map to be recognized; finally, the character map to be recognized is converted into a vector form after feature extraction and encoding, and the feature sequence of the character to be recognized is obtained. After the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network, the context information of the image-based character sequence is captured from the front and back directions, the repeated characters in the predicted sequence are merged, and the blank marks are deleted to obtain the character recognition result; it is applied to image restoration and character recognition in industrial scenes to improve the quality of image restoration as well as the accuracy and efficiency of character recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] By reading the detailed description of the preferred embodiment below, the scheme and advantages of the present application will become clear to those skilled in the art. The accompanying drawings are only for the purpose of illustrating the preferred embodiment and are not to be considered as limiting the present invention.

[0016] In the attached picture: Figure 1 : is a comparison diagram of the image restoration effect between the method of this embodiment and the existing method in the embodiment; Figure 2 In the embodiment, the character enhancement process effect diagram of this embodiment is shown; Figure 2 In the figure, (a) is the image to be recognized, (b) is the texture background restoration image, (c) is the foreground character image, (d) is the denoised character image obtained after denoising in the horizontal texture direction, (e) is the denoised character image obtained after denoising in the horizontal and vertical texture directions, (f) is the contour enhanced character image, and (g) is the character image to be recognized; Figure 3 This is a diagram showing the character recognition effect before and after removing the texture background in this embodiment. DETAILED DESCRIPTION

[0017] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings.

[0018] This embodiment provides a texture background character recognition method based on image restoration, including the following operations: S1. The image to be identified is subjected to several convolution processes with successively reduced scales, long-distance feature interaction processes, and several convolution processes with successively increased scales to obtain an initial feature map to be identified; the initial feature map to be identified is subjected to several convolution processes with successively reduced scales, several residual connection processes based on channel attention, and several convolution processes with successively increased scales to obtain a texture background restoration map; The operation of long-distance feature interaction processing is as follows: the convolution image to be identified is subjected to several fixed-window-based mask context information fusion processing and moving-window-based mask update context information fusion processing to obtain a long-distance feature interaction map, which is used to perform several convolution processing operations with successively increasing scales; The operation of the fixed window-based mask context information fusion processing is as follows: after the input is processed by the fixed window-based mask multi-head attention, it is feature concatenated with the input, and processed by the full connection processing and the multi-layer perceptron to obtain the output; in the fixed window-based mask multi-head attention processing, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the sum of the pixel mask values; S2, obtaining a difference image between the image to be recognized and the texture background restoration image to obtain a foreground character image; performing texture direction denoising on the foreground character image according to the texture direction in the texture background image to obtain a denoised character image; after contour enhancement processing on the denoised character image, adding the target background style to obtain a character image to be recognized; The operation of texture direction denoising is as follows: based on the normalized value of each pixel of the foreground character image and the corresponding color filter adaptive value, the respective color filter mask values ​​are obtained; based on the color filter mask value of each pixel and the pixel value of the corresponding position in the texture direction variant template image, the texture background image is subjected to color filter processing to obtain a denoised character image; S3. After feature extraction and encoding, the character image to be recognized is converted into a vector form to obtain a feature sequence of the character to be recognized; after the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network, repeated characters in the prediction sequence are merged, and blank marks are deleted to obtain the character recognition result.

[0019] S1. The image to be identified is processed by several convolutions with successively decreasing scales, long-distance feature interaction processing, and several convolutions with successively increasing scales to obtain an initial feature map to be identified; the initial feature map to be identified is processed by several convolutions with successively decreasing scales, several residual connection processing based on channel attention, and several convolutions with successively increasing scales to obtain a texture background restoration map.

[0020] The image to be recognized is processed by convolution, long-distance feature interaction, and residual connection based on channel attention. Visible information is used to fill the missing areas in the image, reduce color differences and blur, and maintain the structural rationality and texture consistency of the filled area, so as to restore the texture background of the part covered by the characters and obtain a texture background repair image.

[0021] The above operation of obtaining the texture background restoration image is achieved by placing the image to be recognized into a training image restoration network for processing. The training image restoration network is obtained by training the image restoration network with a training set formed by a number of original character images and character defect images.

[0022] The details of the processing process of the image to be identified in the training image restoration network are as follows.

[0023] First, the image to be identified undergoes several convolution processes with successively decreasing scales, long-distance feature interaction processes, and several convolution processes with successively increasing scales. After the spatial resolution is upsampled to the input size, long-distance feature interaction of the image is performed to extract texture background detail features and obtain the initial feature map to be identified.

[0024] Among them, the operation of long-distance feature interaction processing is: the convolution image to be identified undergoes several times of mask context information fusion processing based on a fixed window and several times of mask update context information fusion processing based on a moving window to obtain a long-distance feature interaction map, which is used to perform several times of convolution processing operations with successively increasing scales.

[0025] Specifically, the convolution image to be identified is sequentially processed by mask context information fusion based on fixed window, mask update context information fusion based on moving window, mask context information fusion based on fixed window, mask update context information fusion based on moving window, and mask context information fusion based on fixed window to obtain a long-distance feature interaction map, so as to obtain accurate and rich texture background detail features.

[0026] The convolution image to be identified is obtained after the image to be identified has been subjected to several convolution processes with successively reduced scales. In the mask context information fusion process based on the fixed window, the shape and size of the fixed window that can slide through the input image remain unchanged; in the mask update context information fusion process based on the moving window, the shape and size of the moving window that can slide through the input image are variable. The shape and size variation of the moving window can be set according to actual needs, or the moving window variation of the SW-MSA network (Window & Shifted Window based Self-Attention) can be preferentially referred to.

[0027] The specific operation of the fixed window-based mask context information fusion processing is as follows: the input (including the convolution image to be identified) is processed by the fixed window-based mask multi-head attention, and then feature concatenated with the input (which can be achieved through superposition operation), and then processed by the fully connected processing and the multi-layer perceptron to obtain the output, which is used to perform the moving window-based mask update context information fusion processing.

[0028] In the fixed window-based masked multi-head attention processing, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the sum of the pixel mask values, that is, the sum of the pixel mask values ​​traversed by the fixed window during the sliding process and the total number of input pixels.

[0029] The current head attention probability distribution coefficient is obtained by the following formula: , For the i The probability distribution coefficient of individual attention, is the softmax function, For the i The query vector of the head, For the i The transpose of the index vector of the head, is the total number of pixels of the input image, is the sum of the pixel mask values ​​of the input image, is the mask map corresponding to the input image, For the i The content vector of the header, is the dimension of the index vector.

[0030] The purpose of the mask update context information fusion processing based on the moving window is to identify only the defective area of ​​the input, that is, to identify the area to be filled, update the mask, and realize the filling of texture background pixels. The specific operation of the mask update context information fusion processing based on the moving window is: after the input is processed by the mask multi-head attention based on the moving window (preferential SW-MSA network), the feature cascade is performed with the input (which can be achieved through superposition operation), and the output is obtained through full connection processing and multi-layer perceptron processing, which is used to perform the mask context information fusion processing based on the fixed window. In the process of mask update context information fusion processing based on the moving window, if the current window token in the input image is valid, that is, there is a filling position, then the mask value of all pixels at the current window is 255. If the token at the current window in the input image is invalid, that is, there is no filling position, then the mask value of all pixels at the current window is 0.

[0031] Then, in order to perform a fine filling process on the local area of ​​the above-mentioned initial feature map to be identified (coarse filling result), the initial feature map to be identified is processed by several convolutions with successively reduced scales, several residual connection processes based on channel attention, and several convolutions with successively increased scales. The surrounding local information is used to properly repair the missing area to obtain a texture background repair map. The repair effect map is shown in Figure 1 .

[0032] Among them, the operation of residual connection processing based on channel attention is specifically as follows: the initial convolution feature map to be identified or the N-1th residual connection feature map is processed by convolution, batch normalization, Relu activation function, convolution, batch normalization and channel attention, and then fused with the initial convolution feature map to be identified or the N-1th residual connection feature map to obtain the Nth residual connection feature map; the last residual connection feature map is used as the residual connection feature map to be identified, and is used to perform several convolution operations with successively increasing scales; the initial convolution feature map to be identified is obtained by performing several convolution operations with successively decreasing scales on the initial feature map to be identified.

[0033] In order to improve the existing channel attention effect, this embodiment considers each channel and the neighboring channels to capture the local cross-channel interaction information. Therefore, in the channel attention processing, the channel attention weight is obtained by the following formula: , , , is the initial normalized feature map to be identified or the N-1th residual connection normalized feature map i The channel attention weight corresponding to the pixel position point, is the sigmoid function, For the i The pixel location point in the neighborhood j The initial channel attention weight of pixels, For the i The pixel location point in the neighborhood j The pixel value of a pixel, J For the i The total number of pixels in the neighborhood of a pixel. The convolution kernel is K One-dimensional convolution processing, is the channel dimension, is the preset value, is the convolution kernel coefficient, is the convolution kernel bias, , is the nearest odd function. The initial normalized feature map to be identified or the N-1th residual connection normalized feature map is obtained by processing the initial convolution feature map to be identified or the N-1th residual connection feature map through convolution, batch normalization, Relu activation function, convolution, and batch normalization.

[0034] S2. Obtain a difference image between the image to be recognized and the texture background restoration image to obtain a foreground character image; perform texture direction denoising on the foreground character image according to the texture direction in the texture background image to obtain a denoised character image; after contour enhancement processing on the denoised character image, add the target background style to obtain a character image to be recognized.

[0035] The foreground character map is obtained by subtracting the image to be recognized from the texture background restoration map, and the foreground character map is denoised to eliminate texture noise in all directions to obtain a denoised character map. After enhancing the character contour details in the denoised character map, the target background style is added to eliminate the micro-pixel scale residual error in the image restoration process, effectively improving the visual perception quality while retaining the character topological structure, and obtaining the character map to be recognized, providing high-quality character data for subsequent character recognition.

[0036] First, obtain the difference map between the image to be recognized and the texture background restoration map to obtain the foreground character map. Specifically, obtain the absolute difference matrix between the image to be recognized and the texture background restoration map as the foreground character map. The effect map can be seen in Figure 2 (c) in.

[0037] Then, according to the texture direction in the texture background image, the foreground character image is subjected to texture direction denoising to obtain a denoised character image. The effect image can be seen in Figure 2 (d) and (e) in the figure.

[0038] The operation of texture direction denoising is as follows: based on the normalized value of each pixel of the foreground character image and the corresponding filter adaptive value, obtain the respective filter mask value; based on the filter mask value of each pixel and the pixel value of the corresponding position in the texture direction variant template image, perform filter processing on the texture background image to obtain the denoised character image. The effect image can be seen in Figure 2 (f) in.

[0039] The filter mask value is obtained by the following formula: when At that time, i The filter mask value of each pixel =1; when At that time, i The filter mask value of each pixel =0; , , Foreground character map i The normalized value of pixels, Foreground character map i The color filter adaptive threshold of pixels, Foreground character map i The pixel value of a pixel, , are the maximum pixel value and the minimum pixel value in the foreground character image, respectively. is the gradient function.

[0040] At the same time, in any direction, the color filtering operation of the texture background image is as follows: if the foreground character image i The filter mask value of each pixel =1, the color filtering process is implemented by the following formula: , if the foreground character image i The filter mask value of each pixel =0, the color filtering process is implemented by the following formula: , The denoised character image i The pixel value of a pixel, Foreground character map i The pixel value of a pixel, For the i The pixel value of the corresponding position of a pixel in the texture direction variant template image.

[0041] The texture direction variant template image is obtained by affine transforming the texture background restoration image containing only texture background information, and the affine transformation angle is 0°, or 45°, or 90°, or 135°, corresponding to the horizontal direction, the positive 45° direction, the vertical direction, and the positive 135° direction, respectively. The texture direction in the texture background image includes but is not limited to being obtained through the grayscale co-occurrence matrix in the texture background image, and when the foreground character image is subjected to texture direction denoising, it includes but is not limited to denoising in one texture direction, and can also be a superposition of denoising in multiple directions.

[0042] Finally, the denoised character image is processed by contour enhancement and the target background style is added to obtain the character image to be recognized. The effect image can be seen in Figure 2 (g) in.

[0043] The specific operation of contour enhancement processing is: create a jitter matrix of size n×n, and the element values ​​in the jitter matrix range from 0 to n 2 −1, and each element value is unique; the jitter matrix is ​​normalized to obtain a normalized jitter matrix; the grayscale image of the denoised character image is divided into n × n The average gray value of each block is added to the element at the corresponding position in the normalized jitter matrix to obtain an initial contour enhancement image; the initial contour enhancement image is binarized to obtain a contour enhancement character image.

[0044] The above-mentioned adding of the target background style can be achieved by superimposing the target style standard image and the contour enhanced character image.

[0045] S3. After feature extraction and encoding, the character image to be recognized is converted into a vector form to obtain a feature sequence of the character to be recognized; after the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network, repeated characters in the prediction sequence are merged, and blank marks are deleted to obtain the character recognition result.

[0046] The character image to be recognized is converted into a vector form after feature extraction and encoding processing to obtain the feature sequence of the character to be recognized; the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network to capture the context information of the image-based character sequence from the front and back directions, and then the repeated characters in the predicted sequence are merged and the blank marks are deleted to obtain the character recognition result.

[0047] First, the character image to be recognized is processed by feature extraction and encoding and then converted into a vector form to obtain a feature sequence of the character to be recognized.

[0048] Among them, the operation of feature extraction and coding processing can be implemented through several convolutions, batch normalization, Relu activation function and maximum pooling. Specifically, the operation of feature extraction and coding processing can be implemented through 6 convolutions, batch normalization, Relu activation function and maximum pooling, and 1 convolution, batch normalization, Relu activation function. In the feature extraction and coding processing, each convolution is composed of a different number of cores. The number of cores increases with the depth of the neural network. The more cores there are, the more deep features are extracted.

[0049] Then, the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network (which can be realized by a BiLSTM network or a BiGRU), the repeated characters in the prediction sequence are merged, and the blank marks are deleted to obtain the character recognition result. The effect diagram can be seen in Figure 3 .

[0050] This embodiment further provides a texture background character recognition system based on image restoration, which is used to implement the above-mentioned texture background character recognition method based on image restoration, including: The texture background repair map generation module is used for the image to be identified to obtain the initial feature map to be identified after several convolution processes with successively reduced scales, long-distance feature interaction processes, and several convolution processes with successively increased scales; the initial feature map to be identified is subjected to several convolution processes with successively reduced scales, several residual connection processes based on channel attention, and several convolution processes with successively increased scales to obtain the texture background repair map; the operation of the long-distance feature interaction process is: the convolution image to be identified is subjected to several mask context information fusion processes based on fixed windows and mask update context information fusion processes based on moving windows to obtain a long-distance feature interaction map, which is used to perform several operations of successively increased scale convolution processes; the operation of the mask context information fusion process based on fixed windows is specifically: after the input is processed by the mask multi-head attention based on fixed windows, it is feature concatenated with the input, and is processed by full connection and multi-layer perceptron to obtain the output; in the mask multi-head attention process based on fixed windows, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the sum of the pixel mask values; The module for generating a character image to be recognized is used to obtain a difference image between the image to be recognized and the texture background restoration image to obtain a foreground character image; according to the texture direction in the texture background image, the foreground character image is subjected to texture direction denoising to obtain a denoised character image; after the denoised character image is subjected to contour enhancement processing, the target background style is added to obtain a character image to be recognized; the operation of texture direction denoising is as follows: based on the normalized value of each pixel point of the foreground character image and the corresponding color filter adaptive value, the respective color filter mask values ​​are obtained; based on the color filter mask value of each pixel point and the pixel value of the corresponding position in the texture direction variant template image, the texture background image is subjected to color filter processing to obtain a denoised character image; The character recognition result generation module is used to convert the character image to be recognized into a vector form after feature extraction and encoding processing, so as to obtain the feature sequence of the character to be recognized; after the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network, the repeated characters in the prediction sequence are merged, and the blank marks are deleted to obtain the character recognition result.

[0051] This embodiment also provides a texture background character recognition device based on image restoration, including a processor and a memory, wherein the processor implements the above-mentioned texture background character recognition method based on image restoration when executing a computer program stored in the memory.

[0052] This embodiment further provides a computer-readable storage medium for storing a computer program, wherein when the computer program is executed by a processor, the above-mentioned texture background character recognition method based on image restoration is implemented.

[0053] The present embodiment provides a texture background character recognition method based on image restoration. First, the image to be recognized is subjected to convolution processing, long-distance feature interaction processing, and residual connection processing based on channel attention, and the missing area in the image is filled with visible information to reduce color differences and blur, and maintain the structural rationality and texture consistency of the filled area, so as to restore the texture background of the character-covered part and obtain a texture background restoration image; then, a foreground character image is obtained by subtracting the image to be recognized from the texture background restoration image, and the foreground character image is denoised to eliminate texture noise in all directions to obtain a denoised character image; after enhancing the character contour details in the denoised character image, The target background style is added to eliminate the micro-pixel scale residual error in the image restoration process, effectively improve the visual perception quality while retaining the character topology, and obtain the character map to be recognized; finally, the character map to be recognized is converted into a vector form after feature extraction and encoding, and the feature sequence of the character to be recognized is obtained. After the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network, the context information of the image-based character sequence is captured from the front and back directions, the repeated characters in the predicted sequence are merged, and the blank marks are deleted to obtain the character recognition result; it is applied to image restoration and character recognition in industrial scenes to improve the quality of image restoration as well as the accuracy and efficiency of character recognition.

[0054] This embodiment provides a texture background character recognition method based on image restoration. It adopts fusion learning in long-distance feature interaction processing, which solves the problem of training instability caused by a large proportion of invalid tokens, and encourages the learning of low-frequency basic features so that high-frequency details can be better learned later, reducing the difficulty of optimization. The attention processing adopts a fixed window and mask update strategy, and uses effective markers to borrow visible information to fill holes, thereby reducing color differences and blur.

[0055] This embodiment provides a texture background character recognition method based on image restoration. In the residual connection processing based on channel attention, the surrounding local information is used to properly repair some missing areas. At the same time, the residual block is used to speed up the image restoration model and improve the image restoration performance.

[0056] This embodiment provides a texture background character recognition method based on image restoration, which realizes the connection between the two stages of image restoration and character recognition, and eliminates the non-negligible errors existing in the micro-pixel scale of the image restoration network; by adding the target background style after the contour enhancement process, the visual perception quality is effectively improved while retaining the character topological structure, thereby providing high-quality character data for subsequent character recognition.

Claims

1. A texture background character recognition method based on image restoration, characterized in that: The following operations are included: S1. The image to be identified is subjected to several convolution processes with successively reduced scales, long-distance feature interaction processes, and several convolution processes with successively increased scales to obtain an initial feature map to be identified; the initial feature map to be identified is subjected to several convolution processes with successively reduced scales, several residual connection processes based on channel attention, and several convolution processes with successively increased scales to obtain a texture background restoration map; The operation of long-distance feature interaction processing is as follows: the convolution image to be identified is subjected to several fixed-window-based mask context information fusion processing and moving-window-based mask update context information fusion processing to obtain a long-distance feature interaction map, which is used to perform several convolution processing operations with successively increasing scales; The operation of the fixed window-based mask context information fusion processing is as follows: after the input is processed by the fixed window-based mask multi-head attention, it is feature concatenated with the input, and processed by the full connection processing and the multi-layer perceptron to obtain the output; in the fixed window-based mask multi-head attention processing, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the sum of the pixel mask values; S2, obtaining a difference image between the image to be recognized and the texture background restoration image to obtain a foreground character image; performing texture direction denoising on the foreground character image according to the texture direction in the texture background image to obtain a denoised character image; After the de-noised character image is processed by contour enhancement, the target background style is added to obtain the character image to be recognized; The operation of texture direction denoising is as follows: based on the normalized value of each pixel of the foreground character image and the corresponding filter adaptive value, obtain the respective filter mask value; Based on the color filter mask value of each pixel and the pixel value of the corresponding position in the texture direction variant template image, the texture background image is subjected to color filter processing to obtain a denoised character image; S3. After feature extraction and encoding, the character image to be recognized is converted into a vector form to obtain a feature sequence of the character to be recognized; after the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network, repeated characters in the prediction sequence are merged, and blank marks are deleted to obtain the character recognition result.

2. The texture background character recognition method based on image restoration according to claim 1 is characterized in that: In S1, the operation of residual connection processing based on channel attention is specifically as follows: The initial convolution feature map to be identified or the N-1th residual connection feature map is processed by convolution, batch normalization, Relu activation function, convolution, batch normalization and channel attention, and then fused with the initial convolution feature map to be identified or the N-1th residual connection feature map to obtain the Nth residual connection feature map; the last residual connection feature map is used as the residual connection feature map to be identified, and is used to perform several convolution operations with successively increasing scales; The initial convolution feature map to be identified is obtained by performing convolution processing on the initial feature map to be identified with the scales successively reduced for several times.

3. The texture background character recognition method based on image restoration according to claim 1, characterized in that: In the channel attention processing in S1, the channel attention weight is obtained by the following formula: , , , For the i The channel attention weight corresponding to the pixel position point, is the sigmoid function, For the i The pixel location point in the neighborhood j The initial channel attention weight of pixels, For the i The pixel location point in the neighborhood j The pixel value of a pixel, J For the i The total number of pixels in the neighborhood of a pixel. The convolution kernel is K One-dimensional convolution processing, is the channel dimension, is the convolution kernel coefficient, is the convolution kernel bias, is the nearest odd function.

4. The texture background character recognition method based on image restoration according to claim 1, characterized in that: In S1, the current head attention probability distribution coefficient is obtained by the following formula: , For the i The probability distribution coefficient of individual attention, is the softmax function, For the i The query vector of the head, For the i The transpose of the index vector of the head, is the total number of pixels of the input image, is the sum of the pixel mask values ​​of the input image. If there is an invalid position in the current window in the input image, is the mask map corresponding to the input image, For the i The content vector of the header, is the dimension of the index vector.

5. The texture background character recognition method based on image restoration according to claim 1, characterized in that: In S2, the color filter mask value is obtained by the following formula: when At that time, i The filter mask value of each pixel =1; when At that time, i The filter mask value of each pixel =0; , , Foreground character map i The normalized value of pixels, Foreground character map i The color filter adaptive threshold of pixels, Foreground character map i The pixel value of a pixel, , are the maximum pixel value and the minimum pixel value in the foreground character image, respectively. is the gradient function.

6. The texture background character recognition method based on image restoration according to claim 1, characterized in that: In S2, the operation of performing color filtering on the texture background image is specifically as follows: If the foreground character image i The filter mask value of each pixel =1, the color filtering is achieved by the following formula: , If the foreground character image i The filter mask value of each pixel =0, the color filtering process is implemented by the following formula: , The denoised character image i The pixel value of a pixel, Foreground character map i The pixel value of a pixel, For the i The pixel value of the corresponding position of a pixel in the texture direction variant template image.

7. The texture background character recognition method based on image restoration according to claim 1 is characterized in that: In S3, the feature extraction and encoding processing operations can be implemented through several times of convolution, batch normalization, ReLU activation function and maximum pooling processing.

8. A texture background character recognition system based on image restoration, used to implement the texture background character recognition method based on image restoration according to claim 1, characterized in that: include: The texture background restoration image generation module is used to obtain the initial feature map to be recognized by subjecting the image to be recognized to several convolution processes with successively reduced scales, long-distance feature interaction processes, and several convolution processes with successively increased scales; the initial feature map to be recognized is subjected to several convolution processes with successively reduced scales, several residual connection processes based on channel attention, and several convolution processes with successively increased scales to obtain the texture background restoration image; The operation of long-distance feature interaction processing is as follows: the convolution image to be identified is subjected to several fixed-window-based mask context information fusion processing and moving-window-based mask update context information fusion processing to obtain a long-distance feature interaction map, which is used to perform several convolution processing operations with successively increasing scales; The operation of the fixed window-based mask context information fusion processing is as follows: after the input is processed by the fixed window-based mask multi-head attention, it is feature concatenated with the input, and processed by the full connection processing and the multi-layer perceptron to obtain the output; in the fixed window-based mask multi-head attention processing, the current head attention probability distribution coefficient is obtained based on the total number of pixels and the sum of the pixel mask values; The module for generating a character image to be recognized is used to obtain a difference image between the image to be recognized and the texture background restoration image to obtain a foreground character image; and to perform texture direction denoising on the foreground character image according to the texture direction in the texture background image to obtain a denoised character image. After the denoised character image is processed by contour enhancement, the target background style is added to obtain the character image to be recognized; the operation of texture direction denoising is: based on the normalized value of each pixel of the foreground character image and the corresponding filter color adaptation value, the respective filter color mask values ​​are obtained; Based on the color filter mask value of each pixel and the pixel value of the corresponding position in the texture direction variant template image, the texture background image is subjected to color filter processing to obtain a denoised character image; The character recognition result generation module is used to convert the character image to be recognized into a vector form after feature extraction and encoding processing, so as to obtain the feature sequence of the character to be recognized; after the feature sequence of the character to be recognized is processed by a bidirectional long short-term memory network, the repeated characters in the prediction sequence are merged, and the blank marks are deleted to obtain the character recognition result.

9. A texture background character recognition device based on image restoration, characterized in that: The method comprises a processor and a memory, wherein the processor implements the texture background character recognition method based on image restoration as described in any one of claims 1 to 7 when executing the computer program stored in the memory.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the texture background character recognition method based on image restoration as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Character recognition system based on Gabor convolution and linear sparse attention

    CN113221874A

  • Inclined character recognition method based on complex background image

    CN115439857A

  • Method for repairing spectrum-space joint repairing network based on shielding vector

    CN118521864A

  • Image segmentation model training method, image segmentation method, and apparatus

    WO2024031219A1