Low-quality steel seal character recognition method based on complex background and electronic device

Through the diffusion model and control network generation data samples, combined with image processing and deep learning technology for preprocessing and recognition, the problem of unclear stamp character recognition under complex backgrounds and low-quality images is solved, and high-precision and reliable recognition results are achieved.

CN119942510APending Publication Date: 2025-05-06中国融通资源开发集团有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411947302.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Under complex backgrounds and low-quality image conditions, it is difficult for the prior art to accurately identify stamped characters, resulting in low recognition accuracy.

Method used

The diffusion model is used to combine the control network to generate data samples, and the sample images are preprocessed through image processing technology and deep learning methods, including contrast enhancement, Gaussian denoising and mitigation of distortion. The character area segmentation network and character recognition model are used based on deep learning for precise segmentation and recognition, and post-processing correction is performed through character context correlation and logic verification.

Benefits of technology

It effectively improves the accuracy and reliability of steel stamp character recognition under complex backgrounds and low-quality image conditions, ensures the logical correctness of the recognition results, and is suitable for a wide range of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942510A_ABST
    Figure CN119942510A_ABST
Patent Text Reader

Abstract

The invention discloses a low-quality steel seal character recognition method based on a complex background and an electronic device. The method comprises the steps that a diffusion model is combined with a control network to generate a data sample; an image processing technology is combined with a deep learning method, and contrast enhancement, Gaussian denoising, distortion reduction and preprocessing of a denoising network are performed on a sample image, so that character features of the sample image are enhanced; performing accurate segmentation on a character region in the preprocessed sample image by using a character region segmentation network based on deep learning; positioning the segmented character region by using an improved target detection algorithm, and extracting accurate position information of the character; a character recognition model based on deep learning is adopted to recognize the character image in the segmented and positioned character region in the sample image and output a character recognition result; in the recognition process, the recognition result is subjected to post-processing correction through character context correlation and logic verification, and the logic correctness of the recognition result is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of character recognition, and in particular to a method and an electronic device for recognizing low-quality steel-printed characters under a complex background. Background Art

[0002] With the development of industrial automation and intelligence, the demand for machine vision systems is growing. In the process of identifying some objects, traditional methods are often limited by factors such as image quality and environmental conditions, resulting in low recognition accuracy.

[0003] At present, although scene text detection and recognition technology based on deep learning has achieved remarkable results in many fields, it still faces severe challenges in handling scenarios in specific complex environments. These challenges mainly include inconsistency in image quality, such as poor lighting conditions, image blur, stain occlusion, and variable object postures, as well as the particularity of the stamped characters themselves, such as multi-line encoding and large font differences. These factors seriously restrict the recognition accuracy and practicality of existing technologies. Summary of the invention

[0004] The purpose of the present invention is to provide a method and electronic device for recognizing low-quality steel-printed characters under complex backgrounds, so as to solve the problem of unclear recognition under complex backgrounds and low-quality image conditions.

[0005] In order to achieve the above tasks, the present invention adopts the following technical solutions:

[0006] A method for recognizing low-quality steel stamp characters under a complex background comprises the following steps:

[0007] Step 1, using a diffusion model combined with a control network to generate data samples; wherein the diffusion model is used to learn data distribution from existing steel-printed character image samples and generate new image samples, thereby expanding the image sample data set; the control network is used to cooperate with the diffusion model to control the generation details of the image samples, including controlling the content of the generated steel-printed characters;

[0008] Step 2: Using image processing technology combined with deep learning methods, the sample image is subjected to contrast enhancement, Gaussian denoising, distortion reduction, and denoising network preprocessing, thereby enhancing the character features of the sample image;

[0009] Step 3, using a deep learning-based character region segmentation network to accurately segment the character region in the preprocessed sample image; and using an improved target detection algorithm to locate the segmented character region and extract the precise position information of the characters;

[0010] Step 4: using a deep learning-based character recognition model to recognize the character image in the segmented and located character area in the sample image and output the character recognition result; during the recognition process, post-processing correction is performed on the recognition result through character context relevance and logic verification to ensure the logical correctness of the recognition result;

[0011] Step 5. For the steel-stamped character image to be recognized, the image is input into the trained denoising network for preprocessing after contrast enhancement, Gaussian denoising and distortion reduction. The preprocessed image is input into the trained character region segmentation network to extract the precise location information of the character. The extraction result is then input into the trained character recognition model to obtain the recognition result of the steel-stamped character.

[0012] Furthermore, the construction process of the diffusion model includes:

[0013] Use the pre-trained Stable Diffusion 2.1 model as the basis and Stable-Diffusion-2-1-Realistic as the base model of the diffusion model;

[0014] Based on the Stable-Diffusion-2-1-Realistic model, the LoRA model is introduced to construct a diffusion model. The existing steel-stamped character image samples are used to train the LoRA model, and the parameters of the diffusion model are adjusted to enable it to better learn data distribution from existing image samples and adapt to the task of generating new image samples.

[0015] Furthermore, the control network includes:

[0016] OCR engine: used to detect text information in image samples;

[0017] Glyph rendering: Render the detected text information at the corresponding position in the whiteboard image to form a glyph image;

[0018] Image VAE encoder and decoder: encode the character image in the input image sample into a latent code, and reconstruct the output character image based on the latent code;

[0019] Text encoder: converts text information into text embedding;

[0020] Glyph control network: encodes glyph information by processing glyph images;

[0021] U-Net encoder and decoder: perform denoising diffusion process, in which the glyph image and character image are fused, and the glyph information is encoded through the glyph control network to generate stamped characters; the generated stamped characters will be provided to the diffusion network, which combines the stamped characters and the data distribution learned from the existing image samples to generate new sample images.

[0022] Furthermore, the contrast enhancement, Gaussian denoising, and distortion reduction of the sample image include:

[0023] Adaptive histogram equalization is used to enhance the contrast of image samples to make characters more prominent;

[0024] A Gaussian filter is used to remove Gaussian noise from image samples;

[0025] The traditional color restoration multi-scale MSRCR algorithm is improved to reduce the distortion of the sample image, including:

[0026] Firstly, the guided filter is used to replace the Gaussian filter in the MSRCR algorithm to reduce the halo artifacts. In addition, an adaptive nonlinear offset is introduced to replace the linear offset in the MSRCR algorithm. The adaptive nonlinear offset refers to dividing the sample image into multiple regions, calculating the image contrast of each region, and using the image contrast in the region to adjust the linear offset in the region.

[0027] Furthermore, the denoising network processes the sample image including:

[0028] (1) First, generate two similar low-resolution images based on the sample image sampling:

[0029] Divide the sample image into squares, select two random pixels a and b from each square, place pixel a at the corresponding position of a new image, and place pixel b at the corresponding position of another new image; repeat this process, so that the two new images generate low-resolution images A and B with the same resolution as the original sample image;

[0030] (2) Three 1×1 convolutional layers are added at the end of the traditional U-Net network as a denoising network to reduce the noise of the image;

[0031] (3) Use image samples and corresponding low-resolution images A and B to train the denoising network:

[0032] For each sample image, the square error between the two corresponding low-resolution images A and B is calculated as the reconstruction loss. Then the sampled low-resolution image A is input into the denoising network, and a new low-resolution image A′ is output through the network. The pixel true value difference between the low-resolution image A′ and the sample image is calculated and regularized, and the result is used as the regularization loss. The regularization loss is multiplied by a preset coefficient and added to the reconstruction loss as the total network loss value.

[0033] Then, start the training process of the denoising network; forward propagation calculates the total loss value of the network, and reverse propagation optimizes the parameters in the network. Repeat this process continuously, and save the parameters in the network when the total loss value of the network drops to a new low point; continue to adjust the learning rate until the total loss value of the network no longer continues to decrease after multiple rounds, and the parameters in the network tend to be stable, and save the denoising network after training.

[0034] Furthermore, a character region segmentation network with a CNN+RPN architecture is used to perform character region segmentation on the sample image, where the CNN network is used to extract the feature map of the image, and the RPN network slides the window on the image feature map to generate multiple region proposals for each window position; the RPN network contains a classifier for determining whether each region in the window contains stamped characters, and a bounding box regressor for fine-tuning the position of the region.

[0035] Furthermore, the improved target detection algorithm includes:

[0036] Based on the traditional non-maximum suppression, the confidence scores of non-maximum detection boxes are reduced proportionally and retained instead of removed.

[0037] Furthermore, the character recognition model includes a convolutional neural network (CNN), a recurrent neural network (RNN), and a connection temporal classification module (CTC), wherein:

[0038] Convolutional neural network CNN uses multiple convolutional layers to extract features of images;

[0039] The recurrent neural network RNN ​​uses a bidirectional LSTM to process the feature sequence in the character image and capture the temporal dependency to achieve context relevance and logic verification; a fully connected layer MLP is set at the last layer of the recurrent neural network module RNN to convert the output of the recurrent neural network module RNN into the final prediction result;

[0040] The connected temporal classification module CTC uses the CTC loss function to transform the alignment problem between the sample image and the output character recognition result into a probability maximization problem.

[0041] An electronic device includes a processor, a memory, and a computer program stored in the memory; when the processor executes the computer program, the method for recognizing low-quality steel-printed characters under a complex background is implemented.

[0042] A computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the method for recognizing low-quality steel-printed characters under a complex background is implemented.

[0043] Compared with the prior art, the present invention has the following technical features:

[0044] 1. The present invention adopts adaptive image enhancement technology to significantly improve image quality in the preprocessing stage, effectively enhance character features, and provide high-quality input data for subsequent steps, thereby improving overall recognition accuracy.

[0045] 2. Through deep learning networks and improved target detection algorithms, the present invention can accurately segment the character area in the image and accurately locate the character position, effectively coping with the challenges brought by complex backgrounds and low-quality images.

[0046] 3. The character enhancement step further improves the recognizability of characters, making them accurately recognized even in complex backgrounds.

[0047] 4. Use deep learning models for character recognition, and combine character context relevance and logical verification rules for post-processing correction to ensure the logical correctness of the recognition results and improve the accuracy and reliability of recognition.

[0048] 5. The present invention also significantly reduces the size of the deep learning model and improves the recognition speed through lightweight model optimization technology, so that the method can meet the needs of rapid response on site while maintaining a high recognition accuracy.

[0049] 6. In addition, the present invention has broad application prospects. It is not only suitable for the recognition of steel-printed characters on objects, but can also be widely used in other fields that require character recognition, such as license plate recognition, bill recognition, etc.

[0050] In summary, the present invention is based on a method for recognizing low-quality steel stamped characters under complex backgrounds. By adopting advanced image processing technology and deep learning algorithms, it effectively solves the problem of unclear recognition under complex backgrounds and low-quality image conditions, improves the accuracy and reliability of recognition, and has good practicality and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0052] The present invention provides a method for recognizing low-quality steel-printed characters under a complex background, comprising the following steps:

[0053] Step 1, in the data preparation stage, use the diffusion model combined with the control network to generate data samples; wherein the diffusion model is used to learn data distribution from existing steel-printed character image samples and generate new image samples, thereby expanding the image sample data set; the control network is used to cooperate with the diffusion model to control the generation details of the image samples, including controlling the content of the generated steel-printed characters.

[0054] Stamped characters are different from traditional characters. When printed on the surface of an object, they will produce a certain degree of depression. At the same time, the encoding method of stamped characters is irregular and the fonts vary greatly. This deformation, the light and shadow changes in the stamped character area, and the color of the object itself will have a great impact on the recognition results. There are currently few public stamped character data sets. For the recognition of stamped characters on different products, it is difficult to obtain sufficient and effective image sample data due to the combined influence of product surface color, material, the stamped characters themselves, lighting conditions, stains, occlusion, image blur, posture changes, etc. Therefore, how to expand new image samples based on limited image samples has an important impact on improving the generalization of the network model.

[0055] In one embodiment of the present invention, the specific implementation steps are as follows:

[0056] 1.1 Build the base mold of the diffusion model.

[0057] Use the pre-trained Stable Diffusion 2.1 model as the basis, and use the more realistic Stable-Diffusion-2-1-Realistic as the base model of the diffusion model to ensure that the generated images are more realistic.

[0058] 1.2 Train the Lora model.

[0059] Based on the Stable-Diffusion-2-1-Realistic model, the LoRA model is introduced to construct a diffusion model. This structure enables the model to quickly adapt to new image sample training tasks while keeping most of the pre-trained weights unchanged. By using the current limited steel-stamped character image samples to train the LoRA model, the limited image samples can be effectively used to adjust the parameters of the diffusion model, so that it can better learn data distribution from existing image samples and adapt to the generation task of new image samples. For specific training methods, the peft package in the python third-party library transformers can be used.

[0060] LoRA (Low-Rank Adaptation) is a model adaptation scaling model that achieves model adaptive adjustment by introducing a low-rank structure into a pre-trained large model without retraining the entire model.

[0061] 1.3 Add control network.

[0062] A control network is added in the process of the diffusion model generating stamped character image samples. The control network is based on visual text and combines glyph information to enhance the performance of the diffusion model in generating accurate visual text. The control network enhances the generation of text to image by introducing glyph input conditions. The glyph image serves as an explicit spatial layout prior, forcing the diffusion model to generate coherent and well-formatted visual text, which can be seamlessly connected with the diffusion model.

[0063] The control network provided by the present invention comprises the following structure:

[0064] OCR engine: used to detect text information in image samples;

[0065] Glyph rendering: Render the detected text information at the corresponding position in the whiteboard image to form a glyph image;

[0066] Image VAE encoder and decoder: encode the character image in the input image sample into a latent code, and reconstruct the output character image based on the latent code;

[0067] Text encoder (OpenAI CLIP text encoder is used by default): converts text information into text embedding;

[0068] Glyph control network: encodes glyph information by processing glyph images;

[0069] U-Net encoder and decoder: perform denoising diffusion process, in which the glyph image and character image are fused, and the glyph information is encoded through the glyph control network to generate stamped characters; the generated stamped characters will be provided to the diffusion network, which combines the stamped characters and the data distribution learned from the existing image samples to generate new sample images.

[0070] 1.4 Batch generate steel stamp character image samples.

[0071] After completing the above steps, realistic steel stamp character image samples can be generated in batches and added to the image sample data set; of course, due to the instability of the diffusion model itself, the generated image samples still require a small amount of manual intervention to remove samples of poor quality; this step greatly expands the image sample data set and lays a good data foundation for subsequent image processing and character recognition.

[0072] Step 2: Use image processing technology combined with deep learning methods to perform contrast enhancement, Gaussian denoising, distortion reduction, and denoising network preprocessing on the sample image, thereby enhancing the character features of the sample image.

[0073] This step can effectively improve the quality of the sample image, solve common problems such as reflection, oil stains, rust spots, etc. in steel stamp character recognition, and provide high-quality input data for subsequent steps, thereby improving the overall recognition accuracy.

[0074] 2.1 Contrast enhancement.

[0075] Adaptive Histogram Equalization (CLAHE) is used to enhance the contrast of image samples, even out image illumination, and effectively reduce the local reflection problem of typical steel-printed character images, making the characters more prominent.

[0076] 2.2 Gaussian denoising.

[0077] A Gaussian filter is used to remove Gaussian noise from image samples. Gaussian noise is usually caused by resistor heating in electronic devices, or by external environmental interference when the sensor is capturing images. It is commonly found in all types of images and will have a certain negative impact on subsequent image detection and character positioning and recognition.

[0078] 2.3 Reduce distortion.

[0079] Improve the traditional color restoration multi-scale MSRCR algorithm to address the problem of local hue distortion of typical steel-printed character images when photographed outdoors and enhance image quality. The specific improvement plan is as follows:

[0080] Firstly, the guided filter is used to replace the Gaussian filter in the classic MSRCR algorithm to reduce the halo artifacts. In addition, the adaptive nonlinear offset is introduced to replace the linear offset in the classic MSRCR algorithm, which effectively alleviates the problem of local hue distortion. The adaptive nonlinear offset refers to dividing the sample image into multiple regions, calculating the image contrast of each region, and using the image contrast in the region to adjust the linear offset in the region.

[0081] 2.4 Denoising network.

[0082] (1) First, two similar low-resolution images are generated by sampling the sample image, as follows:

[0083] Divide the sample image into 2×2 squares, select two random pixels a and b from each 2×2 square, place pixel a at the corresponding position of a new image, and place pixel b at the corresponding position of another new image; repeat this process so that the two new images generate low-resolution images A and B with the same resolution as the original sample image.

[0084] (2) Use the improved U-Net network as the denoising network to denoise the image. U-Net is a commonly used convolutional neural network (CNN) architecture, which is used here for image segmentation and denoising tasks. The characteristic of U-Net is that there are jump connections between the downsampling path and the upsampling path of the network, which helps to retain more detail information during the image reconstruction process.

[0085] Based on the traditional U-Net network, the present invention adds three 1×1 convolutional layers at the end of the network. This design can further refine features and improve denoising performance.

[0086] (3) Perform the above operation on the sample images in the image sample data set to obtain low-resolution images A and B of each sample image, and train the U-Net network as follows:

[0087] For each sample image, the square error of the two corresponding low-resolution images A and B is calculated as the reconstruction loss; then the sampled low-resolution image A is input into the U-Net network, and a new low-resolution image A′ is output through the network. The pixel true value difference between the low-resolution image A′ and the sample image is calculated and regularized, and the result is used as the regularization loss; the regularization loss is multiplied by a preset coefficient, and the reconstruction loss is added as the total network loss value. The purpose of the regularization term here is to reduce the over-smoothing problem caused by the non-zero true value difference introduced in the subsampling process.

[0088] Then, start the training process of the U-Net network; forward propagation calculates the total network loss value, and reverse propagation optimizes the parameters in the network. Repeat this process continuously, and save the parameters in the network when the total network loss value drops to a new low point; continue to adjust the learning rate until the total network loss value no longer continues to decrease after multiple rounds, and the parameters in the network tend to be stable, and save the trained U-Net network model; this model can remove noise in the image very well.

[0089] Step 3: Use a deep learning-based character region segmentation network to accurately segment the character region in the preprocessed sample image; and use an improved target detection algorithm to locate the segmented character region and extract the precise position information of the characters.

[0090] The present invention adopts a deep learning network with a CNN+RPN architecture to perform character region segmentation on sample images, wherein the CNN network is used to extract the feature map of the image, and the RPN network slides the window on the image feature map to generate multiple region suggestions for each window position; the RPN network contains a classifier for determining whether each region in the window contains stamped characters, and a bounding box regressor for fine-tuning the position of the region.

[0091] After completing the character area segmentation using the CNN+RPN architecture, an improved target detection algorithm is used. Based on the traditional non-maximum suppression (NMS), this algorithm does not simply suppress all overlapping detection boxes and completely remove them. Instead, it proportionally reduces the confidence scores of non-maximum detection boxes and retains them instead of removing them. This helps to retain more detection boxes and improve the recall rate of the model, which is of great value for steel stamp character recognition scenarios.

[0092] In the context of character detection and positioning, CNN can learn features from images that help distinguish characters from non-character areas. In the selection of CNN network, the classic ResNet50 is used. It is a deep residual network, a variant of the ResNet series. The characteristics of residual learning help solve the degradation problem in deep networks, that is, the problem that the training error increases with the increase of network layers; ResNet50 contains 50 layers of depth, which includes 4 stages, and each stage consists of multiple residual blocks. This hierarchical structure helps the network learn features of different scales.

[0093] Step 4: Use a deep learning-based character recognition model to recognize the character image in the segmented and located character area in the sample image and output the character recognition result; during the recognition process, post-process and correct the recognition result through character context relevance and logical verification to ensure the logical correctness of the recognition result.

[0094] The character recognition model proposed in the present invention includes a convolutional neural network (CNN), a recurrent neural network (RNN), and a connection temporal classification module (CTC), wherein:

[0095] The convolutional neural network CNN uses multiple convolutional layers (ResNet50 is selected in this embodiment) to extract image features; these network layers can extract local features in character images, such as the edges, corners and textures of stamped characters; a pooling layer is added at the end of this convolutional neural network module CNN to reduce the spatial dimension of the feature map while retaining important feature information.

[0096] The recurrent neural network RNN ​​adopts a bidirectional LSTM to process the feature sequence in the character image and capture the temporal dependency to achieve context relevance and logic verification; and a fully connected layer MLP is set at the last layer of the recurrent neural network module RNN to convert the output of the recurrent neural network module RNN into the final prediction result. In the present invention, the output result is the probability distribution of the character category.

[0097] The connection time classification module CTC uses the CTC loss function to transform the alignment problem between the sample image and the output character recognition result into a probability maximization problem. It is a commonly used loss function in the CRNN structure. This is a method for unsupervised learning sequence prediction that allows the model to output sequences of variable length and can handle the misalignment problem between the input and output sequences. This is designed for the special scenario of steel stamp character recognition, where the length of the output result is uncertain.

[0098] Step 5. For the steel-stamped character image to be recognized, the image is input into the trained denoising network for preprocessing after contrast enhancement, Gaussian denoising and distortion reduction. The preprocessed image is input into the trained character region segmentation network to extract the precise location information of the character. The extraction result is then input into the trained character recognition model to obtain the recognition result of the steel-stamped character.

[0099] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for recognizing low-quality steel stamp characters under complex background, characterized in that: The following steps are involved: Step 1, using a diffusion model combined with a control network to generate data samples; wherein the diffusion model is used to learn data distribution from existing steel-printed character image samples and generate new image samples, thereby expanding the image sample data set; the control network is used to cooperate with the diffusion model to control the generation details of the image samples, including controlling the content of the generated steel-printed characters; Step 2: Using image processing technology combined with deep learning methods, the sample image is subjected to contrast enhancement, Gaussian denoising, distortion reduction, and denoising network preprocessing, thereby enhancing the character features of the sample image; Step 3, using a deep learning-based character region segmentation network to accurately segment the character region in the preprocessed sample image; and using an improved target detection algorithm to locate the segmented character region and extract the precise position information of the characters; Step 4: using a deep learning-based character recognition model to recognize the character image in the segmented and located character area in the sample image and output the character recognition result; during the recognition process, post-processing correction is performed on the recognition result through character context relevance and logic verification to ensure the logical correctness of the recognition result; Step 5. For the steel-stamped character image to be recognized, the image is input into the trained denoising network for preprocessing after contrast enhancement, Gaussian denoising and distortion reduction. The preprocessed image is input into the trained character region segmentation network to extract the precise location information of the character. The extraction result is then input into the trained character recognition model to obtain the recognition result of the steel-stamped character.

2. The method for recognizing low-quality steel stamp characters under complex background according to claim 1 is characterized in that: The construction process of the diffusion model includes: Use the pre-trained Stable Diffusion 2.1 model as the basis and Stable-Diffusion-2-1-Realistic as the base model of the diffusion model; Based on the Stable-Diffusion-2-1-Realistic model, the LoRA model is introduced to construct a diffusion model. The existing steel-stamped character image samples are used to train the LoRA model, and the parameters of the diffusion model are adjusted to enable it to better learn data distribution from existing image samples and adapt to the task of generating new image samples.

3. The method for recognizing low-quality steel stamp characters under complex background according to claim 1, characterized in that: The control network comprises: OCR engine: used to detect text information in image samples; Glyph rendering: Render the detected text information at the corresponding position in the whiteboard image to form a glyph image; Image VAE encoder and decoder: encode the character image in the input image sample into a latent code, and reconstruct the output character image based on the latent code; Text encoder: converts text information into text embedding; Glyph control network: encodes glyph information by processing glyph images; U-Net encoder and decoder: perform denoising diffusion process, in which the glyph image and character image are fused, and the glyph information is encoded through the glyph control network to generate stamped characters; the generated stamped characters will be provided to the diffusion network, which combines the stamped characters and the data distribution learned from the existing image samples to generate new sample images.

4. The method for recognizing low-quality steel stamp characters under complex background according to claim 1, characterized in that: The step of performing contrast enhancement, Gaussian denoising, and distortion reduction on the sample image includes: Adaptive histogram equalization is used to enhance the contrast of image samples to make characters more prominent; A Gaussian filter is used to remove Gaussian noise from image samples; The traditional color restoration multi-scale MSRCR algorithm is improved to reduce the distortion of the sample image, including: Firstly, the guided filter is used to replace the Gaussian filter in the MSRCR algorithm to reduce the halo artifacts. In addition, an adaptive nonlinear offset is introduced to replace the linear offset in the MSRCR algorithm. The adaptive nonlinear offset refers to dividing the sample image into multiple regions, calculating the image contrast of each region, and using the image contrast in the region to adjust the linear offset in the region.

5. The method for recognizing low-quality steel stamp characters under complex background according to claim 1 is characterized in that: The denoising network processes the sample image in the following way: (1) First, generate two similar low-resolution images based on the sample image sampling: Divide the sample image into squares, select two random pixels a and b from each square, place pixel a at the corresponding position of a new image, and place pixel b at the corresponding position of another new image; repeat this process, so that the two new images generate low-resolution images A and B with the same resolution as the original sample image; (2) Three 1×1 convolutional layers are added at the end of the traditional U-Net network as a denoising network to reduce the noise of the image; (3) Use image samples and corresponding low-resolution images A and B to train the denoising network: For each sample image, the square error between the two corresponding low-resolution images A and B is calculated as the reconstruction loss. Then the sampled low-resolution image A is input into the denoising network, and a new low-resolution image A′ is output through the network. The pixel true value difference between the low-resolution image A′ and the sample image is calculated and regularized, and the result is used as the regularization loss. The regularization loss is multiplied by a preset coefficient and added to the reconstruction loss as the total network loss value. Then, start the training process of the denoising network; forward propagation calculates the total loss value of the network, and reverse propagation optimizes the parameters in the network. Repeat this process continuously, and save the parameters in the network when the total loss value of the network drops to a new low point; continue to adjust the learning rate until the total loss value of the network no longer continues to decrease after multiple rounds, and the parameters in the network tend to be stable, and save the denoising network after training.

6. The method for recognizing low-quality steel stamp characters under complex background according to claim 1, characterized in that: A character region segmentation network with a CNN+RPN architecture is used to perform character region segmentation on sample images. The CNN network is used to extract the feature map of the image, and the RPN network slides the window on the image feature map to generate multiple region proposals for each window position. The RPN network contains a classifier to determine whether each region in the window contains stamped characters, and a bounding box regressor to fine-tune the position of the region.

7. The method for recognizing low-quality steel stamp characters under complex background according to claim 1, characterized in that: The improved target detection algorithm includes: Based on the traditional non-maximum suppression, the confidence scores of non-maximum detection boxes are reduced proportionally and retained instead of removed.

8. The method for recognizing low-quality steel stamp characters under complex background according to claim 1, characterized in that: The character recognition model includes a convolutional neural network (CNN), a recurrent neural network (RNN), and a connection temporal classification module (CTC), wherein: Convolutional neural network CNN uses multiple convolutional layers to extract features of images; The recurrent neural network RNN ​​uses a bidirectional LSTM to process the feature sequence in the character image and capture the temporal dependency to achieve context relevance and logic verification; a fully connected layer MLP is set at the last layer of the recurrent neural network module RNN to convert the output of the recurrent neural network module RNN into the final prediction result; The connected temporal classification module CTC uses the CTC loss function to transform the alignment problem between the sample image and the output character recognition result into a probability maximization problem.

9. An electronic device comprising a processor, a memory, and a computer program stored in the memory; characterized in that: When the processor executes the computer program, it implements the method for recognizing low-quality steel-printed characters under a complex background according to any one of claims 1 to 8.

10. A computer-readable storage medium, wherein a computer program is stored in the medium; characterized in that: When the computer program is executed by a processor, the method for recognizing low-quality steel-printed characters under a complex background according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Drug package multi-scale OCR (optical character recognition) and matching method oriented to complex background

    CN121121718A