OCR (Optical Character Recognition) method based on adaptive mapping and perception adjustment
Through the OCR recognition method of adaptive mapping and perceptual adjustment, the convolution kernel offset and receptive field are dynamically adjusted, which solves the recognition accuracy and efficiency of traditional OCR systems in complex environments, and achieves higher text recognition accuracy and real-time processing capabilities.
Patent Information
- Application Number
- CN202510489692.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-18
AI Technical Summary
Traditional OCR systems are not very accurate and efficient in processing texts of complex backgrounds, different sizes or densities, and are difficult to adapt, especially in text deformation, distortion or high dynamic range environments, resulting in increased recognition error rate and inaccurate positioning of bounding boxes.
The OCR recognition method of adaptive mapping and perceptual adjustment is adopted to extract feature maps through neural networks, calculate the convolution kernel offset, dynamically adjust the receptive field size and shape, and combine text density and character spacing to optimize the neural network model to improve recognition accuracy and efficiency.
It significantly improves the accuracy and efficiency of OCR recognition, can better adapt to different text layouts and formats, reduce background noise interference, and enhances the generalization ability and recognition rate of the model.
Smart Images

Figure CN120340050A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to an OCR recognition method based on adaptive mapping and perceptual adjustment. Background Art
[0002] In traditional optical character recognition (OCR) systems, fixed receptive fields and standard convolutional networks are usually used to process text recognition tasks in images.
[0003] The main problems faced by the prior art are that when dealing with text having complex backgrounds, different sizes or densities, the recognition accuracy and efficiency are often not high. Especially when the text is deformed, distorted or in a high dynamic range environment, the fixed receptive field and standard convolutional processing methods are difficult to adapt to these changes, resulting in an increase in the recognition error rate and inaccurate bounding box localization. In addition, existing OCR systems often lack an effective optimization mechanism when co-processing classification and regression tasks, making it difficult to achieve high-precision character classification and accurate bounding box regression simultaneously in practical applications, especially when the text size and spacing vary greatly. These problems limit the performance and practicality of traditional OCR technologies in a wider range of application scenarios. Summary of the Invention
[0004] To this end, the present invention provides an OCR recognition method based on adaptive mapping and perceptual adjustment to overcome the adaptability problems in the prior art when dealing with text deformation, distortion and in dynamic environments, thereby significantly improving the classification accuracy of characters and the positioning accuracy of bounding boxes.
[0005] To achieve the above object, the present invention provides an OCR recognition method based on adaptive mapping and perceptual adjustment, including:
[0006] Step S1, obtaining an OCR image to be processed, using a neural network model to extract features of the OCR image to obtain an original feature map, and also obtaining a historical OCR image as historical data;
[0007] Step S2, calculating the deviation between the original feature map and a predefined morphological model to adjust the offset of the convolutional kernel, and adjusting the feature point positions according to the offset;
[0008] Step S3, obtaining the text region of the OCR image, and calculating the text density and the average spacing between characters according to each text region;
[0009] Step S4, calculating the receptive field size according to the text density, and adjusting the receptive field shape according to the average spacing;
[0010] Step S5: Convolve the original feature map according to the receptive field size and the receptive field shape to obtain a target feature map;
[0011] Step S6: Train the neural network model according to the target feature map, the text density, and the average spacing to obtain a target OCR recognition model.
[0012] Further, the process of calculating the deviation between the original feature map and a predetermined morphological model to adjust the offset of the convolution kernel includes:
[0013] Compare the original feature map with the template in the morphological model to obtain a comparison result;
[0014] Analyze the comparison result to calculate the deviation between the original feature map and the template and adjust the offset.
[0015] Further, the process of analyzing the comparison result to calculate the deviation between the original feature map and the template and adjust the offset includes:
[0016] Offset(x,y) = σ(Conv(Input, W offset ) + Bias)
[0017] where Offset(x,y) represents the offset that the convolution kernel needs to adjust at the position of coordinates (x,y) on the original feature map. The Sigmoid function is used to map the convolution result to a target range. Conv(Input, W_offset) is the convolution operation performed on the OCR image Input using the weight matrix W_offset to calculate the offset, and Bias is the bias term used to adjust the baseline of the offset.
[0018] Further, the process of adjusting the position of the feature points according to the offset includes:
[0019] P adjusted (x,y) = Interpolation(P, x + Scale factor ·Δx, y + Scale factor ·Δy)
[0020] Among them, P_adjusted(x,y) represents the position of the adjusted feature point, used to obtain a more accurate character boundary. Interpolation(P,x',y') performs interpolation calculation on the original feature map P, where x' and y' are the new coordinates after offset adjustment. Scale_factor is the scale factor, and deltax and deltay are the specific values obtained from the offset calculation, representing the horizontal and vertical adjustments to be made on the original coordinates (x,y).
[0021] Furthermore, the process of step S3 includes:
[0022] Perform binarization and denoising on the OCR image to obtain a processed image;
[0023] Use edge detection to determine the edges of the processed image to obtain an edge result;
[0024] Perform dilation, erosion, opening, and closing operations on the edge result to obtain an operation result;
[0025] Detect the contours of the operation result to obtain a detection result;
[0026] Filter the text region according to the detection result and determine the text region area;
[0027] Determine the number of text pixels according to the processed image;
[0028] Take the ratio of the number of text pixels to the text region area to obtain the text density;
[0029] Use the vertical projection technique to determine the character positions according to the text region;
[0030] Measure the distance between adjacent characters according to the character positions to obtain a distance result, and take the average of all the distances to obtain the average spacing.
[0031] Furthermore, the process of calculating the receptive field size according to the text density includes:
[0032] R dynamic = α·log(1 + Density measure (P))·R initial
[0033] Among them, R_dynamic represents the dynamically adjusted receptive field size, Alpha represents the adjustment coefficient, log(1 + Density_measure(P)) represents converting the text density to a slowly growing logarithmic scale, Density_measure(P) represents the text density measurement function, and R_initial represents the initial receptive field size.
[0034] Further, the process of adjusting the receptive field shape according to the average spacing includes:
[0035] If the character spacing is uniform, a rectangular or square receptive field is adopted;
[0036] If the character spacing is non-uniform, an elliptical or adaptive-shaped receptive field is adopted;
[0037] For a specific font or layout style, the shape of the receptive field is customized according to its characteristics.
[0038] Further, the process of step S5 includes:
[0039] Generate a corresponding convolutional kernel according to the receptive field size and the receptive field shape;
[0040] Traverse the positions of the original feature map with the convolutional kernel;
[0041] Calculate the dot product of the convolutional kernel and the local area in the original feature map according to each position, and add the bias term to obtain the elements of the target feature map;
[0042] Perform activation and pooling processing on the elements of the target feature map to obtain the target feature map.
[0043] Further, the process of step S6 includes:
[0044] Divide the historical data into a training set and a validation set;
[0045] Initialize the loss weights lambda1 and lambda2 to balance the classification loss and the regression loss;
[0046] Define the weighted sum of the classification loss and the regression loss as the total loss;
[0047] Train the neural network model according to the training set to calculate the total loss;
[0048] Update the weights of the neural network model using gradient descent according to the total loss to obtain the target weights;
[0049] Validate the neural network model according to the target weights and the validation set to obtain the target OCR recognition model.
[0050] Further, the process of training the neural network model according to the training set to calculate the total loss includes:
[0051] Loss total = λ1·L cls(Softmax(Output cls )) + λ2·L reg (SmoothL1(Output reg ))
[0052] Among them, Loss_total represents the total loss, L_cls represents the classification loss function, Softmax(Output_cls) represents converting the original output of the classification task into a probability distribution through the Softmax function, L_reg represents the regression loss function, SmoothL1(Output_reg) represents applying the SmoothL1 loss to the output of the regression task, and lambda1 and lambda2 are loss weights used to adjust the classification loss and the regression loss.
[0053] Compared with the prior art, the beneficial effects of the present invention are as follows. The present invention can more accurately capture the features of the text by adaptively mapping to adjust the offset of the convolutional kernel, thereby improving the accuracy of OCR recognition. The perception adjustment mechanism can dynamically adjust the receptive field according to the text density and character spacing, enabling the model to better adapt to different text layouts and formats. By integrating these adaptive and perception adjustment mechanisms into the neural network model, while maintaining a high recognition rate, the dependence on complex post-processing steps can be reduced, thereby improving the efficiency of real-time OCR recognition. The adaptive adjustment of the feature map helps the model to better focus on the text area and reduce the influence of background noise and other non-text elements.
[0054] In particular, by comparing with the template in the morphological model, the key features in the original feature map can be more accurately identified and located, which helps to improve the accuracy of the OCR system in the feature extraction stage. Adjusting the offset of the convolutional kernel according to the comparison result enables the feature extraction process to adapt to different types of text features and layouts. Accurate feature matching and extraction directly affect the subsequent classification and recognition stages, thereby improving the overall OCR recognition rate.
[0055] In particular, by finely adjusting the offset of the convolutional kernel, the model can more accurately locate the position of the text features, thereby improving the accuracy of feature extraction. The text in the image may have various deformations, such as tilting and distortion. Dynamically adjusting the offset of the convolutional kernel can help the model better adapt to these deformations and improve the recognition rate.
[0056] In particular, by finely adjusting the position of the feature points, the model can more precisely locate the boundaries and key features of the characters, thereby improving the accuracy of feature extraction. In an image with a complex background or a lot of noise, by adjusting the position of the feature points, the model can better focus on the text area and reduce the interference of background noise.
[0057] In particular, through binarization and denoising processing, the noise and interference in the image are reduced, and the clarity of the image is improved. Through binarization and denoising processing, the noise and interference in the image are reduced, and the clarity of the image is improved. By measuring the distance between adjacent characters and calculating the average spacing, important information about the text layout can be provided for the model, which helps to improve the model's adaptability to text spacing changes. By measuring the distance between adjacent characters and calculating the average spacing, important information about the text layout can be provided for the model, which helps to improve the model's adaptability to text spacing changes.
[0058] In particular, by adjusting the receptive field size according to the text density, the model can more accurately capture the features of texts with different densities, thereby improving the recognition accuracy, especially in regions with dense or sparse text. The dynamic receptive field helps the model learn more general and robust features, rather than just adapting to the specific text density in the training data, thus reducing the risk of overfitting. Using a smaller receptive field in sparse text regions can reduce unnecessary computations, while using a larger receptive field in dense text regions can better capture complex features, thereby improving the computational efficiency overall.
[0059] In particular, by adopting a receptive field shape that matches the character spacing, the model can more accurately segment adjacent characters, especially when the character spacing is uniform or non-uniform. In the case of non-uniform character spacing, using an elliptical or adaptive-shaped receptive field can reduce misrecognition caused by character spacing changes. By using a receptive field that is more suitable for character spacing and shape, unnecessary computations can be reduced, thus speeding up the recognition process.
[0060] In particular, by designing an appropriate receptive field size and shape, the convolutional kernel can efficiently extract key features in the image, improving the efficiency of feature extraction. By calculating the dot product of the convolutional kernel and the original feature map and combining the bias term, the model can learn a richer and more representative feature representation, which helps to improve the recognition accuracy. The introduction of the activation function adds non-linear characteristics to the feature map, enabling the model to capture more complex patterns and relationships and enhancing the generalization ability of the model. The pooling operation reduces the spatial size of the feature map, lowers the computational complexity of the model, and at the same time retains the most important feature information, helping to prevent overfitting.
[0061] In particular, by dividing the data into a training set and a validation set, the model can verify its performance on independent data, which helps to improve the generalization ability of the model and reduce overfitting. By defining the total loss as a weighted sum of the classification loss and the regression loss, the model can simultaneously focus on the two tasks of text classification and location regression, thereby optimizing the overall performance. Using the gradient descent algorithm to update the model weights according to the total loss can effectively optimize the model parameters and reduce recognition errors. By transforming the complex optimization problem into the minimization problem of the total loss, the training process is simplified, making the model training more efficient.
[0062] In particular, by simultaneously optimizing the classification and regression losses, the model can more comprehensively learn the data features, so that while identifying the text category, it can also accurately predict the location information of the text, which is particularly important in OCR applications. By combining the classification and regression tasks in a total loss function, end-to-end training can be achieved, simplifying the training process and potentially improving the training efficiency. Brief Description of the Drawings
[0063] Figure 1 It is a schematic flow chart of an OCR recognition method based on adaptive mapping and perception adjustment provided by an embodiment of the present invention;
[0064] Figure 2 It is a schematic flow chart of step S3 in an OCR recognition method based on adaptive mapping and perception adjustment provided by an embodiment of the present invention;
[0065] Figure 3 It is a schematic flow chart of step S5 in an OCR recognition method based on adaptive mapping and perception adjustment provided by an embodiment of the present invention;
[0066] Figure 4 It is a schematic flow chart of step S6 in an OCR recognition method based on adaptive mapping and perception adjustment provided by an embodiment of the present invention; Detailed Embodiments
[0067] In order to make the objectives and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0068] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0069] It should be noted that in the description of the present invention, the terms indicating the direction or positional relationship such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the direction or positional relationship shown in the drawings. This is only for convenience of description, rather than indicating or implying that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.
[0070] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and defined, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0071] Please refer to Figure 1 As shown, an OCR recognition method based on adaptive mapping and perception adjustment provided by an embodiment of the present invention includes:
[0072] Step S1, obtaining an OCR image to be processed, using a neural network model to extract the features of the OCR image to obtain an original feature map, and also obtaining a historical OCR image as historical data;
[0073] Step S2, calculating the deviation between the original feature map and a predefined morphological model to adjust the offset of the convolution kernel, and adjusting the feature point positions according to the offset;
[0074] Step S3, obtaining the text region of the OCR image, and calculating the text density and the average spacing between characters according to each text region;
[0075] Step S4, calculating the receptive field size according to the text density, and adjusting the receptive field shape according to the average spacing;
[0076] Step S5, performing convolution on the original feature map according to the receptive field size and the receptive field shape to obtain a target feature map;
[0077] Step S6, training the neural network model according to the target feature map, the text density, and the average spacing to obtain a target OCR recognition model.
[0078] Specifically, load the OCR image to be processed. Ensure that the image format, size, and resolution are suitable for neural network input. Use a pre-trained neural network model (such as CNN) to extract features from the OCR image. Obtain the original feature map, which contains a high-level abstract representation of the image. Collect and annotate historical OCR image data. Preprocess the historical data to match the processing flow of the current OCR image. Calculate the difference between the feature map and the morphological model to obtain a deviation value. Adjust the offset of the convolutional kernel according to the deviation value. Apply the adjusted convolutional kernel to recalculate the position of the feature points. Use image segmentation techniques (such as connected component analysis) to identify the text regions in the OCR image. For each text region, calculate the ratio of the pixels occupied by the characters to the total area of the region. Identify the gaps between characters, calculate and average these spacings. Determine the size of the receptive field according to the text density to adapt to texts with different densities. Adjust the shape of the receptive field according to the average spacing between characters to ensure that the receptive field can adapt to characters with different spacings. Perform a convolutional operation on the original feature map using the adjusted receptive field size and shape. Obtain the target feature map, which is more suitable for subsequent recognition tasks. Use the target feature map, text density, and average spacing as input data. Design a loss function that combines classification loss (such as cross-entropy) and regression loss (such as SmoothL1). Set loss weights (lambda1 and lambda2) to balance the classification and regression tasks. Use backpropagation and optimization algorithms (such as Adam) to update the weights. Monitor the performance of the validation set during training and perform hyperparameter tuning. When the performance of the validation set reaches the best, save the model parameters.
[0079] Specifically, by adaptively mapping to adjust the offset of the convolutional kernel, it is possible to more accurately capture the features of the text, thereby improving the accuracy of OCR recognition. The perception adjustment mechanism can dynamically adjust the receptive field according to the text density and character spacing, enabling the model to better adapt to different text layouts and formats. By integrating these adaptive and perception adjustment mechanisms into the neural network model, it is possible to reduce the dependence on complex post-processing steps while maintaining a high recognition rate, thereby improving the efficiency of real-time OCR recognition. The adaptive adjustment of the feature map helps the model better focus on the text region, reducing the influence of background noise and other non-text elements.
[0080] Specifically, the process of calculating the deviation between the original feature map and a predetermined morphological model to adjust the offset of the convolutional kernel includes:
[0081] Compare the original feature map with the template in the morphological model to obtain a comparison result;
[0082] Analyze the comparison result to calculate the deviation between the original feature map and the template and adjust the offset.
[0083] Specifically, define templates in the morphological model. These templates represent patterns of different text features, such as lines, corners, textures, etc. Ensure that the size of the template matches the area to be compared in the original feature map. For each area of the original feature map, perform a sliding window comparison using the template. At each window position, calculate the similarity between the template and the feature map area. Correlation coefficient, mutual information, or other similarity metrics can be used. Record the comparison results at each window position to form a comparison result map, where each point represents the similarity between the template and the feature map. Analyze the comparison result map to identify areas with low similarity, which may indicate deviations between the feature map and the template. Determine the degree and direction of the deviation, which can be achieved by calculating the local gradient or using other image processing techniques. Quantify the analyzed deviation into numerical values, which will be used to adjust the offset of the convolutional kernel. The quantization of the deviation can be a simple numerical mapping or the result of a more complex regression analysis. According to the quantified deviation value, map it to the adjustment of the convolutional kernel offset. For example, if the deviation indicates that the text line in the feature map is shifted to the right, then adjust the offset of the convolutional kernel accordingly. Implement the adjustment of the offset in the neural network, which may involve the generation of dynamic convolutional kernels or the real-time update of existing convolutional kernel parameters. Ensure that the adjusted convolutional kernel can be correctly applied to the original feature map to capture the correct feature point positions.
[0084] Specifically, by comparing with the templates in the morphological model, key features in the original feature map can be more accurately identified and located, which helps to improve the accuracy of the OCR system in the feature extraction stage. Adjust the offset of the convolutional kernel according to the comparison results, so that the feature extraction process can adapt to different types of text features and layouts. Accurate feature matching and extraction directly affect the subsequent classification and recognition stages, thus improving the overall OCR recognition rate.
[0085] Specifically, the process of analyzing and calculating the deviation between the original feature map and the template according to the comparison results to adjust the offset includes:
[0086] Offset(x,y)=σ(Conv(Input,W offset )+Bias)
[0087] Where Offset(x,y) represents the offset that the convolutional kernel needs to adjust at the position of coordinates (x,y) on the original feature map. The Sigmoid function is used to map the convolutional result to a target range. Conv(Input,W_offset) is the convolutional operation performed on the OCR image Input using the weight matrix W_offset to calculate the offset, and Bias is the bias term used to adjust the baseline of the offset.
[0088] Specifically, the target range is usually from -1 to 1. The Sigmoid function makes the offset more smooth and easy to control, and the Bias ensures the centrality and balance of the offset. During the training process, the weight matrix W_offset and the bias term Bias are updated through the backpropagation algorithm. The goal is to minimize the error of OCR recognition. Through multiple iterations, the model will learn how to automatically adjust the offset of the convolution kernel according to the features of the input image.
[0089] Specifically, by finely adjusting the offset of the convolution kernel, the model can more accurately locate the position of text features, thereby improving the accuracy of feature extraction. Texts in images may have various deformations, such as tilting, distortion, etc. Dynamically adjusting the offset of the convolution kernel can help the model better adapt to these deformations and improve the recognition rate.
[0090] Specifically, the process of adjusting the position of the feature points according to the offset includes:
[0091] P adjusted (x,y) = Interpolation(P,x+Scale factor ·Δx,y+Scale factor ·Δy)
[0092] Among them, P_adjusted(x,y) represents the adjusted position of the feature point, used to obtain a more accurate character boundary. Interpolation(P,x',y') is to perform interpolation calculation on the original feature map P. x' and y' are the new coordinates after offset adjustment. Scale_factor is the scale factor, and deltax and deltay are the specific values obtained from the offset calculation, indicating the horizontal and vertical adjustments that need to be made on the original coordinates (x,y).
[0093] Specifically, a scale factor is determined, which is used to control the influence degree of the offset on the adjustment of the feature point position. The scale factor is usually a positive number less than 1 to avoid excessive adjustment. Using the previously calculated offset formula, calculate the horizontal (Δx) and vertical (Δy) offsets for each feature point. For each feature point P(x,y) in the original feature map P, calculate the adjusted coordinates (x’,y’) according to the following formula:
[0094] x’ = x + Scale_factor * Δx; y’ = y + Scale_factor * Δy.
[0095] Use the interpolation method (such as bilinear interpolation, bicubic interpolation, etc.) to calculate the value of the adjusted coordinates (x’,y’) on the original feature map P. This process can be expressed as:
[0096] P_adjusted(x,y) = Interpolation(P,x’,y’)
[0097] Steps of bilinear interpolation: Determine the pixel grid where the adjusted coordinates (x’, y’) are located. Find the four adjacent pixel points (i, j) of this pixel grid: The pixel point at the upper left corner of the target coordinates. (i, j + 1): The pixel point at the upper right corner of the target coordinates. (i + 1, j): The pixel point at the lower left corner of the target coordinates. (i + 1, j + 1): The pixel point at the lower right corner of the target coordinates. Among them, i and j are integers, representing the integer coordinates closest to x’ and y’ respectively.
[0098] Calculate the horizontal and vertical distances of the target coordinates (x’, y’) relative to its upper left pixel point (i, j), that is, the interpolation weights u and v:
[0099] u = x’ - i; v = y’ - j.
[0100] These weights represent the relative position of the target coordinates (x’, y’) within the pixel grid. According to the interpolation weights and the values of the adjacent pixel points, calculate the interpolation result:
[0101] P_adjusted(x’, y’) = (1 - u)(1 - v)P(i, j) + u(1 - v)P(i, j + 1) + (1 - u)vP(i + 1, j) + uvP(i + 1, j + 1)
[0102] Among them, this formula means that the interpolation result is obtained by weighted averaging the values of the four adjacent pixel points, where the weights (1 - u) and (1 - v) correspond to the horizontal and vertical distances within the pixel grid respectively.
[0103] Specifically, by making fine-grained adjustments to the positions of the feature points, the model can more accurately locate the boundaries and key features of the characters, thereby improving the accuracy of feature extraction. In images with complex backgrounds or a lot of noise, by adjusting the positions of the feature points, the model can better focus on the text area and reduce the interference of background noise.
[0104] Specifically, as Figure 2 shown, the process of step S3 includes:
[0105] Step S31, perform binarization and denoising processing on the OCR image to obtain a processed image;
[0106] Step S32, use edge detection to determine the edges of the processed image to obtain an edge result;
[0107] Step S33, perform operations of dilation, erosion, opening, and closing on the edge result to obtain an operation result;
[0108] Step S34, detecting the contour of the operation result to obtain a detection result;
[0109] Step S35, screening the text area according to the detection result and determining the text area size, and determining the number of text pixels according to the processed image;
[0110] Step S36, calculating the ratio of the number of text pixels to the text area size to obtain the text density;
[0111] Step S37, determining the character positions according to the text area using the vertical projection technique;
[0112] Step S38, measuring the distance between adjacent characters according to the character positions to obtain a distance result, and averaging all the distances to obtain the average spacing.
[0113] Specifically, read the OCR image and convert it into a grayscale image. Apply a binarization algorithm, such as Otsu's method, to select a threshold to convert the grayscale image into a binary image. In the binary image, the pixel values are usually 0 (black) or 255 (white). Denoise the binary image using median filtering or Gaussian filtering to reduce the noise caused by scanning or shooting. Use Gaussian filtering to smooth the image to remove noise. Calculate the gradient magnitude and direction of each pixel point in the image. Apply non-maximum suppression to thin the edges. Use double thresholds to determine real and potential edges. Determine the edges through edge tracking and hysteresis thresholding. Dilate the edge image using a structuring element to connect adjacent edges. Erode the edge image using a structuring element to eliminate small noise points. Erode first and then dilate to remove small objects and smooth the boundaries of larger objects. Dilate first and then erode to fill small holes or breaks and connect adjacent objects. Use the findContours function in OpenCV to find the contours on the morphologically processed image. The contours are a sequence of continuous points, and each point has the same color or intensity. Screen out the external contours and ignore the internal hole contours. Screen out the text area according to the characteristics of the contour area, aspect ratio, shape, etc. Use the contourArea function to calculate the area of each screened contour. Traverse each pixel in the binary image: if the pixel belongs to the text area (usually white pixels), then count. Accumulate to get the total number of text pixels. Calculate the text density using the following formula:
[0114] Text density = number of text pixels / total text area.
[0115] Perform a vertical projection on each text line and count the number of white pixels in each column. The peaks in the projection usually correspond to the positions of the characters. Based on the results of the vertical projection, determine the starting and ending positions of the characters. Calculate the distance between the centers of adjacent characters. Take the average of the distances between all adjacent characters to obtain the average spacing.
[0116] Specifically, through binarization and denoising processing, the noise and interference in the image are reduced, and the clarity of the image is improved. By measuring the distance between adjacent characters and calculating the average spacing, important information about the text layout can be provided for the model, which helps to improve the adaptability of the model to text spacing changes. By measuring the distance between adjacent characters and calculating the average spacing, important information about the text layout can be provided for the model, which helps to improve the adaptability of the model to text spacing changes.
[0117] Specifically, the process of calculating the receptive field size according to the text density includes:
[0118] R dynamic = α · log(1 + Density measure (P)) · R initial
[0119] where R_dynamic represents the dynamically adjusted receptive field size, Alpha represents the adjustment coefficient, log(1 + Density_measure(P)) represents converting the text density to a slowly growing logarithmic scale, Density_measure(P) represents the text density measurement function, and R_initial represents the initial receptive field size.
[0120] Specifically, determine the initial receptive field size (R_initial), which is the default receptive field size used by the model when there is no density information. Determine the adjustment coefficient (α), which is a preset coefficient used to balance the change range of the receptive field size and ensure that the adjustment does not overly affect other parts of the image. Substitute the measured text density value into the logarithmic function to convert it to a slowly growing logarithmic scale. This ensures that even with large changes in text density, the adjustment of the receptive field size is smooth and gradual. The calculation formula is: log(1 + Density_measure(P)). Here, 1 is to ensure that the logarithmic function is still defined when Density_measure(P) is 0 and to avoid the undefined problem of the logarithmic function. Combine the adjustment coefficient, the result of text density conversion, and the initial receptive field size to obtain a receptive field size adjusted according to the current text density. Apply the calculated R_dynamic to the convolutional layer in the neural network to adjust its receptive field size. This can be achieved by changing the size of the convolutional kernel, the stride, or using convolutional kernels of different sizes.
[0121] Specifically, for text-dense regions, the system reduces the size of the receptive field. This is because in character-dense regions, a smaller receptive field can more precisely capture the details of individual characters and avoid interference from adjacent characters. For text-sparse regions, the system increases the size of the receptive field. A larger receptive field helps the system capture more background information, which is especially important for understanding the context of characters and improving the accuracy of character recognition.
[0122] Specifically, by adjusting the receptive field size according to text density, the model can more precisely capture the features of text with different densities, thereby improving the recognition accuracy, especially in text-dense or sparse regions. The dynamic receptive field helps the model learn more general and robust features instead of just adapting to the specific text density in the training data, thus reducing the risk of overfitting. Using a smaller receptive field in text-sparse regions can reduce unnecessary calculations, while using a larger receptive field in text-dense regions can better capture complex features, thereby improving the overall computational efficiency.
[0123] Specifically, the process of adjusting the receptive field shape according to the average spacing includes:
[0124] If the character spacing is uniform, use a rectangular or square receptive field;
[0125] If the character spacing is non-uniform, use an oval or adaptive-shaped receptive field;
[0126] For a specific font or layout style, customize the shape of the receptive field according to its characteristics.
[0127] Specifically, using the average inter-character spacing calculated previously, analyze the uniformity of the character spacing. The uniformity of the spacing can be evaluated by calculating the standard deviation or coefficient of variation of the character spacing. Set a threshold to determine whether the character spacing is uniform. If the standard deviation or coefficient of variation is less than the threshold, the character spacing is considered uniform; otherwise, it is considered non-uniform. If the character spacing is uniform, select a rectangular or square receptive field. A rectangular receptive field is suitable for cases where there is a large difference in character height and width, and a square receptive field is suitable for cases where the character sizes are relatively consistent. Adjust the size of the receptive field to match the size of the characters. If the character spacing is non-uniform: Adopt an elliptical receptive field, the major axis of which can better adapt to the variation of the character spacing. Alternatively, use a receptive field with an adaptive shape, which can be dynamically adjusted according to the actual shape and spacing of the characters. Analyze the characteristics of a specific font or typesetting style, such as font width, inclination, decoration, etc. Design a customized receptive field shape based on these characteristics. For example: For italic fonts, a parallelogram receptive field can be adopted. For highly decorative fonts, a more complex polygon receptive field can be used. Parameterize the shape of the receptive field. For example, for an ellipse, it can be parameterized as the major axis, minor axis, and rotation angle. Develop an algorithm to adjust the receptive field according to the parameterized shape. The algorithm may include the following steps: Select a basic shape (rectangle, square, ellipse, or custom shape) according to the character spacing analysis result. Adjust the shape parameters according to the character features (such as size, spacing, inclination). Apply the shape parameters to the convolutional kernel to form the final receptive field.
[0128] Specifically, by adopting a receptive field shape that matches the character spacing, the model can more accurately segment adjacent characters, especially in cases where the character spacing is uniform or non-uniform. In the case of non-uniform character spacing, using an elliptical or adaptive-shaped receptive field can reduce misidentifications caused by variations in character spacing. By using a receptive field that is more suitable for character spacing and shape, unnecessary calculations can be reduced, thereby accelerating the recognition process.
[0129] Specifically, as Figure 3 shown, the process of step S5 includes:
[0130] Step S51, generate a corresponding convolutional kernel according to the receptive field size and the receptive field shape;
[0131] Step S52, traverse the positions of the original feature map with the convolutional kernel;
[0132] Step S53, calculate the dot product of the convolutional kernel and the local region in the original feature map according to each position, and add the bias term to obtain the elements of the target feature map;
[0133] Step S54, perform activation and pooling operations on the elements of the target feature map to obtain the target feature map.
[0134] Specifically, according to the calculated dynamic receptive field size and shape, determine the size and shape of the convolutional kernel. Randomly initialize the weights of the convolutional kernel, or use pre-trained weights. If necessary, further adjust the weights of the convolutional kernel according to the text density or character spacing to ensure that it can effectively capture text features. Select an appropriate stride to traverse the original feature map. The stride determines the moving distance of the convolutional kernel on the feature map. Place the convolutional kernel at the starting position of the original feature map. For the local area of the original feature map covered by the convolutional kernel, calculate the dot product of the convolutional kernel weights and the corresponding elements of the feature map. Add a bias term to the calculated dot product result. The bias term is a learnable parameter used to adjust the activation level of the output feature map. Pass the result of adding the bias term to the dot product through a non-linear activation function, such as ReLU, Sigmoid, or Tanh, to introduce non-linearity and enhance the expressive power of the model. Select max pooling, average pooling, or other pooling methods as needed. Apply the pooling operation to the activated feature map to reduce the size of the feature map while retaining the most important information. Summarize the elements of the activated and pooled feature map to form the target feature map.
[0135] Specifically, by designing an appropriate receptive field size and shape, the convolutional kernel can efficiently extract key features in the image and improve the efficiency of feature extraction. By calculating the dot product of the convolutional kernel and the original feature map and combining the bias term, the model can learn richer and more representative feature representations, which helps to improve the recognition accuracy. The introduction of the activation function adds non-linear characteristics to the feature map, enabling the model to capture more complex patterns and relationships and enhancing the generalization ability of the model. The pooling operation reduces the spatial size of the feature map, reduces the computational complexity of the model, and retains the most important feature information, which helps to prevent overfitting.
[0136] Specifically, as Figure 4 shown, the process of step S6 includes:
[0137] Step S61, divide the historical data into a training set and a validation set;
[0138] Step S62, initialize the loss weights lambda1 and lambda2 to balance the classification loss and the regression loss;
[0139] Step S63, define the weighted sum of the classification loss and the regression loss as the total loss;
[0140] Step S64: Train the neural network model according to the training set to calculate the total loss;
[0141] Step S65: Update the weights of the neural network model using gradient descent according to the total loss to obtain the target weights;
[0142] Step S66: Validate the neural network model according to the target weights and the validation set to obtain the target OCR recognition model.
[0143] Specifically, collect a large amount of OCR recognition historical data, including images and corresponding correct text labels. Use the method of random partitioning to divide the historical data into two main parts: the training set and the validation set. Usually, the training set accounts for 80%-90% of the total data, while the validation set accounts for the remaining part. Initialize two loss weights lambda1 and lambda2. These weights are used to balance the importance of classification loss and regression loss during the training process. For example, lambda1 can be set to 1 and lambda2 can be set to 0.5. The specific values need to be adjusted according to the specific requirements of the task and the experimental results. For each batch in the training set, perform forward propagation through the network to calculate the predicted values. Calculate the classification loss Lcls and the regression loss Lreg according to the predicted values and the actual labels. Calculate the total loss of the current batch. Perform backpropagation on the calculated total loss to calculate the gradient of the loss with respect to the network parameters. Use a gradient descent algorithm (such as Stochastic Gradient Descent SGD, Adam, etc.) to update the weights of the network according to the calculated gradient. This process can be expressed as:
[0144] weights = weights - learning_rate * gradient
[0145] where weights are the weights of the neural network, learning_rate is the learning rate, and gradient is the gradient. Evaluate the performance of the model on the validation set, calculate the total loss and other relevant performance metrics (such as accuracy, F1 score, etc.). According to the performance feedback on the validation set, adjust hyperparameters such as loss weights and learning rate to optimize the model. The model weights that perform best on the validation set are selected as the final target weights for constructing the target OCR recognition model.
[0146] Specifically, by dividing the data into a training set and a validation set, the model can verify its performance on independent data, which helps improve the generalization ability of the model and reduce overfitting. By defining the total loss as a weighted sum of the classification loss and the regression loss, the model can simultaneously focus on the two tasks of text classification and location regression, thereby optimizing the overall performance. Using the gradient descent algorithm to update the model weights according to the total loss can effectively optimize the model parameters and reduce recognition errors. By transforming the complex optimization problem into the minimization problem of the total loss, the training process is simplified, making the model training more efficient.
[0147] Specifically, the process of training the neural network model according to the training set to calculate the total loss includes:
[0148] Loss total = λ1·L cls (SOftmax(Output cls )) + λ2·L reg (SmoothL1(Output reg ))
[0149] Where Loss_total represents the total loss, L_cls represents the classification loss function, Softmax(Output_cls) represents converting the original output of the classification task into a probability distribution through the Softmax function, L_reg represents the regression loss function, SmoothL1(Output_reg) represents applying the SmoothL1 loss to the output of the regression task, and lambda1 and lambda2 are loss weights used to adjust the classification loss and the regression loss.
[0150] Specifically, by simultaneously optimizing the classification and regression losses, the model can more comprehensively learn the data features, so that while identifying the text category, it can also accurately predict the location information of the text, which is particularly important in OCR applications. By combining the classification and regression tasks in a total loss function, end-to-end training can be achieved, simplifying the training process and potentially improving the training efficiency.
[0151] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
[0152] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention; for those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An OCR recognition method based on adaptive mapping and perception adjustment, characterized in that Including: Step S1: Obtain the OCR image to be processed, use a neural network model to extract the features of the OCR image to obtain an original feature map, and also obtain a historical OCR image as historical data; Step S2: Calculate the deviation between the original feature map and a predefined morphological model to adjust the offset of the convolutional kernel, and adjust the feature point positions according to the offset; Step S3: Obtain the text regions of the OCR image, and calculate the text density and the average spacing between characters according to each text region; Step S4: Calculate the receptive field size according to the text density, and adjust the receptive field shape according to the average spacing; Step S5: Convolve the original feature map according to the receptive field size and the receptive field shape to obtain a target feature map; Step S6: Train the neural network model according to the target feature map, the text density, and the average spacing to obtain a target OCR recognition model.
2. The OCR recognition method based on adaptive mapping and perception adjustment according to claim 1, characterized in that The process of calculating the deviation between the original feature map and a predefined morphological model to adjust the offset of the convolutional kernel includes: Comparing the original feature map with the template in the morphological model to obtain a comparison result; Analyzing according to the comparison result to calculate the deviation between the original feature map and the template to adjust the offset.
3. The OCR recognition method based on adaptive mapping and perception adjustment according to claim 2, wherein The process of analyzing according to the comparison result to calculate the deviation between the original feature map and the template to adjust the offset includes: Offset(x,y)=σ(Conv(Input,W offset )+Bias) Among them, Offset(x,y) represents the offset that the convolutional kernel needs to adjust at the position of coordinates (x,y) on the original feature map. The Sigmoid function is used to map the convolutional result to a target range. Conv(Input,W_offset) is the convolution operation performed on the OCR image Input using the weight matrix W_offset to calculate the offset, and Bias is the bias term used to adjust the baseline of the offset.
4. The OCR recognition method based on adaptive mapping and perception adjustment according to claim 3, wherein The process of adjusting the feature point positions according to the offset includes: P adjusted (x,y) = Interpolation(P, x + Scale factor ·Δx, y + Scale factor ·Δy). Here, P_adjusted(x,y) represents the position of the adjusted feature point, used to obtain a more accurate character boundary. Interpolation(P,x',y') is to perform interpolation calculation on the original feature map P, where x' and y' are the new coordinates after offset adjustment, Scale_factor is the scale factor, and deltax and deltay are the specific values obtained from the offset calculation, indicating the horizontal and vertical adjustments to be made on the original coordinates (x,y).
5. The OCR recognition method based on adaptive mapping and perception adjustment according to claim 4, wherein The process of Step S3 includes: Performing binarization and denoising processing on the OCR image to obtain a processed image; Using edge detection to determine the edges of the processed image to obtain an edge result; Performing dilation, erosion, opening, and closing operations on the edge result to obtain an operation result; Detecting the contours of the operation result to obtain a detection result; Screening the text regions according to the detection result and determining the text region area; Determining the number of text pixels according to the processed image; Calculating the ratio of the number of text pixels to the text region area to obtain the text density; Using the vertical projection technique to determine the character positions according to the text regions; Measuring the distances between adjacent characters according to the character positions to obtain a distance result, and averaging all the distances to obtain the average spacing.
6. The OCR recognition method based on adaptive mapping and perception adjustment according to claim 5, characterized in that The process of calculating the receptive field size according to the text density includes: R dynamic = α·log(1 + Density measure (P))·R initial Among them, R_dynamic represents the size of the receptive field after dynamic adjustment, Alpha represents the adjustment coefficient, log(1 + Density_measure(P)) represents converting the text density to a slowly growing logarithmic scale, Density_measure(P) represents the text density measurement function, and R_initial represents the initial receptive field size.
7. The OCR recognition method based on adaptive mapping and perception adjustment according to claim 6, wherein The process of adjusting the receptive field shape according to the average spacing includes: If the character spacing is uniform, use a rectangular or square receptive field; If the character spacing is non-uniform, use an elliptical or adaptive-shaped receptive field; For a specific font or layout style, customize the shape of the receptive field according to its characteristics.
8. The OCR recognition method based on adaptive mapping and perception adjustment according to claim 7, wherein The process of step S5 includes: Generate a corresponding convolution kernel according to the receptive field size and the receptive field shape; Traverse the positions of the original feature map with the convolution kernel; Calculate the dot product of the convolution kernel and the local area in the original feature map according to each position, and add the bias term to obtain the elements of the target feature map; Perform activation and pooling processing on the elements of the target feature map to obtain the target feature map.
9. The OCR recognition method based on adaptive mapping and perception adjustment according to claim 8, wherein The process of step S6 includes: Divide the historical data into a training set and a validation set; Initialize the loss weights lambda1 and lambda2 to balance the classification loss and the regression loss; Define the weighted sum of the classification loss and the regression loss as the total loss; Train the neural network model according to the training set to calculate the total loss; Update the weights of the neural network model using gradient descent according to the total loss to obtain the target weights; Validate the neural network model according to the target weights and the validation set to obtain the target OCR recognition model.
10. The OCR recognition method based on adaptive mapping and perception adjustment according to claim 9, characterized in that, The process of training the neural network model according to the training set to calculate the total loss includes: Loss total = λ1·L cls (Softmax(Output cls )) + λ2·L reg (SmoothL1(Output reg )) Among them, Loss_total represents the total loss, L_cls represents the classification loss function, Softmax(Output_cls) represents converting the original output of the classification task into a probability distribution through the Softmax function, L_reg represents the regression loss function, SmoothL1(Output_reg) represents applying the SmoothL1 loss to the output of the regression task, and lambda1 and lambda2 are loss weights used to adjust the classification loss and the regression loss.
Citation Information
Patent Citations
OCR-based image analysis method, system and device, and medium
CN111539412A
Super-resolution text image recognition method and device, equipment and storage medium
CN114170608A
Traditional Chinese text recognition method and system for historical materials
CN118411727A
Text region determination method and device, electronic equipment and storage medium
CN118887677A
Cited By
Automatic data rechecking system based on visual scanning
CN120997638A