Method for Identifying Readable Texts of Commodities Based on Pixel-Level Attention Mechanism
By introducing a pixel point-level attention mechanism in OCR technology, the readable confidence of text blocks is calculated and readable text blocks are screened, and the problem of identifying errors in the prior art is solved, and the recognition accuracy and screening ability are improved.
Patent Information
- Application Number
- CN202111550512.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-12-17
AI Technical Summary
When existing OCR technology recognizes text in product pictures, it is easy to misidentify patterns or unreadable parts as text, resulting in incorrect text sequences and affects blind people's understanding of information.
Using a method based on the pixel point-level attention mechanism, the coordinates and content of text blocks are obtained through OCR technology, the readable confidence of each text block is calculated using the pixel point-level attention mechanism, and the readable text blocks are filtered out through adaptive thresholds.
It improves the accuracy of text recognition, reduces the problem of single correction results caused by corpus restrictions, and enhances the ability to screen readable text for product images.
Smart Images

Figure CN114359908B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a method for discriminating readable text of commodities based on a pixel-level attention mechanism. Background Art
[0002] Identifying the text in a commodity picture and converting it into a readable text sequence is beneficial to conveying the information of the commodity to the blind. However, when the existing OCR technology identifies the text in a commodity picture, it will misidentify some patterns or other unreadable parts as text, and finally form an incorrect text sequence, which affects the blind's understanding of the information.
[0003] The solution of the existing technology is to eliminate the incoherent information through a pre-trained language model when converting the picture into a readable text sequence. However, this method is highly dependent on the corpus, and the correction result is single, without fully utilizing the visual information in the picture.
[0004] The above problems need to be solved urgently. Summary of the Invention
[0005] The present invention aims to overcome the above-mentioned drawbacks of the prior art and provides a method for discriminating readable text of commodities based on a pixel-level attention mechanism.
[0006] To solve the above technical problems, the method for discriminating readable text of commodities based on a pixel-level attention mechanism includes the following steps:
[0007] S110. Obtain the coordinates and corresponding text of the text blocks on the commodity picture through OCR technology;
[0008] S120. Use the pixel-level attention mechanism to obtain the attention weights of all pixel units in the picture;
[0009] S130. Calculate the readable confidence of each text block using the attention weights;
[0010] S140. Set a confidence threshold, and filter out the readable text blocks according to the threshold when constructing the text output corresponding to the commodity picture.
[0011] Further, the obtaining of the coordinates and corresponding text of the text blocks on the commodity picture through OCR technology in step S110 specifically includes:
[0012] S1101. Preprocess the commodity picture by using Gaussian filtering for noise reduction and mean smoothing;
[0013] S1102. Use the PSENet model to perform text detection on the commodity picture to obtain the coordinate set of n text blocks in the commodity picture
[0014] S1103, based on the text block coordinate set Cut out the sub-image set corresponding to each text block
[0015] S1104, using the CRNN model to perform text recognition on the cut text block sub-graphs, and obtain the text content corresponding to the n text blocks in the product image
[0016] Furthermore, the step S120 of using the pixel-level attention mechanism to obtain the attention weights of all pixel units in the image specifically includes:
[0017] S1201, for the product image pixel set Z, pixel z i The initial attention weight corresponding to ∈Z is determined by other pixels in the whole image. The calculation method of the initial attention weight matrix of the whole image is A=ZWZ T ;
[0018] S1202, after obtaining the preliminary attention weight matrix of the whole image, the preliminary attention weight value of each pixel point is processed by an activation function to obtain the attention weight value corresponding to each pixel point, and the calculation method is A′=sigmoid(A).
[0019] Furthermore, step S130 uses the attention weight to calculate the readability confidence of each text block, which specifically includes:
[0020] S1301, obtain the center point coordinates T corresponding to each text block = {(x1, y1)…(x n ,y n )};
[0021] S1302, for the text block k, with the center point as the origin, apply the bivariate normal distribution formula to obtain the bivariate normal distribution probability value P corresponding to each pixel point inside the text block k = {p1…p n};
[0022] S1303, for each pixel in the text block, the probability is weighted averaged using the weight obtained by the attention mechanism, that is, The calculation result is used as the readability confidence of each text block;
[0023] Furthermore, the setting of the confidence threshold in step S140 specifically includes:
[0024] S141: Setting a confidence threshold, including:
[0025] S1401, setting the initial confidence threshold t0 to 0.5;
[0026] S1402, During the training process, fix other parameters of the algorithm, attempt to reduce the threshold with the initial amplitude t0, and test the screening results of readable text using this threshold;
[0027] S1403, If the algorithm result is good, then reduce the value of the amplitude t, reduce the gradient of threshold attenuation, and retest the screening results of readable text under this threshold;
[0028] S1404, The final selectable threshold oscillates within a certain interval range [t min , t max , and take the median of the interval as the final confidence threshold;
[0029] S142, In the method of screening readable text blocks according to the threshold when constructing the text output corresponding to the commodity picture, when constructing the text output corresponding to the commodity picture, when the readable confidence of the currently traversed text block is less than the threshold, suppress the recognized text content from appearing in the finally constructed text output content.
[0030] The beneficial effects of the present invention are that the present invention provides a method for discriminating readable text of commodity pictures based on a pixel - level attention mechanism. Among them, the method obtains the text block coordinates and text content of commodity pictures through OCR technology; obtains the readable confidence of text blocks through a pixel - level attention mechanism; and conducts screening of readable text through an adaptive threshold and the readable confidence of each text block, improving the problem of single correction results caused by the limitation of the corpus when using a pre - trained language model to screen text blocks in the prior art, thereby improving the accuracy of screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] , Figure 1 is a flowchart of the method for discriminating readable text of commodity pictures based on a pixel - level attention mechanism provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0033] Embodiment 1
[0034] As Figure 1 shown, Embodiment 1 of the present invention provides a method for discriminating readable text of commodity pictures based on a pixel - level attention mechanism. The method includes: avoiding the phenomenon of semantic incoherence caused by mixing unreadable text during the process of converting commodity pictures into readable text sequences using existing OCR technology.
[0035] Specifically, the method includes:
[0036] S110: Obtain the coordinates and corresponding text of the text blocks on the commodity picture through OCR technology;
[0037] Specifically, the position, i.e., the coordinate set, of the text blocks in the commodity picture is located through PSENet The text content of the text block sub - pictures is recognized by using the CRNN text recognition method
[0038] S120: Use the pixel - level attention mechanism to obtain the attention weights of all pixel units in the picture;
[0039] Specifically, for the set of pixel points Z of the commodity picture, the preliminary attention weight corresponding to the pixel point z i ∈Z is jointly determined by other pixel points in the whole picture. The calculation method of the preliminary attention weight matrix of the whole picture is A = ZWZ T After obtaining the preliminary attention weight matrix of the whole picture, the preliminary attention weight values of each pixel point are processed by using an activation function to obtain the attention weight values corresponding to each pixel point. The calculation method is A′ = sigmoid(A).
[0040] S130: Calculate the readable confidence of each text block by using the attention weights;
[0041] Specifically, the method for calculating the readable confidence of each text block by using the attention weights includes:
[0042] S1301. Obtain the center point coordinates T = {(x1, y1)…(x n , y n )} corresponding to each text block;
[0043] S1302. For the text block k, with the center point as the origin, apply the bivariate normal distribution formula to obtain the bivariate normal distribution probability values P k = {p1…p n} corresponding to each pixel point inside the text block;
[0044] S1303. For each pixel point inside the text block, perform a weighted average operation on the probability with the weight obtained by the attention mechanism, that is The calculation result is used as the readable confidence of each text block.
[0045] S140: Set a confidence threshold, and screen out the readable text blocks according to the threshold when constructing the text output corresponding to the commodity picture.
[0046] Specifically, it includes:
[0047] S141: The method for setting the confidence threshold includes:
[0048] S1401, set the initial confidence threshold t0 to 0.5;
[0049] S1402, during the training process, fix other parameters of the algorithm, attempt to reduce the threshold with the initial amplitude t0, and test the result of screening readable text using this threshold;
[0050] S1403, if the algorithm result is good, reduce the value of the amplitude t, reduce the gradient of threshold attenuation, and retest the result of screening readable text under this threshold;
[0051] S1404, the final selectable threshold oscillates within a certain range [t min , t max , and take the median of the range as the final confidence threshold.
[0052] S142: When constructing the text output corresponding to the product picture, screen out readable text blocks according to the threshold;
[0053] Specifically, when constructing the text output corresponding to the product picture, when the readable confidence of the currently traversed text block is less than the threshold, suppress the recognized text content from appearing in the finally constructed text output content.
[0054] In summary, the method for discriminating readable text in product pictures based on the pixel - level attention mechanism provided by the present invention, wherein the method obtains the text block coordinates and text content of the product picture through OCR technology; obtains the readable confidence of the text block through the pixel - level attention mechanism; and screens readable text through the adaptive threshold and the readable confidence of each text block, improving the problem of single correction result caused by the limitation of the corpus when using the pre - trained language model to screen text blocks in the prior art, thereby improving the accuracy of screening.
[0055] Inspired by the ideal embodiments of the present invention described above, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. A method for identifying readable text in product images based on a pixel - level attention mechanism, characterized in that, It includes the following steps: S110. Obtain the coordinates and corresponding texts of text blocks on the commodity picture through OCR technology; S120. Use the pixel-level attention mechanism to obtain the attention weights of all pixel units in the picture; Specifically including: S1201. For the set of pixels Z of the product image and the pixel z i ∈Z, the initial attention weight corresponding thereto is jointly determined by the other pixels of the entire image. The calculation method of the initial attention weight matrix of the entire image is A = ZWZ T ; S1202. After obtaining the preliminary attention weight matrix of the whole picture, process the preliminary attention weight values of each pixel point with an activation function to obtain the attention weight values corresponding to each pixel point. The calculation method is A' = sigmoid(A); S130. Calculate the readable confidence of each text block using the attention weights; Specifically including: S1301, obtain the center point coordinates T = {(x1, y1) … (x n , y n )} corresponding to each text block; S1302. For the text block k, with the center point as the origin, apply the bivariate normal distribution formula to obtain the bivariate normal distribution probability value P corresponding to each pixel point inside the text block k ={p1…p n}; S1303, for each pixel point within the text block, perform a weighted average operation on the probability using the weights obtained by the attention mechanism, that is The calculation result is used as the readable confidence of each text block; S140. Set a confidence threshold, and filter out readable text blocks according to the threshold when constructing the text output corresponding to the commodity picture.
2. The method for identifying readable text in commodity pictures based on a pixel-level attention mechanism according to claim 1, wherein Step S110 includes: S1101. Preprocess the commodity picture by using Gaussian filtering for noise reduction and mean smoothing; S1102, Use the PSENet model to perform text detection on the product image, and obtain the coordinate set of n text blocks in the product image S1103, according to the text block coordinate set Cut out the sub-graph set corresponding to each text block S1104, use the CRNN model to perform text recognition on the cropped text block sub-images, and obtain the text content corresponding to the n text blocks in the product image 3. The method for identifying readable text in product images based on a pixel-level attention mechanism according to claim 1, wherein, Step S140 includes: S141: Set a confidence threshold, specifically including: S1401. Set the initial confidence threshold t0 to 0.5; S1402. During the training process, fix other parameters of the algorithm, attempt to reduce the threshold with the initial amplitude t0, and test the readable text screening results using this threshold; S1403. If the algorithm result is good, reduce the value of the amplitude t, reduce the gradient of the threshold attenuation, and retest the readable text screening results under this threshold; S1404, the final optional threshold oscillates within a certain range [t min , t max , and the median value of the interval is taken as the final confidence threshold; S142: Filter out readable text blocks according to the threshold when constructing the text output corresponding to the commodity picture; When constructing the text output corresponding to the commodity picture, when the readable confidence of the currently traversed text block is less than the threshold, suppress the recognized text content from appearing in the finally constructed text output content.
Citation Information
Patent Citations
License plate real-time detection method based on edge guiding sparse attention mechanism
CN111444913A
Text recognition method and device based on improved AttentionOCR
CN112580738A