A method and device for recognizing non-significant text with posterior probability of weighted semantic relevance
The Hough transform corrects image tilt, uses wavelet denoising and enhancement methods to process non-striking text, and combines semantic correlation and multiple predictions of preset weights, the accuracy of non-striking text recognition is solved and the application effect of text images is improved.
Patent Information
- Application Number
- CN202211227212.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-09
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-10-09
AI Technical Summary
The prior art is difficult to accurately identify non-significant texts with incline, significant noise, partial absence or influenced by light, which reduces the practical application value of text images.
Image tilt is corrected by Hough transform, and non-striking text is processed using wavelet denoising and wavelet multi-scale image enhancement methods. Combining semantic correlation and multiple predictions of preset weights, accurate recognition of non-striking text is achieved.
It improves the accuracy of recognition of non-significant texts caused by tilt, significant noise, partial loss or lighting effects, and guarantees the practical application value of text images.
Smart Images

Figure CN115620296B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of character recognition, and in particular, to a method and device for recognizing non-significant characters with weighted semantic relevance posteriori. Background Art
[0002] In the information age, a large number of documents such as contracts and agreements are often stored and transmitted in the form of pictures, and are widely used in many fields such as finance, commerce, and law, playing an important role. We can very conveniently read the text content in the image, but cannot conveniently edit the image text. How to recognize the target text from the image has become a very meaningful task, and traditional methods can already accurately recognize the text in the picture.
[0003] However, in some cases, some characters in the picture cannot be ideally displayed. Slanting, significant noise, partial loss, and light influence will make the characters become non-significant characters. For non-significant characters, traditional recognition methods often cannot accurately recognize them, reducing the practical application value of the text image. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and device for recognizing non-significant characters with weighted semantic relevance posteriori, so as to improve the problem in the prior art that non-significant characters in a text image cannot be accurately recognized, thereby reducing the practical application value of the text image.
[0005] The embodiments of the present invention are implemented as follows:
[0006] First aspect, an embodiment of the present application provides a method for recognizing non-significant text with weighted semantic relevance posteriori, which includes the following steps: Step S110: Obtain the text image to be recognized, and perform skew correction on the text image to be recognized by using the Hough transform. Step S120: Use a preselected text recognition method to perform text recognition on the text image to be recognized, where the unrecognizable part in the text image to be recognized is non-significant text. Step S130: Use the wavelet denoising method to remove noise from the non-significant text, and then use the wavelet multi-scale image enhancement method to enhance the non-significant text. Step S140: Use the preselected text recognition method to perform text recognition on the enhanced non-significant text. If the recognition is successful, the non-significant text recognition result is obtained. Step S150: If the recognition is not successful, intercept the regions of a preset number of texts adjacent to the non-significant text, and use the semantic relevance between the non-significant text and the texts to predict the non-significant text to obtain a prediction result. Step S160: Repeat Step S150 to obtain multiple prediction results, where each time Step S150 is executed, the preset number is adjusted. Step S170: Weight the multiple prediction results respectively according to the preset weights to obtain the final recognition result of the non-significant text.
[0007] In some embodiments of the present invention, the above Step S160 includes the following steps: Intercept the regions of 4 texts adjacent to the non-significant text, and use the semantic relevance between the non-significant text and the 4 texts to predict the non-significant text to obtain a first prediction result. Intercept the regions of 6 texts adjacent to the non-significant text, and use the semantic relevance between the non-significant text and the 6 texts to predict the non-significant text to obtain a second prediction result. Intercept the regions of 8 texts adjacent to the non-significant text, and use the semantic relevance between the non-significant text and the 8 texts to predict the non-significant text to obtain a third prediction result. Intercept the regions of 10 texts adjacent to the non-significant text, and use the semantic relevance between the non-significant text and the 10 texts to predict the non-significant text to obtain a fourth prediction result. Intercept the regions of 12 texts adjacent to the non-significant text, and use the semantic relevance between the non-significant text and the 12 texts to predict the non-significant text to obtain a fifth prediction result.
[0008] In some embodiments of the present invention, the above Step S170 includes: The weights of the first prediction result, the second prediction result, the third prediction result, the fourth prediction result, and the fifth prediction result decrease in turn.
[0009] In some embodiments of the present invention, before the above step S150, the method further includes: collecting a plurality of text samples. Querying the co-occurrence times of different characters in all text samples, and calculating the co-occurrence probability of different characters according to the co-occurrence times. Predicting non-significant characters according to the co-occurrence probability of different characters.
[0010] In some embodiments of the present invention, the above steps of querying the co-occurrence times of different characters in all text samples and calculating the co-occurrence probability of different characters according to the co-occurrence times include: counting the first co-occurrence times of different characters within a first preset quantity, and calculating the first co-occurrence probability according to the first co-occurrence times. Counting the second co-occurrence times of different characters within a second preset quantity, and calculating the second co-occurrence probability according to the second co-occurrence times. Predicting non-significant characters according to the first co-occurrence probability and the second co-occurrence probability.
[0011] In some embodiments of the present invention, the above step S170 includes the following steps: determining a preset weight according to the category of the text image to be recognized.
[0012] In some embodiments of the present invention, the above preselected text recognition method includes at least Single-shot textdetector.
[0013] In a second aspect, an apparatus for recognizing non-significant characters with weighted semantic relevance posterior provided by an embodiment of the present application includes: a skew correction module, configured to obtain a text image to be recognized and perform skew correction on the text image to be recognized by using the Hough transform. A text image recognition module to be recognized, configured to perform text recognition on the text image to be recognized by using a preselected text recognition method, where the unrecognizable part in the text image to be recognized is a non-significant character. A non-significant character enhancement module, configured to remove noise from the non-significant character by using a wavelet denoising method, and then enhance the non-significant character by using a wavelet multi-scale image enhancement method. A non-significant character recognition module, configured to perform text recognition on the enhanced non-significant character by using a preselected text recognition method, and if successful recognition is achieved, obtain a non-significant character recognition result. A non-significant character prediction module, configured to, if successful recognition cannot be achieved, intercept a region of a preset number of characters adjacent to the non-significant character, and predict the non-significant character by using the semantic relevance between the non-significant character and the character to obtain a prediction result. A multiple prediction module, configured to repeatedly execute the non-significant character prediction module to obtain multiple prediction results, where each time the non-significant character prediction module is executed, the preset number is adjusted. A final non-significant character recognition module, configured to weight the multiple prediction results respectively according to a preset weight to obtain a final non-significant character recognition result.
[0014] In some embodiments of the present invention, the above-mentioned multiple prediction module includes: a first prediction unit, configured to intercept a region of 4 characters adjacent to the non-significant character, and use the semantic relevance between the non-significant character and the 4 characters to predict the non-significant character to obtain a first prediction result. A second prediction unit, configured to intercept a region of 6 characters adjacent to the non-significant character, and use the semantic relevance between the non-significant character and the 6 characters to predict the non-significant character to obtain a second prediction result. A third prediction unit, configured to intercept a region of 8 characters adjacent to the non-significant character, and use the semantic relevance between the non-significant character and the 8 characters to predict the non-significant character to obtain a third prediction result. A fourth prediction unit, configured to intercept a region of 10 characters adjacent to the non-significant character, and use the semantic relevance between the non-significant character and the 10 characters to predict the non-significant character to obtain a fourth prediction result. A fifth prediction unit, configured to intercept a region of 12 characters adjacent to the non-significant character, and use the semantic relevance between the non-significant character and the 12 characters to predict the non-significant character to obtain a fifth prediction result.
[0015] In some embodiments of the present invention, the above-mentioned non-significant character final recognition module includes: a weight assignment unit, configured to make the weights of the first prediction result, the second prediction result, the third prediction result, the fourth prediction result, and the fifth prediction result decrease in sequence.
[0016] In some embodiments of the present invention, the above-mentioned non-significant character recognition device based on weighted semantic relevance further includes: a character sample collection module, configured to collect multiple character samples. A joint occurrence probability calculation module, configured to query the joint occurrence times of different characters in all character samples, and calculate the joint occurrence probability of different characters according to the joint occurrence times. A joint occurrence probability prediction module, configured to predict the non-significant character according to the joint occurrence probability of different characters.
[0017] In some embodiments of the present invention, the above-mentioned joint occurrence probability calculation module includes: a first joint probability calculation unit, configured to count the first joint times of different characters within a first preset quantity, and calculate a first joint probability according to the first joint times. A second joint probability calculation unit, configured to count the second joint times of different characters within a second preset quantity, and calculate a second joint probability according to the second joint times. A comprehensive prediction unit, configured to predict the non-significant character according to the first joint probability and the second joint probability.
[0018] In some embodiments of the present invention, the above-mentioned non-significant character final recognition module includes: a preset weight determination unit, configured to determine a preset weight according to the category of the text image to be recognized.
[0019] In some embodiments of the present invention, the above preselected text recognition method at least includes Single-shot textdetector.
[0020] In a third aspect, an embodiment of the present application provides an electronic device, which includes a memory for storing one or more programs; and a processor. When the one or more programs are executed by the processor, the method according to any one of the above first aspects is implemented.
[0021] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method according to any one of the above first aspects is implemented.
[0022] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects:
[0023] The present invention provides a method and apparatus for recognizing non-significant text with weighted semantic relevance posterior, which includes the following steps: Step S110: Obtain a text image to be recognized, and perform skew correction on the text image to be recognized by using the Hough transform. Step S120: Use the preselected text recognition method to perform text recognition on the text image to be recognized, where the unrecognizable part in the text image to be recognized is non-significant text. Step S130: Use the wavelet denoising method to remove noise from the non-significant text, and then use the wavelet multi-scale image enhancement method to enhance the non-significant text. Step S140: Use the preselected text recognition method to perform text recognition on the enhanced non-significant text. If the recognition is successful, the non-significant text recognition result is obtained. Step S150: If the recognition is not successful, intercept the regions of a preset number of texts adjacent to the non-significant text, and use the semantic relevance between the non-significant text and the text to predict the non-significant text to obtain a prediction result. Step S160: Repeat Step S150 to obtain multiple prediction results, where each time Step S150 is executed, the preset number is adjusted. Step S170: Weight the multiple prediction results respectively according to a preset weight to obtain the final recognition result of the non-significant text.
[0024] The method and device identify the line features in the text image to be recognized by using the Hough transform, so as to perform skew correction on the text image to be recognized, thereby obtaining the corrected text image to be recognized, avoiding the text in the text image to be recognized from becoming non-significant text due to the skew of the text image to be recognized, and making the text image to be recognized more accurate. The text in the text image to be recognized can be recognized by a preselected text recognition method to distinguish the non-significant text in the text image to be recognized. The noise is removed by wavelet denoising without affecting the important information of the non-significant text, making the non-significant text clearer. The non-significant text is enhanced by the wavelet multi-scale image enhancement method. While improving the image contrast, the problem of noise enhancement is effectively solved, making the enhanced non-significant text image convenient for image analysis and post-processing. If the non-significant text is recognized by the preselected text recognition method after enhancement and the non-significant text recognition result is successfully obtained, the purpose of recognizing the non-significant text is achieved. If the non-significant text cannot be successfully recognized, a region of a preset number of texts around the non-significant text is intercepted, and prediction is performed using semantic relevance to obtain a prediction result, so as to reflect a primary prediction judgment on the non-significant text. And the above preset number is adjusted, and step S150 is repeatedly executed to obtain multiple prediction results, so as to reflect multiple prediction judgments on the non-significant text. And according to the preset weights, the multiple prediction results are weighted respectively, and the final recognition result of the non-significant text can be obtained. Thus, the purpose of accurately recognizing the non-significant text caused by skew, significant noise, partial loss, and illumination influence is realized, and the practical application value of the text image to be recognized is guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a flowchart of a method for recognizing non-significant text with weight-based semantic relevance posteriori provided by an embodiment of the present invention;
[0027] Figure 2 It is a structural block diagram of a device for recognizing non-significant text with weight-based semantic relevance posteriori provided by an embodiment of the present invention;
[0028] Figure 3 It is a schematic structural block diagram of an electronic device provided by an embodiment of the present invention.
[0029] Icons: 100 - Non - significant text recognition device for weighted semantic relevance posterior; 110 - Skew correction module; 120 - Text image recognition module to be recognized; 130 - Non - significant text enhancement module; 140 - Non - significant text recognition module; 150 - Non - significant text prediction module; 160 - Multiple prediction module; 170 - Final non - significant text recognition module; 101 - Memory; 102 - Processor; 103 - Communication interface. Detailed implementation manners
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0031] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0032] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, if terms such as "first", "second", etc. are used only for distinguishing descriptions, they cannot be understood as indicating or implying relative importance.
[0033] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or sequence between these entities or operations. Moreover, if terms such as "include", "comprise" or any other variant thereof are intended to cover non - exclusive inclusion, a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article or device. Without further limitation, if an element is defined by the statement "including one...", it does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0034] In the description of this application, it should be noted that if terms such as "upper", "lower", "inner", "outer", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the products of this application are usually placed during use. This is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or component referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to this application.
[0035] In the description of this application, it should also be noted that unless otherwise clearly specified and limited, if terms such as "set" and "connect" are used, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific situations.
[0036] The following will describe in detail some embodiments of this application with reference to the drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0037] Embodiment
[0038] Please refer to Figure 1 , Figure 1 as shown in the flowchart of a non-significant text recognition method with posterior weight-based semantic relevance provided by an embodiment of this application. An embodiment of this application provides a non-significant text recognition method with posterior weight-based semantic relevance, which includes the following steps:
[0039] Step S110: Obtain the text image to be recognized, and use the Hough transform to perform skew correction on the text image to be recognized;
[0040] Specifically, the text image to be recognized can be obtained through image acquisition devices such as cameras, scanners, and cameras. The Hough transform can identify the line features in the text image to be recognized to perform skew correction on the text image to be recognized, so as to obtain the corrected text image to be recognized, and avoid the text in the text image to be recognized becoming non-significant due to the skew of the text image to be recognized, making the text image to be recognized more accurate.
[0041] Step S120: Use a preselected text recognition method to perform text recognition on the text image to be recognized, where the unrecognizable part in the text image to be recognized is non-significant text;
[0042] Specifically, the above preselected text recognition method may include text recognition methods such as Single-shot text detector. For non-significant texts in the text image to be recognized, such as those with significant noise, partial loss, or affected by light, if the above non-significant texts are recognized through the preselected text recognition method, the accuracy of the recognition result cannot be guaranteed, and even the preselected text recognition method cannot recognize the non-significant texts. Then, the non-significant texts in the text image to be recognized can be distinguished through the preselected text recognition method.
[0043] Step S130: Use the wavelet denoising method to remove noise from the non-significant text, and then use the wavelet multi-scale image enhancement method to enhance the non-significant text;
[0044] Specifically, the non-significant text is deeply optimized through wavelet denoising and wavelet multi-scale image enhancement. Among them, in the above wavelet denoising, after the non-significant text is wavelet-transformed, the generated wavelet coefficients contain important information of the non-significant text. After the non-significant text is wavelet-decomposed, the wavelet coefficients of the non-significant text are larger, and the wavelet coefficients of the noise are smaller. By selecting a suitable threshold, those wavelet coefficients smaller than the threshold are considered to be generated by noise, and then the noise is removed to achieve the purpose of denoising. After denoising the non-significant text through the wavelet denoising method, the important information of the non-significant text can be retained. Thus, on the basis of not affecting the important information of the non-significant text, the noise is removed, making the non-significant text clearer.
[0045] In addition, by enhancing the non-significant text through the above wavelet multi-scale image enhancement method, while improving the image contrast, the problem of noise being enhanced is effectively solved, making the enhanced image of the non-significant text convenient for image analysis and post-processing.
[0046] Step S140: Use the preselected text recognition method to recognize the enhanced non-significant text. If the recognition is successful, the non-significant text recognition result is obtained;
[0047] Step S150: If the recognition is not successful, intercept the areas of a preset number of texts adjacent to the non-significant text, and use the semantic relevance between the non-significant text and the text to predict the non-significant text to obtain a prediction result;
[0048] Specifically, the above situation where the recognition is not successful may include that the non-significant text still cannot be recognized through the preselected text recognition method or the confidence level of the non-significant text recognition result is lower than the preset value. When the non-significant text cannot be successfully recognized, the areas of a preset number of texts around the non-significant text are intercepted and predicted using semantic relevance to obtain a prediction result to reflect a primary prediction and judgment of the non-significant text.
[0049] Step S160: Repeat step S150 to obtain multiple prediction results, wherein the preset number is adjusted each time step S150 is performed;
[0050] Specifically, the user can select the preset number of adjustments and the specific value of the preset number based on the actual category of the text image to be recognized. The preset number is adjusted and step S150 is repeated to obtain multiple prediction results to reflect multiple predictions of non-salient text.
[0051] Step S170: weighting the plurality of prediction results respectively according to preset weights to obtain a final recognition result of the non-salient text.
[0052] Specifically, users can set preset weights for each prediction result based on the actual category of the text image to be recognized. Since multiple prediction results reflect multiple predictions for non-significant text, these prediction results are weighted according to the preset weights to obtain the final recognition result for the non-significant text. This achieves relatively accurate recognition of non-significant text caused by tilt, significant noise, partial omissions, and lighting effects, ensuring the practical application value of the text image to be recognized.
[0053] In some implementations of this embodiment, step S160 includes the following steps: intercepting an area of four characters adjacent to the non-significant character, and predicting the non-significant character using the semantic association between the non-significant character and the four characters to obtain a first prediction result; intercepting an area of six characters adjacent to the non-significant character, and predicting the non-significant character using the semantic association between the non-significant character and the six characters to obtain a second prediction result; intercepting an area of eight characters adjacent to the non-significant character, and predicting the non-significant character using the semantic association between the non-significant character and the eight characters to obtain a third prediction result; intercepting an area of ten characters adjacent to the non-significant character, and predicting the non-significant character using the semantic association between the non-significant character and the ten characters to obtain a fourth prediction result; intercepting an area of twelve characters adjacent to the non-significant character, and predicting the non-significant character using the semantic association between the non-significant character and the twelve characters to obtain a fifth prediction result. Specifically, the preset number is set to 4, 6, 8, 10, and 12, and step S150 is repeated in sequence to obtain the first prediction result, the second prediction result, the third prediction result, the fourth prediction result, and the fifth prediction result, thereby achieving the purpose of multiple prediction judgments on non-significant text.
[0054] In some implementations of this embodiment, step S170 includes: weighting the first prediction result, the second prediction result, the third prediction result, the fourth prediction result, and the fifth prediction result in descending order. Specifically, the smaller the preset number, the more accurate the corresponding prediction result obtained based on the semantic association between the non-significant text and the preset number of texts. As the preset numbers corresponding to the first through fifth prediction results decrease in sequence, the weights corresponding to the first through fifth prediction results also decrease in sequence, further ensuring the accuracy of the final recognition result of the non-significant text.
[0055] In some implementations of this embodiment, before step S150, the method further includes collecting multiple text samples to ensure the accuracy of the joint occurrence probabilities of different characters obtained based on the text samples. The number of joint occurrences of different characters in all text samples is queried, and the joint occurrence probabilities of different characters are calculated based on the joint occurrence numbers. Non-significant characters are predicted based on the joint occurrence probabilities of different characters. Specifically, the non-significant characters are predicted based on the combined joint occurrence probabilities of different characters in all text samples, minimizing the possibility that the characters are not included in the text samples, thereby ensuring the accuracy of the prediction of non-significant characters.
[0056] In some implementations of this embodiment, the step of querying the joint occurrence counts of different characters in all text samples and calculating the joint occurrence probabilities of the different characters based on the joint occurrence counts includes: counting a first joint occurrence count of different characters within a first preset number and calculating a first joint probability based on the first joint occurrence count; counting a second joint occurrence count of different characters within a second preset number and calculating a second joint probability based on the second joint occurrence count; and predicting non-significant characters based on the first joint probability and the second joint probability. Specifically, the second preset number is greater than the first preset number. Since the fewer the number of different characters in all text samples, the more accurate the probability of occurrence of each character obtained based on the joint occurrence probabilities of different characters, the joint occurrence probabilities of characters in all text samples where the number of different characters does not exceed the first preset number are primarily counted and used as a judgment basis to ensure the accuracy of prediction for non-significant character recognition. Furthermore, the joint occurrence probabilities of characters in all text samples where the number of different characters does not exceed the second preset number are counted and used as a judgment reference. Combining the judgment basis and the judgment reference allows for more accurate prediction of non-significant characters.
[0057] In some implementations of this embodiment, step S170 includes the following step: determining preset weights based on the category of the text image to be recognized. Specifically, the user may set the preset weights for each prediction result based on the actual category of the text image to be recognized, thereby ensuring that the preset weights are appropriate for the category of the text image to be recognized.
[0058] In some embodiments of this embodiment, the above preselected text recognition method at least includes Single-shot text detector. Specifically, the user can select any one or more preselected text recognition methods in the Single-shot text detector according to the actual category of the text image to be recognized, so as to be closer to the actual category of the text image to be recognized.
[0059] Please refer to Figure 2 , Figure 2 The following figure shows a structural block diagram of a non-significant text recognition device 100 with weighted semantic relevance posterior provided by an embodiment of the present application. A non-significant text recognition device 100 with weighted semantic relevance posterior includes: a skew correction module 110, configured to obtain a text image to be recognized and perform skew correction on the text image to be recognized by using the Hough transform. A text image recognition module 120 to be recognized, configured to perform text recognition on the text image to be recognized by using a preselected text recognition method, where the unrecognizable part in the text image to be recognized is non-significant text. A non-significant text enhancement module 130, configured to remove noise from the non-significant text by using a wavelet denoising method, and then enhance the non-significant text by using a wavelet multi-scale image enhancement method. A non-significant text recognition module 140, configured to perform text recognition on the enhanced non-significant text by using a preselected text recognition method. If the recognition is successful, a non-significant text recognition result is obtained. A non-significant text prediction module 150, configured to, if the recognition is not successful, intercept a region of a preset number of texts adjacent to the non-significant text, and predict the non-significant text by using the semantic relevance between the non-significant text and the text to obtain a prediction result. A multiple prediction module 160, configured to repeatedly execute the non-significant text prediction module 150 to obtain multiple prediction results, where each time the non-significant text prediction module 150 is executed, the preset number is adjusted. A non-significant text final recognition module 170, configured to weight the multiple prediction results respectively according to a preset weight to obtain a non-significant text final recognition result.
[0060] Specifically, the device uses the Hough transform to identify the line features in the text image to be recognized, so as to perform skew correction on the text image to be recognized, thereby obtaining the corrected text image to be recognized, avoiding the text in the text image to be recognized from becoming non-significant text due to the skew of the text image to be recognized, and making the text image to be recognized more accurate. The text in the text image to be recognized can be recognized by a preselected text recognition method to distinguish the non-significant text in the text image to be recognized. Wavelet denoising is used to remove the noise on the basis of not affecting the important information of the non-significant text, making the non-significant text clearer. The non-significant text is enhanced by the wavelet multi-scale image enhancement method. While improving the image contrast, the problem of enhanced noise is effectively solved, making the enhanced non-significant text image convenient for image analysis and post-processing. If the non-significant text is successfully recognized by the preselected text recognition method after enhancement, the purpose of recognizing the non-significant text is achieved. If the non-significant text cannot be successfully recognized, a region of a preset number of characters around the non-significant text is intercepted and predicted using semantic relevance to obtain a prediction result, so as to reflect a prediction judgment on the non-significant text. The above preset number is adjusted, and the non-significant text prediction module 150 is repeatedly executed to obtain multiple prediction results, so as to reflect multiple prediction judgments on the non-significant text. And according to the preset weights, the multiple prediction results are weighted respectively to obtain the final recognition result of the non-significant text. Thus, the purpose of accurately recognizing non-significant text caused by skew, significant noise, partial loss, and illumination influence is realized, ensuring the practical application value of the text image to be recognized.
[0061] In some implementations of this embodiment, the multiple prediction module 160 includes: a first prediction unit configured to extract a region of four characters adjacent to a non-significant character and predict the non-significant character using the semantic association between the non-significant character and the four characters to obtain a first prediction result; a second prediction unit configured to extract a region of six characters adjacent to the non-significant character and predict the non-significant character using the semantic association between the non-significant character and the six characters to obtain a second prediction result; a third prediction unit configured to extract a region of eight characters adjacent to the non-significant character and predict the non-significant character using the semantic association between the non-significant character and the eight characters to obtain a third prediction result; a fourth prediction unit configured to extract a region of ten characters adjacent to the non-significant character and predict the non-significant character using the semantic association between the non-significant character and the ten characters to obtain a fourth prediction result; and a fifth prediction unit configured to extract a region of twelve characters adjacent to the non-significant character and predict the non-significant character using the semantic association between the non-significant character and the twelve characters to obtain a fifth prediction result. Specifically, the preset number is set to 4, 6, 8, 10, and 12, and step S150 is repeated in sequence to obtain the first prediction result, the second prediction result, the third prediction result, the fourth prediction result, and the fifth prediction result, thereby achieving the purpose of multiple prediction judgments on non-significant text.
[0062] In some implementations of this embodiment, the non-significant text final recognition module 170 includes a weight assignment unit configured to assign weights to the first prediction result, the second prediction result, the third prediction result, the fourth prediction result, and the fifth prediction result in descending order. Specifically, the smaller the preset number, the more accurate the corresponding prediction result obtained based on the semantic association between the non-significant text and the preset number of characters. As the preset numbers corresponding to the first through fifth prediction results decrease in order, the weights corresponding to the first through fifth prediction results also decrease in order, further ensuring the accuracy of the final recognition result of the non-significant text.
[0063] In some implementations of this embodiment, the weighted semantic relevance posteriori non-significant text recognition device 100 further includes: a text sample collection module for collecting multiple text samples; a joint occurrence probability calculation module for querying the number of joint occurrences of different texts in all text samples and calculating the joint occurrence probability of different texts based on the joint occurrence number; and a joint occurrence probability prediction module for predicting non-significant texts based on the joint occurrence probabilities of different texts. Specifically, the non-significant text is predicted based on the joint occurrence probabilities of different texts in all text samples, minimizing the possibility that the text samples do not contain the text, thereby ensuring the accuracy of the non-significant text prediction.
[0064] In some embodiments of the present embodiment, the above-mentioned joint occurrence probability calculation module includes: a first joint probability calculation unit, configured to count the first joint times of different characters within a first preset quantity, and calculate a first joint probability according to the first joint times. A second joint probability calculation unit, configured to count the second joint times of different characters within a second preset quantity, and calculate a second joint probability according to the second joint times. A comprehensive prediction unit, configured to predict non-significant characters according to the first joint probability and the second joint probability. Specifically, the above-mentioned second preset quantity is greater than the first preset quantity. Since the fewer the number of different characters in all text samples, the more accurate the probability of each character obtained according to the joint occurrence probability of different characters. Then, focus on counting the joint occurrence probability of text samples with the number of different characters not exceeding the first preset quantity, and use it as a judgment basis to ensure the prediction accuracy of non-significant character recognition. And count the joint occurrence probability of text samples with the number of different characters not exceeding the second preset quantity, and use it as a judgment reference. Combining the judgment basis and the judgment reference can more accurately predict non-significant characters.
[0065] In some embodiments of the present embodiment, the above-mentioned non-significant character final recognition module 170 includes: a preset weight determination unit, configured to determine a preset weight according to the category of the text image to be recognized. Specifically, the user can set the preset weights of each prediction result according to the actual category of the text image to be recognized, so as to ensure that the preset weights are adapted to the category of the text image to be recognized.
[0066] In some embodiments of the present embodiment, the above-mentioned preselected character recognition method includes at least Single-shot text detector. Specifically, the user can select any one or more preselected character recognition methods in Single-shot text detector according to the actual category of the text image to be recognized, so as to be closer to the actual category of the text image to be recognized.
[0067] Please refer to Figure 3 , Figure 3A schematic structural block diagram of an electronic device provided by an embodiment of the present application. The electronic device includes a memory 101, a processor 102, and a communication interface 103. The memory 101, the processor 102, and the communication interface 103 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory 101 can be used to store software programs and modules, such as program instructions / modules corresponding to a non-significant text recognition device 100 with weighted semantic relevance posteriori provided by an embodiment of the present application. The processor 102 executes various functional applications and data processing by executing the software programs and modules stored in the memory 101. The communication interface 103 can be used for signaling or data communication with other node devices.
[0068] Among them, the memory 101 can be, but is not limited to, a random access memory 101 (Random Access Memory, RAM), a read-only memory 101 (Read Only Memory, ROM), a programmable read-only memory 101 (Programmable Read-Only Memory, PROM), an erasable programmable read-only memory 101 (Erasable Programmable Read-Only Memory, EPROM), an electrically erasable programmable read-only memory 101 (Electric Erasable Programmable Read-Only Memory, EEPROM), etc.
[0069] The processor 102 can be an integrated circuit chip with signal processing capabilities. The processor 102 can be a general-purpose processor 102, including a central processing unit 102 (Central Processing Unit, CPU), a network processor 102 (Network Processor, NP), etc.; it can also be a digital signal processor 102 (Digital Signal Processing, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0070] It can be understood that Figure 3 The structure shown is only for illustration, and the electronic device may further include more or fewer components than those shown in Figure 3 or have a different configuration from that shown in Figure 3 shown.Figure 3 Each component shown can be implemented using hardware, software, or a combination thereof.
[0071] In the embodiments provided in the present application, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the apparatus, method, and computer program product according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0072] In addition, the various functional modules in the embodiments of the present application can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0073] If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory 101 (ROM, Read-Only Memory), random access memory 101 (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0074] In summary, a method and apparatus for recognizing non-significant text with weight-based semantic relevance posteriori provided by an embodiment of the present application include the following steps: Step S110: Obtain a text image to be recognized, and perform skew correction on the text image to be recognized using the Hough transform. Step S120: Perform text recognition on the text image to be recognized using a preselected text recognition method, where the unrecognizable part in the text image to be recognized is non-significant text. Step S130: Use the wavelet denoising method to remove noise from the non-significant text, and then use the wavelet multi-scale image enhancement method to enhance the non-significant text. Step S140: Perform text recognition on the enhanced non-significant text using the preselected text recognition method. If the recognition is successful, the non-significant text recognition result is obtained. Step S150: If the recognition is not successful, intercept the regions of a preset number of texts adjacent to the non-significant text, and use the semantic relevance between the non-significant text and the text to predict the non-significant text to obtain a prediction result. Step S160: Repeat Step S150 to obtain multiple prediction results, where each time Step S150 is executed, the preset number is adjusted. Step S170: Weight each of the multiple prediction results according to a preset weight to obtain the final recognition result of the non-significant text. This method and apparatus use the Hough transform to identify the line features in the text image to be recognized, so as to perform skew correction on the text image to be recognized, thereby obtaining the corrected text image to be recognized, avoiding the text in the text image to be recognized becoming non-significant text due to the skew of the text image to be recognized, and making the text image to be recognized more accurate. Through the preselected text recognition method, text recognition can be performed on the text image to be recognized to distinguish the non-significant text in the text image to be recognized. By wavelet denoising, the noise is removed on the basis of not affecting the important information of the non-significant text, making the non-significant text clearer. The non-significant text is enhanced by the wavelet multi-scale image enhancement method. While improving the image contrast, the problem of noise being enhanced is effectively solved, making the image of the enhanced non-significant text convenient for image analysis and post-processing. If the enhanced non-significant text is successfully recognized by the preselected text recognition method to obtain the non-significant text recognition result, the purpose of recognizing the non-significant text is achieved. If the non-significant text cannot be successfully recognized, the regions of a preset number of texts around the non-significant text are intercepted, and the semantic relevance is used for prediction to obtain a prediction result to reflect a prediction judgment on the non-significant text. And the above preset number is adjusted, and Step S150 is repeatedly executed to obtain multiple prediction results to reflect multiple prediction judgments on the non-significant text. And according to the preset weight, each of the multiple prediction results is weighted to obtain the final recognition result of the non-significant text. Thus, the purpose of accurately recognizing non-significant text caused by skew, significant noise, partial loss, and illumination influence is achieved, ensuring the practical application value of the text image to be recognized.
[0075] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0076] For those skilled in the art, it is obvious that the present application is not limited to the details of the above-described exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A non-significant text recognition method with weighted semantic relevance posterior, characterized in that It includes the following steps: Step S110: Obtain the text image to be recognized, and perform skew correction on the text image to be recognized by using the Hough transform; Step S120: Perform text recognition on the text image to be recognized by using a preselected text recognition method, wherein the unrecognizable part in the text image to be recognized is non-significant text; Step S130: Remove the noise from the non-significant text by using the wavelet denoising method, and then enhance the non-significant text by using the wavelet multi-scale image enhancement method; Step S140: Perform text recognition on the enhanced non-significant text by using the preselected text recognition method. If the recognition is successful, the non-significant text recognition result is obtained; Step S150: If the recognition is not successful, intercept the region of a preset number of characters adjacent to the non-significant text, and predict the non-significant text by using the semantic relevance between the non-significant text and the characters to obtain a prediction result; Step S160: Repeat Step S150 to obtain multiple prediction results. When performing Step S150 each time, adjust the preset number; Step S170: Weight the multiple prediction results respectively according to a preset weight to obtain the final recognition result of the non-significant text.
2. The method for identifying non-significant texts with posterior weight-based semantic relevance according to claim 1, characterized in that, Step S160 includes the following steps: Intercept the region of 4 characters adjacent to the non-significant text, and predict the non-significant text by using the semantic relevance between the non-significant text and the 4 characters to obtain a first prediction result; Intercept the region of 6 characters adjacent to the non-significant text, and predict the non-significant text by using the semantic relevance between the non-significant text and the 6 characters to obtain a second prediction result; Intercept the region of 8 characters adjacent to the non-significant text, and predict the non-significant text by using the semantic relevance between the non-significant text and the 8 characters to obtain a third prediction result; Intercept the region of 10 characters adjacent to the non-significant text, and predict the non-significant text by using the semantic relevance between the non-significant text and the 10 characters to obtain a fourth prediction result; Intercept the region of 12 characters adjacent to the non-significant text, and predict the non-significant text by using the semantic relevance between the non-significant text and the 12 characters to obtain a fifth prediction result.
3. The non-significant text recognition method with posterior weight-based semantic relevance according to claim 2, characterized in that, Step S170 includes: The weights of the first prediction result, the second prediction result, the third prediction result, the fourth prediction result, and the fifth prediction result decrease in sequence.
4. The method for identifying non-significant text with posterior weight-based semantic relevance according to claim 1, wherein Before Step S150, it further includes: Collect multiple character samples; Query the joint occurrence times of different characters in all the character samples, and calculate the joint occurrence probability of different characters according to the joint occurrence times; Predict the non-significant text according to the joint occurrence probability of different characters.
5. The non-significant text recognition method with posteriori weighted semantic relevance according to claim 4, characterized in that The step of querying the joint occurrence times of different characters in all the character samples and calculating the joint occurrence probability of different characters according to the joint occurrence times includes: Count the first joint occurrences of different characters within the first preset quantity, and calculate the first joint probability according to the first joint occurrences; Count the second joint occurrences of different characters within the second preset quantity, and calculate the second joint probability according to the second joint occurrences; Predict the non-significant characters according to the first joint probability and the second joint probability.
6. The method for identifying non-significant texts with weighted semantic relevance posteriori according to claim 1, characterized in that Step S170 includes the following steps: Determine the preset weight according to the category of the text image to be recognized.
7. The non-significant text recognition method with posterior weight-based semantic relevance according to claim 1, characterized in that The preselected text recognition method includes at least Single-shot text detector.
8. A non-significant text recognition device with posteriori weight-based semantic relevance, characterized in that, It includes: An inclination correction module, configured to obtain a text image to be recognized, and perform inclination correction on the text image to be recognized by using the Hough transform; A text image to be recognized module, configured to perform text recognition on the text image to be recognized by using the preselected text recognition method, wherein the unrecognizable part in the text image to be recognized is non-significant characters; A non-significant character enhancement module, configured to remove noise from the non-significant characters by using the wavelet denoising method, and then enhance the non-significant characters by using the wavelet multi-scale image enhancement method; A non-significant character recognition module, configured to perform text recognition on the enhanced non-significant characters by using the preselected text recognition method. If the recognition is successful, a non-significant character recognition result is obtained; A non-significant character prediction module, configured to, if the recognition is not successful, intercept a region of a preset quantity of characters adjacent to the non-significant characters, and predict the non-significant characters by using the semantic relevance between the non-significant characters and the characters to obtain a prediction result; A multiple prediction module, configured to repeatedly execute the non-significant character prediction module to obtain multiple prediction results. Wherein, each time the non-significant character prediction module is executed, the preset quantity is adjusted; A non-significant character final recognition module, configured to weight the multiple prediction results respectively according to the preset weight to obtain a non-significant character final recognition result.
9. An electronic device, characterized in that, It includes: A memory, configured to store one or more programs; A processor; When the one or more programs are executed by the processor, the method according to any one of claims 1-7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Semantic association character recognition method and device
CN113221904A
Training method of text recognition model, text recognition method, and apparatus
KR1020220127189A