Text image processing method, processing device, processing system and readable storage medium
By dividing the text image set into multiple sets and generating labels, and using a trained model to correct the text image rotation angle, the problem of inaccurate text image rotation angle detection in existing technologies is solved, achieving efficient and accurate text image processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YONYOU NETWORK TECH CO LTD
- Filing Date
- 2022-10-24
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies for detecting the rotation angle of text images have low accuracy and are complex and cumbersome.
The text image set is divided into multiple sets, and labels are generated based on the rotation angle range and the rotation angle of the target field. The rotation angle in the text image is then corrected by training a model.
It improves the accuracy and speed of target field recognition in text images, simplifies the processing, enhances the generalization ability of the model, supports detection from any angle, and has small errors.
Smart Images

Figure CN115690808B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, and more specifically, to a text image processing method, a text image processing apparatus, a text image processing system, and a readable storage medium. Background Technology
[0002] In existing technologies, methods for correcting the rotation angle of text in text images have the following technical problems: the accuracy of detecting the rotation angle of text is not high, and the methods are complex and cumbersome. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems existing in the prior art or related art.
[0004] Therefore, a first aspect of the present invention is to provide a text image processing method.
[0005] A second aspect of the present invention is to provide a text image processing apparatus.
[0006] A third aspect of the present invention is to provide a text image processing system.
[0007] A fourth aspect of the present invention is to provide a readable storage medium.
[0008] In view of this, according to one aspect of the present invention, a text image processing method is proposed, comprising: acquiring a text image set, wherein multiple text images in the text image set include multiple target fields with rotation angles, and the rotation angles of the target fields in each text image are the same; determining multiple text image sets, wherein the rotation angle ranges corresponding to the multiple text image sets are different; dividing the multiple text images into multiple text image sets according to the rotation angles and rotation angle ranges; generating multiple labels according to the multiple text image sets; adding the labels to the multiple text images in the corresponding text image sets; training a preset model based on the multiple text images after adding multiple labels; and correcting the rotation angles of the target fields in the text images to be processed using the trained preset model.
[0009] First, a set of text images for training is obtained, which includes multiple types of text images. Specifically, the text images can be different types of ticket images, such as taxi receipts, train tickets, etc., with different text distributions and layout structures. Then, the obtained text images are designed and made into text images that can be used to train the preset model.
[0010] Furthermore, the text image includes multiple target fields with rotation angles. The target fields in the text image can be fields with feature information in the text image. Users can limit the information of the text image as needed, and thus limit the target fields in the text image. The rotation angle is the tilt angle between the arrangement direction of the target fields and the horizontal direction. In the same image, the rotation angle of the target fields is the same. Furthermore, any field in the same image, that is, the text field in the text image, also has the same rotation angle.
[0011] After acquiring multiple types of text images, multiple text image sets are first determined and then divided into multiple text image sets. Specifically, each text image set represents a category of text images. The rotation angle of the target field differs between text images of different categories. Since the target field of each text image has a rotation angle, in the downstream invoice OCR (Optical Character Recognition) task, when the rotation angle of the target field in the text image differs from that in the horizontal direction by less than 10 degrees, the impact on the final recognition effect of the recognition task is reduced, that is, the impact on the recognition accuracy is small. Therefore, in order to ensure the recognition accuracy and reduce the computational load, text images with similar rotation angles should be divided into one category. When recognizing the rotation angle of multiple text images in this category, a single rotation angle is output, thereby ensuring the recognition accuracy, reducing computation time, and improving the cost-effectiveness of this technical solution.
[0012] Subsequently, multiple labels are generated based on multiple text image sets.
[0013] In the technical solution of this application, the label is used to indicate the rotation angle characteristics of the text image in different text image categories. Since the rotation angle of the target field of text images belonging to different text image categories is different, by adding different labels to text images of different categories, the difference between the rotation angles of the target field in text images of different text image categories can be better indicated.
[0014] After adding the corresponding label to each of the multiple text images, use it as the training text image for the preset model. Train the preset model so that it can learn target fields with different rotation angles in text images of different categories as learning targets and correct the rotation angle of the target fields in the text images to be processed through the trained preset model.
[0015] In the technical solution of this application, correction means adjusting the tilt angle between the arrangement direction of the target field in the text image and the horizontal direction, so that the direction of the target field in the text image is consistent with the horizontal direction.
[0016] In the technical solution of this application, the accuracy of identifying target fields in text images is improved by dividing multiple text images into multiple text image sets.
[0017] By adding multiple different labels to multiple second text image sets and multiple first text image sets, different identifiers can be added to the rotation angles of different target fields, which facilitates the training of the preset model and improves the training speed and training effect. At the same time, the text image processing method of the present invention simplifies the text image processing process by classifying multiple text images, adding labels, and training the preset model with the labeled text images, thereby further improving the text image processing speed and reducing resource consumption.
[0018] Furthermore, when the type of text image changes or the category of the target field in the text image changes, multiple acquired text images can be updated to obtain the changed text image, thereby improving the generalization ability of the preset model, meeting the timeliness requirements of the preset model for text image processing, improving the processing effect of the text image processing method and the processing performance of the trained preset model. At the same time, the output is the corrected text image to be processed, which meets the accuracy requirements of the downstream invoice OCR task for text images, thus facilitating OCR task integration.
[0019] Furthermore, after dividing multiple text images into multiple text image sets, the scale of the multiple text images also needs to be adjusted, that is, the text images are adjusted to a standard size, so as to avoid the text images with inconsistent scales affecting the training of the preset model, and thus affecting the accuracy of the preset model in detecting the rotation angle of the target field in the text image to be processed.
[0020] The text image processing method of this technical solution can quickly output the rotation angle of the target field in the current text image to be detected. At the same time, it has a high accuracy rate and can support detection at any angle with small error.
[0021] Meanwhile, this text image processing method can update the images in the text image set at any time, realizing the automatic expansion and generation of the text image set. Thus, this text image processing method can meet the needs of various types of text image detection. The trained model has strong generalization ability and a wide range of applications.
[0022] The text image processing method of the present invention may also have the following technical features:
[0023] In the above technical solution, multiple text images are divided into multiple text image sets according to the rotation angle and the rotation angle range. Specifically, this includes: determining multiple rotation angles of multiple target fields in multiple text images; determining multiple rotation angle ranges of multiple text image sets; and determining the multiple text image sets in which the multiple text images are located based on the multiple rotation angle ranges and multiple rotation angles.
[0024] This technical solution further explains the process of dividing multiple text images into multiple text image sets. Specifically, it determines multiple rotation angles of multiple target fields in multiple text images, that is, the rotation angle of each target field. Then, it determines the rotation angle range of each text image set and compares the rotation angle of the target field with the rotation angle range of the text image set. When the rotation angle of the target field falls within the rotation angle range of a certain text image set, the text image represented by this target field is divided into the corresponding text image set. By determining the text image set to which the text image is divided based on the rotation angle range of the text image set and the rotation angle of the target field, accurate division of text images is achieved, the number of specific angle types of text images is reduced, and the accuracy of target field recognition in text images is improved.
[0025] In any of the above technical solutions, before adding the label to multiple text images in the corresponding text image set, the text image processing method further includes: determining the image type corresponding to the multiple text images based on multiple rotation angles of the multiple text images, wherein the image type includes non-filled image type and filled image type; and performing filling processing on the multiple text images of the filled image type.
[0026] In this technical solution, before adding labels to multiple text images, it is necessary to further classify the multiple text images that have already been classified into rotation angle categories. Specifically, the multiple text images are divided into non-filled image types and filled image types.
[0027] Specifically, since text images may contain other irrelevant information besides the information indicated by the target field, and the fields used to indicate this information can interfere with the target field, and in filled image type text images, the target field itself may affect the detection of the target field rotation angle due to sparse fields or complex field layout, it is necessary to fill the target region so that the target field can be displayed more clearly in the text image, thereby optimizing the target field in the filled image type, making the target field in the filled image type easier to detect, and thus reducing the training difficulty.
[0028] In any of the above technical solutions, the image type corresponding to the multiple text images is determined based on the multiple rotation angles of the multiple text images. Specifically, this includes: determining the target angle and preset range of the non-filled image type; determining that the multiple text images belong to the non-filled image type when the multiple rotation angles of the multiple text images are equal to the target angle; and determining that the multiple text images belong to the filled image type when the multiple rotation angles of the multiple text images are not equal to the target angle. The target angle is N times 90°, and the value of N is in the range of 1≤N≤4.
[0029] In this technical solution, the determination of the non-filled image type is further disclosed, specifically, the target angle and preset range are determined.
[0030] Specifically, the target angles are set to achieve omnidirectional angle correction of the text images in the text image set. The target angles are 90°, 180°, 270°, and 360°. The reason for choosing these four angles is that when the rotation angle of the target field of the text image is one of these four angles, the pixel distribution of the text image is consistent with the pixel distribution of the corrected text image. Therefore, the text image can clearly represent the target field without the need for padding. However, when the rotation angle of the target field of the text image is another angle, the text image needs to be padded to eliminate interference from other image data in the text image, thereby improving the accuracy of the target field identification.
[0031] Once the target angle is determined, multiple text images with the same target angle are grouped into a single text image set, i.e., non-filled image type. Subsequently, multiple text images with different target angles are grouped into a single text image set, i.e. filled image type. This grouping method enables the judgment of the angle of text images in the text image set from all directions. At the same time, by setting the target angle, the accuracy of the text image processing method of the present invention in judging and correcting the rotation angle is improved when the rotation angle is the target angle.
[0032] In any of the above technical solutions, the filling process for multiple text images of the filling image type specifically includes: determining multiple target regions of the multiple text images of the filling image type, wherein the multiple target regions are used to define target fields in the text images; determining multiple text image parameters of the multiple text images; and performing filling processing on the multiple target regions in the multiple text images according to the multiple text image parameters.
[0033] In this technical solution, firstly, multiple target regions of multiple text images belonging to the filled image type are determined. The target region is the region used to define the target field in the text image. In other words, the target region is the region outside the region where the target field is located in the text image.
[0034] Once the area to be filled is determined, the filling method for the target area is then determined. Specifically, multiple text image parameters belonging to the same text image type are identified, and the target area is filled according to these different text image parameters. Since each category of text images differs, an appropriate filling method needs to be selected based on the text image parameters of each of the multiple text images belonging to the same text image type.
[0035] The embodiments of this application fill multiple target regions of multiple text images belonging to the filled image type, so that the target fields of the text images belonging to the filled image type can be displayed more clearly on the corresponding text images, eliminating the interference of other information in the text images on the target fields. By filling the second text image with different text image parameters, the accuracy of identifying the target fields in each text image belonging to the filled image type and the accuracy of correcting the rotation angle of the target fields are further improved.
[0036] In any of the above technical solutions, the filling process for multiple target regions in multiple text images based on multiple text image parameters specifically includes: when the multiple text image parameters are first text image parameters, zero-padding processing is performed on the multiple target regions; when the multiple text image parameters are second text image parameters, pixel filling processing is performed on the multiple target regions; the text image parameters are used to represent the integrity of the text image edges.
[0037] In this technical solution, the text image parameters are used to represent the integrity of the edges of the second text image, that is, whether the edges of the text image are complete.
[0038] Specifically, the first text image parameter indicates that the edges of the text image are complete; the second text image parameter indicates that the edges of the text image are incomplete. For text images with complete edges, zero-padding is used to enhance the recognizability of the rotation angle of the target field in the text image. For text images with incomplete edges, pixel-padding is used to mitigate the adverse effects of the edges of the text image on the rotation angle of the target field in the text image.
[0039] The embodiments of this application optimize the target fields in the text images by filling multiple target regions in multiple text images, thereby making the target fields in the text images more clearly displayed, thus improving the accuracy of target field recognition in text images and the accuracy of target field rotation angle correction.
[0040] Specifically, zero-padding makes the edge features of text images more prominent, thereby enhancing the ability of the preset model to identify different types of target fields in the text image, and the operation is simple; pixel-padding reduces the impact of the edge features of the text image on the preset model, improving the preset model's ability to determine the rotation angle of the target field in the text image. This leads to optimization of multiple text images belonging to the padded image type and to the performance of the preset model.
[0041] In any of the above technical solutions, after correcting the rotation angle of the target field in the text image to be processed by the trained preset model, the text image processing method further includes: outputting the rotation angle of the target field in the text image to be processed.
[0042] In this technical solution, the preset model not only outputs the corrected text image to be processed, but also outputs the rotation angle of the target field in the text image to be processed. This enables the preset model to achieve the technical effect of detecting the rotation angle of the target field in the text image to be processed, making it easier for users to understand and thus improving the user experience.
[0043] According to a second aspect of the present invention, a text image processing apparatus is provided, comprising: an acquisition unit for acquiring a set of text images; a processing unit for determining multiple sets of text images; the processing unit is further configured to: divide the multiple text images into multiple sets of text images according to a rotation angle and a range of rotation angles; generate multiple labels according to the multiple sets of text images; add the labels to the multiple text images in the corresponding sets of text images; a training unit for training a preset model based on the multiple text images after adding the multiple labels; and an application unit for correcting the rotation angle of a target field in the text image to be processed using the trained preset model.
[0044] The text image processing apparatus provided by this invention includes an acquisition unit, a processing unit, a training unit, and an application unit. The acquisition unit first acquires a set of text images for training, which includes multiple types of text images. Specifically, the text images can be images of different types of receipts, such as taxi receipts and train tickets, which have different text distributions and layout structures. Subsequently, the acquired text images are designed and processed into text images that can be used to train a preset model.
[0045] Furthermore, the text image includes multiple target fields with rotation angles. The target fields in the text image can be fields with feature information in the text image. Users can limit the information of the text image as needed, and thus limit the target fields in the text image. The rotation angle is the tilt angle between the arrangement direction of the target fields and the horizontal direction. In the same image, the rotation angle of the target fields is the same. Furthermore, any field in the same image, that is, the text field in the text image, also has the same rotation angle.
[0046] After acquiring multiple types of text images, the processing unit first determines multiple text image sets and divides the multiple text images into multiple text image sets. Specifically, each text image set represents a category of text images. The rotation angle of the target field is different between text images of different categories. Since the target field of each text image has a rotation angle, in the downstream invoice OCR (Optical Character Recognition) task, when the rotation angle of the target field in the text image differs from that in the horizontal direction by less than 10 degrees, the impact on the final recognition effect of the recognition task is reduced, that is, the impact on the recognition accuracy is small. Therefore, in order to ensure the recognition accuracy and reduce the amount of computation, the processing unit divides text images with similar rotation angles into one category. When the processing unit recognizes the rotation angle of multiple text images in this category, it outputs a rotation angle, thereby ensuring the recognition accuracy, reducing the computation time, and improving the cost-effectiveness of this technical solution.
[0047] Subsequently, the processing unit generates multiple labels based on multiple text image sets.
[0048] In the technical solution of this application, the label is used to indicate the rotation angle characteristics of the text image in different text image categories. Since the rotation angle of the target field of text images belonging to different text image categories is different, the processing unit can better indicate the difference between the rotation angles of the target field in text images of different text image categories by adding different labels to text images of different categories.
[0049] After the processing unit adds the corresponding label to each of the multiple text images, the training unit uses these text images as training text images for a pre-defined model. The pre-defined model is then trained to learn target fields with different rotation angles in text images of different categories. Subsequently, the application unit uses the trained pre-defined model to correct the rotation angles of the target fields in the text images to be processed.
[0050] In the technical solution of this application, correction means adjusting the tilt angle between the arrangement direction of the target field in the text image and the horizontal direction, so that the direction of the target field in the text image is consistent with the horizontal direction.
[0051] In the technical solution of this application, the accuracy of identifying target fields in text images is improved by dividing multiple text images into multiple text image sets.
[0052] The processing unit adds multiple different labels to the text images in multiple second text image sets and multiple first text image sets, enabling the addition of different identifiers for the rotation angles of different target fields. This facilitates the training of the preset model, improving training speed and effectiveness. Furthermore, the text image processing device of this invention simplifies the text image processing process by having the processing unit classify and label multiple text images, and the training unit uses the labeled text images to train the preset model. This further improves the text image processing speed and reduces resource consumption.
[0053] Furthermore, when the type of text image changes or the category of the target field in the text image changes, the acquired multiple text images can be updated so that the acquisition unit can acquire the changed text image, thereby improving the generalization ability of the preset model, meeting the timeliness requirements of the preset model for text image processing, improving the processing effect of the text image processing device and the processing performance of the preset model after training by the training unit. At the same time, the output of the training unit is the corrected text image to be processed, which meets the accuracy requirements of the downstream invoice OCR task for text images, thus facilitating OCR task integration.
[0054] Furthermore, after dividing multiple text images into multiple text image sets, the processing unit also needs to scale the multiple text images, that is, adjust the text images to a standard size, so as to avoid the text images with inconsistent scales affecting the training of the preset model, and thus affecting the accuracy of the preset model in detecting the rotation angle of the target field in the text image to be processed.
[0055] The text image processing device of this technical solution can quickly output the rotation angle of the target field in the current text image to be detected, while having a high accuracy rate. In addition, this method can support detection at any angle with small error.
[0056] Meanwhile, this text image processing device can update the images in the text image set at any time, realizing the automatic expansion and generation of the text image set. This enables the text image processing method to meet the needs of various types of text image detection. The trained model has strong generalization ability and a wide range of applications.
[0057] According to a third aspect of the present invention, a text image processing system is proposed, comprising a memory storing a program; and a processor that, when executing the program, implements the text image processing method as described in any of the above-described technical solutions.
[0058] The text image processing system provided by this invention includes a memory and a processor. The memory stores a program; the processor is connected to the memory and executes the program, implementing the text image processing method as described in any of the above technical solutions when the processor executes the program. Therefore, it possesses all the beneficial effects of the text image processing method as described in any of the above technical solutions.
[0059] According to a fourth aspect of the present invention, a readable storage medium is provided on which a program is stored, which, when executed by a processor, implements the text image processing method as described in any of the above technical solutions.
[0060] The readable storage medium provided by the present invention is used to store a program that, when executed by a processor, implements the text image processing method as described in any of the above technical solutions, and therefore has all the beneficial effects of the text image processing method as described in any of the above technical solutions.
[0061] Additional aspects and advantages of the invention will become apparent in the following description or may be learned by practice of the invention. Attached Figure Description
[0062] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:
[0063] Figure 1 One of the flowcharts of the text image processing method according to an embodiment of the present invention is shown;
[0064] Figure 2 A second schematic flowchart of the text image processing method according to an embodiment of the present invention is shown;
[0065] Figure 3 The third schematic flowchart of the text image processing method according to an embodiment of the present invention is shown;
[0066] Figure 4 The fourth schematic flowchart of the text image processing method according to an embodiment of the present invention is shown;
[0067] Figure 5 The fifth schematic flowchart of the text image processing method according to an embodiment of the present invention is shown;
[0068] Figure 6 The sixth schematic flowchart of the text image processing method according to an embodiment of the present invention is shown;
[0069] Figure 7 The seventh flowchart of the text image processing method according to an embodiment of the present invention is shown;
[0070] Figure 8 This illustration shows a schematic diagram of a text image processing method according to an embodiment of the present invention, in which the image type corresponding to the multiple text images is determined based on multiple rotation angles of the multiple text images;
[0071] Figure 9 A structural block diagram of a text image processing apparatus according to an embodiment of the present invention is shown;
[0072] Figure 10 A structural block diagram of a text image processing apparatus according to yet another embodiment of the present invention is shown. Detailed Implementation
[0073] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other.
[0074] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.
[0075] The following reference Figures 1 to 10 This invention describes a text image processing method, a text image processing apparatus, a text image processing system, and a readable storage medium according to some embodiments of the present invention.
[0076] Example 1:
[0077] One embodiment of the present invention proposes a text image processing method. Figure 1 One of the flowcharts of a text image processing method according to an embodiment of this application is shown, such as... Figure 1 As shown, text image processing methods include:
[0078] Step 102, obtain the text image set;
[0079] Step 104: Determine multiple sets of text images;
[0080] Step 106: Divide the multiple text images into multiple text image sets according to the rotation angle and the range of the rotation angle;
[0081] Step 108: Generate multiple labels based on multiple text image sets;
[0082] Step 110: Add the label to multiple text images within the corresponding text image set;
[0083] Step 112: Train the preset model based on multiple text images with added labels;
[0084] Step 114: Correct the rotation angle of the target field in the text image to be processed using the trained preset model.
[0085] Among them, the multiple text images in the text image set include multiple target fields with rotation angles, and the rotation angle of the target fields in each text image is the same; the rotation angle ranges corresponding to the multiple text image sets are different.
[0086] In this embodiment of the application, a set of text images for training is first obtained. The set of text images includes multiple types of text images. Specifically, the text images can be different types of ticket images, such as taxi invoices, train tickets, etc., which have different text distributions and layout structures. Then, the multiple text images obtained are designed and made into text images that can be used to train a preset model.
[0087] Furthermore, the text image includes multiple target fields with rotation angles. The target fields in the text image can be fields with feature information in the text image. Users can limit the information of the text image as needed, and thus limit the target fields in the text image. The rotation angle is the tilt angle between the arrangement direction of the target fields and the horizontal direction. In the same image, the rotation angle of the target fields is the same. Furthermore, any field in the same image, that is, the text field in the text image, also has the same rotation angle.
[0088] After acquiring multiple types of text images, multiple text image sets are first determined and divided into multiple text image sets. Specifically, each text image set represents a category of text images. The rotation angle of the target field differs between text images of different categories. Since the target field of each text image has a rotation angle, in the downstream invoice OCR (Optical Character Recognition) task, when the rotation angle of the target field in the text image differs from that in the horizontal direction by less than 10 degrees, the impact on the final recognition effect of the recognition task is reduced, that is, the impact on the recognition accuracy is small. Therefore, in order to ensure the recognition accuracy and reduce the computational load, text images with similar rotation angles should be divided into one category. When recognizing the rotation angle of multiple text images in this category, a rotation angle is output, thereby ensuring the recognition accuracy, reducing the computation time, and improving the cost-effectiveness of the solution in this embodiment.
[0089] Subsequently, multiple labels are generated based on multiple text image sets.
[0090] In this embodiment of the application, the label is used to indicate the rotation angle characteristics of the text image in different text image categories. Since the rotation angle of the target field is different between text images belonging to different text image categories, by adding different labels to text images of different categories, the difference between the rotation angles of the target field in text images of different text image categories can be better indicated.
[0091] After adding the corresponding label to each of the multiple text images, use it as the training text image for the preset model. Train the preset model so that it can learn target fields with different rotation angles in text images of different categories as learning targets and correct the rotation angle of the target fields in the text images to be processed through the trained preset model.
[0092] In this embodiment of the application, correction means adjusting the tilt angle between the arrangement direction of the target field in the text image and the horizontal direction, so that the direction of the target field in the text image is consistent with the horizontal direction.
[0093] For example, when a user wants to correct the destination information on a train ticket, the system first automatically acquires a large number of train ticket images and determines the target field, which is the destination information text on the train ticket. Then, it determines the rotation angle of the destination information text, which is the tilt angle between the destination information text and the horizontal direction. The multiple train ticket images are divided into multiple categories according to the rotation angle of the target field. Different categories of train ticket images have different rotation angles of the target field. Then, different labels are added to the train ticket images of different categories to add different identifiers for different rotation angles of the target field. The preset model is trained based on multiple labeled train ticket images. After the preset model is trained, the rotation angle of the train ticket with the destination information text to be corrected can be corrected based on the rotation angle data of the destination information text learned during training.
[0094] In this embodiment of the application, by dividing multiple text images into multiple text image sets, the accuracy of identifying target fields in text images is improved.
[0095] By adding multiple different labels to multiple second text image sets and multiple first text image sets, different identifiers can be added to the rotation angles of different target fields, which facilitates the training of the preset model and improves the training speed and training effect. At the same time, the text image processing method of the present invention simplifies the text image processing process by classifying multiple text images, adding labels, and training the preset model with the labeled text images, thereby further improving the text image processing speed and reducing resource consumption.
[0096] Furthermore, when the type of text image changes or the category of the target field in the text image changes, multiple acquired text images can be updated to obtain the changed text image, thereby improving the generalization ability of the preset model, meeting the timeliness requirements of the preset model for text image processing, improving the processing effect of the text image processing method and the processing performance of the trained preset model. At the same time, the output is the corrected text image to be processed, which meets the accuracy requirements of the downstream invoice OCR task for text images, thus facilitating OCR task integration.
[0097] Furthermore, after dividing multiple text images into multiple text image sets, the scale of the multiple text images also needs to be adjusted, that is, the text images are adjusted to a standard size, so as to avoid the text images with inconsistent scales affecting the training of the preset model, and thus affecting the accuracy of the preset model in detecting the rotation angle of the target field in the text image to be processed.
[0098] For example, text images may include images of documents such as general VAT invoices, special VAT invoices, electronic general VAT invoices, electronic special VAT invoices, taxi invoices, train tickets, fixed-amount invoices, airline passenger itineraries, and electronic statements.
[0099] Furthermore, the default model uses the lightweight MobileNetV3-Small model (the MobileNetV3-Small model is a model obtained through neural architecture search proposed by Google on March 21, 2019) as the backbone network to identify and obtain the rotation angle features of the target field in the text image. A global average pooling layer and a fully connected layer are connected in the output head to construct the image processing network.
[0100] The text image processing method in this embodiment can quickly output the rotation angle of the target field in the current text image to be detected, while having a high accuracy rate. In addition, this method can support detection at any angle with small error.
[0101] Meanwhile, this text image processing method can update the images in the text image set at any time, realizing the automatic expansion and generation of the text image set. Thus, this text image processing method can meet the needs of various types of text image detection. The trained model has strong generalization ability and a wide range of applications.
[0102] Example 2
[0103] An embodiment of the present invention provides a text image processing method, such as... Figure 2 As shown, text image processing methods include:
[0104] Step 202: Obtain the text image set;
[0105] Step 204: Determine multiple sets of text images;
[0106] Step 206: Determine multiple rotation angles for multiple target fields in multiple text images;
[0107] Step 208: Determine multiple rotation angle ranges for multiple sets of text images;
[0108] Step 210: Determine the set of multiple text images containing multiple text images based on multiple rotation angle ranges and multiple rotation angles;
[0109] Step 212: Generate multiple labels based on multiple text image sets;
[0110] Step 214: Add the label to the multiple text images in the corresponding text image set;
[0111] Step 216: Train the preset model based on multiple text images with multiple labels added;
[0112] Step 218: Correct the rotation angle of the target field in the text image to be processed using the trained preset model.
[0113] In this embodiment, the process of dividing multiple text images into multiple text image sets is further explained. Specifically, multiple rotation angles of multiple target fields in multiple text images are determined, that is, the rotation angle of each target field. Then, the rotation angle range of each text image set is determined, and the rotation angle of the target field is compared with the rotation angle range of the text image set. When the rotation angle of the target field falls within the rotation angle range of a certain text image set, the text image represented by this target field is divided into the corresponding text image set. By determining the text image set to which the text image is divided based on the rotation angle range of the text image set and the rotation angle of the target field, accurate division of text images is achieved, the specific angle types of text images are reduced, and the accuracy of target field recognition in text images is improved.
[0114] For example, the rotation angle range is set to 5 degrees, thereby dividing the entire rotation angle space into 72 categories of text image sets.
[0115] Example 3
[0116] An embodiment of the present invention provides a text image processing method, such as... Figure 3 As shown, text image processing methods include:
[0117] Step 302: Obtain the text image set;
[0118] Step 304: Determine multiple sets of text images;
[0119] Step 306: Divide the multiple text images into multiple text image sets according to the rotation angle and the range of the rotation angle;
[0120] Step 308: Generate multiple labels based on multiple text image sets;
[0121] Step 310: Determine the image type corresponding to the multiple text images based on the multiple rotation angles of the multiple text images;
[0122] Step 312: Perform fill processing on multiple text images of the fill image type;
[0123] Step 314: Add the label to multiple text images within the corresponding text image set;
[0124] Step 316: Train the preset model based on multiple text images with multiple labels added;
[0125] Step 318: Correct the rotation angle of the target field in the text image to be processed using the trained preset model.
[0126] Among them, image types include non-filled image types and filled image types;
[0127] In this embodiment, before adding labels to multiple text images, it is necessary to further divide the multiple text images that have already been classified into rotation angle categories. Specifically, the multiple text images are divided into non-filled image types and filled image types.
[0128] Specifically, since text images may contain other irrelevant information besides the information indicated by the target field, and the fields used to indicate this information can interfere with the target field, and in filled image type text images, the target field itself may affect the detection of the target field rotation angle due to sparse fields or complex field layout, it is necessary to fill the target region so that the target field can be displayed more clearly in the text image, thereby optimizing the target field in the filled image type, making the target field in the filled image type easier to detect, and thus reducing the training difficulty.
[0129] Example 4
[0130] An embodiment of the present invention provides a text image processing method, such as... Figure 4 As shown, text image processing methods include:
[0131] Step 402, obtain the text image set;
[0132] Step 404: Determine multiple sets of text images;
[0133] Step 406: Divide the multiple text images into multiple text image sets according to the rotation angle and the range of the rotation angle;
[0134] Step 408: Generate multiple labels based on multiple text image sets;
[0135] Step 410: Determine the target angle and preset range for non-filled image types;
[0136] Step 412: If multiple rotation angles of multiple text images are equal to the target angle, determine that the multiple text images belong to the non-filled image type.
[0137] Step 414: When multiple rotation angles of multiple text images are not equal to the target angle, determine that multiple text images belong to the filled image type;
[0138] Step 416: Perform fill processing on multiple text images of the fill image type;
[0139] Step 418: Add the label to multiple text images within the corresponding text image set;
[0140] Step 420: Train the preset model based on multiple text images with added labels;
[0141] Step 422: Correct the rotation angle of the target field in the text image to be processed using the trained preset model.
[0142] The target angle is N times 90°, and the value of N ranges from 1 to N to 4.
[0143] In this embodiment, the determination of the non-filled image type is further disclosed, specifically, the target angle and the preset range are determined.
[0144] In this embodiment, the target angle is set to achieve omnidirectional angle correction of the text images in the text image set. Specifically, the target angles are 0°, 90°, 180°, and 270°. The reason for selecting these four angles as target angles is that when the rotation angle of the target field of the text image is one of these four angles, the image pixel distribution of the text image is consistent with the image pixel distribution of the corrected text image. Therefore, the text image can clearly represent the target field without the need for filling. However, when the rotation angle of the target field of the text image is another angle, in order to eliminate the interference of other image data in the text image besides the target field, it is necessary to fill the text image, thereby improving the accuracy of the target field judgment.
[0145] Once the target angle is determined, multiple text images with the same target angle are grouped into a single text image set, i.e., non-filled image type. Subsequently, multiple text images with different target angles are grouped into a single text image set, i.e. filled image type. This grouping method enables the judgment of the angle of text images in the text image set from all directions. At the same time, by setting the target angle, the accuracy of the text image processing method of the present invention in judging and correcting the rotation angle is improved when the rotation angle is the target angle.
[0146] For example, Figure 8 This illustration shows a schematic diagram of a text image processing method according to an embodiment of the present invention, in which multiple text images are classified into multiple non-filled image types and multiple filled image types, as shown below. Figure 8 As shown, firstly, four target angles are determined: 0°, 90°, 180°, and 270°. Since the pixel distribution of these four types of text images is basically the same as that of the original image (i.e., the horizontal image), no filling processing is required for these text images. In order to symmetrically and evenly divide the rotation angles of the target fields in multiple text images, the above four angles are used as the standard angles for the corresponding non-filled image types, and text images that meet these four target angles are classified as non-filled image types. When the rotation angle of a text image does not meet the target angle, it is classified as a filled image type. By setting target angles, the accuracy of judging text images with rotation angles of 90°, 180°, and 270° is improved, while the interference of other image data in the text image besides the target field is eliminated, thereby improving the accuracy of text image correction.
[0147] Example 5
[0148] An embodiment of the present invention provides a text image processing method, such as... Figure 5 As shown, text image processing methods include:
[0149] Step 502: Obtain the text image set;
[0150] Step 504: Determine multiple sets of text images;
[0151] Step 506: Divide the multiple text images into multiple text image sets according to the rotation angle and the range of the rotation angle;
[0152] Step 508: Generate multiple labels based on multiple text image sets;
[0153] Step 510: Determine the image type corresponding to the multiple text images based on the multiple rotation angles of the multiple text images;
[0154] Step 512: Determine multiple target regions for multiple text images of the fill image type, the multiple target regions being used to define target fields in the text images;
[0155] Step 514: Determine multiple text image parameters for multiple text images;
[0156] Step 516: Fill multiple target regions in multiple text images according to multiple text image parameters;
[0157] Step 518: Add the label to multiple text images within the corresponding text image set;
[0158] Step 520: Train the preset model based on multiple text images with multiple labels added;
[0159] Step 522: Correct the rotation angle of the target field in the text image to be processed using the trained preset model.
[0160] In this embodiment, firstly, multiple target regions of multiple text images belonging to the filled image type are determined. The target region is the region used to define the target field in the text image. That is, the target region is the region outside the region where the target field is located in the text image.
[0161] Once the area to be filled is determined, the filling method for the target area is then determined. Specifically, multiple text image parameters belonging to the same text image type are identified, and the target area is filled according to these different text image parameters. Since each category of text images differs, an appropriate filling method needs to be selected based on the text image parameters of each of the multiple text images belonging to the same text image type.
[0162] The embodiments of this application fill multiple target regions of multiple text images belonging to the filled image type, so that the target fields of the text images belonging to the filled image type can be displayed more clearly on the corresponding text images, eliminating the interference of other information in the text images on the target fields. By filling the second text image with different text image parameters, the accuracy of identifying the target fields in each text image belonging to the filled image type and the accuracy of correcting the rotation angle of the target fields are further improved.
[0163] Example 6
[0164] An embodiment of the present invention provides a text image processing method, such as... Figure 6 As shown, text image processing methods include:
[0165] Step 602: Obtain the text image set;
[0166] Step 604: Determine multiple sets of text images;
[0167] Step 606: Divide the multiple text images into multiple text image sets according to the rotation angle and the range of the rotation angle;
[0168] Step 608: Generate multiple labels based on multiple text image sets;
[0169] Step 610: Determine the image type corresponding to the multiple text images based on the multiple rotation angles of the multiple text images;
[0170] Step 612: Determine multiple target regions for multiple text images of the fill image type, the multiple target regions being used to define target fields in the text images;
[0171] Step 614: Determine multiple text image parameters for multiple text images;
[0172] Step 616: When multiple text image parameters are the same as the first text image parameters, zero-padding is performed on multiple target regions.
[0173] Step 618: When multiple text image parameters are the second text image parameters, perform pixel filling processing on multiple target regions;
[0174] Step 620: Add the label to the multiple text images in the corresponding text image set;
[0175] Step 622: Train the preset model based on multiple text images with multiple labels added;
[0176] Step 624: Correct the rotation angle of the target field in the text image to be processed using the trained preset model.
[0177] Among them, the text image parameters are used to represent the integrity of the text image edges.
[0178] In this embodiment, the text image parameter is used to represent the integrity of the edges of the second text image, that is, whether the edges of the text image are complete.
[0179] Specifically, the first text image parameter indicates that the edges of the text image are complete; the second text image parameter indicates that the edges of the text image are incomplete. For text images with complete edges, zero-padding is used to enhance the recognizability of the rotation angle of the target field in the text image. For text images with incomplete edges, pixel-padding is used to mitigate the adverse effects of the edges of the text image on the rotation angle of the target field in the text image.
[0180] The embodiments of this application optimize the target fields in the text images by filling multiple target regions in multiple text images, thereby making the target fields in the text images more clearly displayed, thus improving the accuracy of target field recognition in text images and the accuracy of target field rotation angle correction.
[0181] Specifically, zero-padding makes the edge features of text images more prominent, thereby enhancing the ability of the preset model to identify different types of target fields in the text image, and the operation is simple; pixel-padding reduces the impact of the edge features of the text image on the preset model, improving the preset model's ability to determine the rotation angle of the target field in the text image. This leads to optimization of multiple text images belonging to the padded image type and to the performance of the preset model.
[0182] Example 7
[0183] An embodiment of the present invention provides a text image processing method, such as... Figure 7 As shown, text image processing methods include:
[0184] Step 702, obtain the text image set;
[0185] Step 704: Determine multiple sets of text images;
[0186] Step 706: Divide the multiple text images into multiple text image sets according to the rotation angle and the range of the rotation angle;
[0187] Step 708: Generate multiple labels based on multiple text image sets;
[0188] Step 710: Add the label to multiple text images in the corresponding text image set;
[0189] Step 712: Train the preset model based on multiple text images with added labels;
[0190] Step 714: Correct the rotation angle of the target field in the text image to be processed using the trained preset model;
[0191] Step 716: Output the rotation angle of the target field in the text image to be processed.
[0192] In this embodiment, the preset model not only outputs the corrected text image to be processed, but also outputs the rotation angle of the target field in the text image to be processed. This enables the preset model to achieve the technical effect of detecting the rotation angle of the target field in the text image to be processed, making it easier for users to understand and thus improving the user experience.
[0193] Example 8
[0194] Another embodiment of the present invention provides a text image processing apparatus 900. Figure 9 A structural block diagram of a text image processing apparatus 900 according to an embodiment of this application is shown, such as... Figure 9 As shown, the text image processing device 900 includes: an acquisition unit 902 for acquiring a set of text images; a processing unit 904 for determining multiple sets of text images; the processing unit 904 is further configured to: divide the multiple text images into multiple sets of text images according to rotation angles and rotation angle ranges; generate multiple labels according to the multiple sets of text images; add the labels to the multiple text images in the corresponding sets of text images; a training unit 906 for training a preset model based on the multiple text images after adding multiple labels; and an application unit 908 for correcting the rotation angle of the target field in the text image to be processed using the trained preset model.
[0195] The text image processing apparatus 900 provided in this application embodiment includes an acquisition unit 902, a processing unit 904, a training unit 906, and an application unit 908. The acquisition unit first acquires a set of text images for training, which includes multiple types of text images. Specifically, the text images can be images of different types of receipts, such as taxi receipts, train tickets, etc., with different text distributions and layout structures. Subsequently, the acquired multiple text images are designed and produced into text images that can be used to train a preset model.
[0196] Furthermore, the text image includes multiple target fields with rotation angles. The target fields in the text image can be fields with feature information in the text image. Users can limit the information of the text image as needed, and thus limit the target fields in the text image. The rotation angle is the tilt angle between the arrangement direction of the target fields and the horizontal direction. In the same image, the rotation angle of the target fields is the same. Furthermore, any field in the same image, that is, the text field in the text image, also has the same rotation angle.
[0197] After acquiring multiple types of text images, the processing unit 904 first determines multiple text image sets and divides the multiple text images into multiple text image sets. Specifically, each text image set represents a category of text images. The rotation angle of the target field is different between text images of different categories. Since the target field of each text image has a rotation angle, in the downstream invoice OCR (Optical Character Recognition) task, when the rotation angle of the target field in the text image differs from that in the horizontal direction by less than 10 degrees, the impact on the final recognition effect of the recognition task is reduced, that is, the impact on the recognition accuracy is small. Therefore, in order to ensure the recognition accuracy and reduce the amount of computation, the processing unit 904 divides text images with similar rotation angles into one category. When the processing unit 904 recognizes the rotation angle of multiple text images in this category, it outputs a rotation angle, thereby ensuring the recognition accuracy, reducing the computation time, and improving the cost-effectiveness of this technical solution.
[0198] Subsequently, the processing unit 904 generates multiple labels based on multiple text image sets.
[0199] In the technical solution of this application, the label is used to indicate the rotation angle characteristics of the text image in different text image categories. Since the rotation angle of the target field of text images belonging to different text image categories is different, the processing unit 904 can better indicate the difference between the rotation angles of the target field in text images of different text image categories by adding different labels to text images of different categories.
[0200] After the processing unit 904 adds the corresponding label to each of the multiple text images, the training unit 906 uses it as the training text image of the preset model to train the preset model so that the preset model can learn the target fields with different rotation angles in the text images of different categories as learning targets. Subsequently, the application unit 908 corrects the rotation angle of the target fields in the text images to be processed through the trained preset model.
[0201] In the technical solution of this application, correction means adjusting the tilt angle between the arrangement direction of the target field in the text image and the horizontal direction, so that the direction of the target field in the text image is consistent with the horizontal direction.
[0202] In the technical solution of this application, the accuracy of identifying target fields in text images is improved by dividing multiple text images into multiple text image sets.
[0203] The processing unit 904 adds multiple different labels to the text images in the multiple second text image sets and the multiple first text image sets, which can add different identifiers to the rotation angles of different target fields, making it easier to train the preset model and improving the training speed and training effect. At the same time, the text image processing device of the present invention, through the processing unit 904 classifying and labeling multiple text images, and the training unit 906 training the preset model with the labeled text images, simplifies the text image processing process, further improves the text image processing speed, and reduces resource consumption.
[0204] Furthermore, when the type of text image changes or the category of the target field in the text image changes, the acquired multiple text images can be updated so that the acquisition unit 902 can acquire the changed text image, thereby improving the generalization ability of the preset model, meeting the timeliness requirements of the preset model for text image processing, improving the processing effect of the text image processing device and the processing performance of the preset model trained by the training unit 906. At the same time, the output of the training unit 906 is the corrected text image to be processed, which meets the accuracy requirements of the downstream invoice OCR task for text images, thus facilitating OCR task integration.
[0205] Furthermore, after dividing multiple text images into multiple text image sets, the processing unit 904 also needs to scale the multiple text images, that is, adjust the text images to a standard size, so as to avoid the text images with inconsistent scales affecting the training of the preset model, and thus affecting the accuracy of the preset model in detecting the rotation angle of the target field in the text image to be processed.
[0206] The text image processing device of this technical solution can quickly output the rotation angle of the target field in the current text image to be detected, while having a high accuracy rate. In addition, this method can support detection at any angle with small error.
[0207] Meanwhile, this text image processing device can update the images in the text image set at any time, realizing the automatic expansion and generation of the text image set. This enables the text image processing method to meet the needs of various types of text image detection. The trained model has strong generalization ability and a wide range of applications.
[0208] Example 9
[0209] Another embodiment of the present invention provides a text image processing apparatus. Figure 10 A structural block diagram of a text image processing apparatus according to an embodiment of this application is shown, such as... Figure 10 As shown, the text image processing device includes an acquisition unit 1010, a processing unit 1020, and a training unit 1030. The acquisition unit 1010 acquires multiple text images of different types; the processing unit 1020 adjusts the pose of the multiple text images of different types; the processing unit 1020 further divides the pose-adjusted text images into two categories: unfilled and filled text images; performs filling processing on target regions in the multiple filled text images; and adds multiple different labels to the multiple filled text images and the multiple unfilled text images. The training unit 1030 trains a preset model based on the multiple text images with added labels.
[0210] Figure 10 The process of training a preset model by the text image processing apparatus according to this embodiment is also illustrated. Specifically, after the acquisition unit 1010 acquires multiple text images of different types, the processing unit 1020 first adjusts the pose of the multiple text images of different types, that is, determines the target field of the text image. The target field can be a field with feature information in the text image. The user can limit the information of the text image as needed, and then limit the target field in the text image. In this embodiment, the target field is 11111.
[0211] After the processing unit 1020 determines the target field of the text image, it determines multiple text image sets and divides the multiple text images into multiple text image sets. Specifically, each text image set represents a category of text images. The rotation angle of the target field is different between text images of different categories. Specifically, the processing unit 1020 determines multiple rotation angles of multiple target fields in multiple text images, that is, the rotation angle of each target field. Subsequently, it determines the rotation angle range of each text image set and compares the rotation angle of the target field with the rotation angle range of the text image set. When the rotation angle of the target field falls within the rotation angle range of a certain text image set, the text image represented by this target field is divided into the corresponding text image set.
[0212] Subsequently, the processing unit 1020 divides multiple text images into non-filled text images and filled text images according to the different rotation angles of the target fields in different text images. Specifically, it determines four target angles, namely 0°, 90°, 180°, and 270°, and uses these four angles as the rotation angles of the corresponding non-filled text images. When the rotation angle of the target field in the text image is equal to the target angle, the text image is classified as a non-filled text image. When the rotation angle of the target field in the text image is not equal to the target angle, the text image is classified as a filled text image.
[0213] Subsequently, processing unit 1020 scales multiple text images to bring them all to a uniform size. This prevents inconsistently scaled text images from affecting the training of the preset model in training unit 1030, which in turn affects the accuracy of the preset model in detecting the rotation angle of the target field in the text image to be processed. Next, the text image is filled. Specifically, the area outside the target field in the text image is defined as the target area, i.e., the area of the text image excluding the target field. The target area is then filled, allowing the target field in the filled text image to be more clearly displayed, eliminating interference from other information in the filled text image.
[0214] Subsequently, the processing unit 1020 adds multiple different labels to the multiple filled text images and multiple unfilled text images after filling processing, so that the preset model can determine the target field in the text image to be processed, that is, the rotation angle of the target field, according to the labels after training. Then, the multiple text images with multiple labels are used to train the preset model. Specifically, the labels of the text images in different text images are different. Since the multiple text images in this embodiment belong to different categories, it is necessary to add different labels to them.
[0215] Subsequently, the training unit 1030 trains the preset model based on three text images with different labels. Specifically, the text images corresponding to multiple text images are placed on the right. The preset model places the text images with the rotation angles of the five unknown target fields on the left into three categories of text images. When the rotation angle of the target field in the text image placed therein meets the rotation angle range of any one of the three text images, the text image is classified into the category to which the corresponding text image belongs.
[0216] The text image processing apparatus of this embodiment divides multiple text images into two categories through the processing unit 1020, and performs filling processing on multiple target regions in multiple filled text images. This avoids the influence of other fields in the text images other than the target field on the recognition of the target field, and also solves the problem of the influence of complex text layout or sparse text on the recognition of the target field, thereby improving the accuracy of target field recognition in filled text images. The processing unit 1020 adds multiple different labels to the text images in multiple filled text images and multiple unfilled text images, which can add different identifiers for the rotation angle of different target fields, making it easier for the training unit 1030 to train the preset model, thereby improving the training speed and training effect of the training unit 1030.
[0217] Example 10
[0218] Another embodiment of the present invention provides a text image processing system, which includes a memory storing a program; and a processor that executes the program to implement the text image processing method as described in any of the above embodiments.
[0219] The text image processing system provided in this application includes a memory and a processor. The memory stores a program; the processor is connected to the memory and executes the program, implementing the text image processing method as described in any of the above embodiments. Therefore, it possesses all the beneficial effects of the text image processing method described in any of the above technical solutions, which will not be elaborated further here.
[0220] Example 11
[0221] Another embodiment of the present invention provides a readable storage medium on which a program is stored, which, when executed by a processor, implements the text image processing method as described in any of the above technical solutions.
[0222] The readable storage medium provided in this application embodiment is used to store a program that implements the text image processing method as described in any one of the above technical solutions when executed by a processor. Therefore, it has all the beneficial effects of the text image processing method as described in any one of the above technical solutions, which will not be repeated here.
[0223] In the description of this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance, unless otherwise expressly specified and limited. The terms "connection," "installation," and "fixing," etc., should be interpreted broadly. For example, "connection" can mean a fixed connection, a detachable connection, or an integral connection; it can mean a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0224] In the description of this specification, the terms "one embodiment," "some embodiments," "specific embodiment," etc., refer to a specific feature, structure, material, or characteristic described in connection with that embodiment or example, which is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0225] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A text image processing method, characterized in that, include: Obtain a text image set, wherein multiple text images in the text image set include multiple target fields with rotation angles, and the rotation angle of the target fields in each text image is the same; Multiple sets of text images are identified, and the rotation angle ranges corresponding to these multiple sets of text images are different; The plurality of text images are divided into the plurality of text image sets according to the rotation angle and the range of the rotation angle; Generate multiple labels based on the multiple text image sets; Add the label to multiple text images within the corresponding text image set; The preset model is trained based on the multiple text images after adding multiple labels; The rotation angle of the target field in the text image to be processed is corrected by the trained preset model; Before adding the label to the multiple text images in the corresponding text image set, the text image processing method further includes: Based on multiple rotation angles of the multiple text images, the image type corresponding to the multiple text images is determined, and the image type includes non-filled image type and filled image type; Perform fill processing on multiple text images of the fill image type; The filling process for multiple text images of the filled image type specifically includes: Determine multiple target regions for the plurality of text images of the filled image type, the plurality of target regions being used to define target fields in the text images, wherein the target regions are regions outside the region where the target field is located in the text images; Determine multiple text image parameters for the multiple text images; The multiple target regions in the multiple text images are filled according to the multiple text image parameters; After correcting the rotation angle of the target field in the text image to be processed using the trained preset model, the text image processing method further includes: Output the rotation angle of the target field in the text image to be processed.
2. The text image processing method according to claim 1, characterized in that, The step of dividing the plurality of text images into the plurality of text image sets according to the rotation angle and the rotation angle range specifically includes: Determine multiple rotation angles for the multiple target fields in the multiple text images; Determine multiple rotation angle ranges for the multiple text image sets; The set of multiple text images is determined based on the multiple rotation angle ranges and the multiple rotation angles.
3. The text image processing method according to claim 1, characterized in that, The step of determining the image type corresponding to the multiple text images based on multiple rotation angles of the multiple text images specifically includes: Determine the target angle and preset range of the non-filled image type; If the rotation angles of the multiple text images are equal to the target angle, then the multiple text images are determined to belong to the non-filled image type. If the rotation angles of the multiple text images are not equal to the target angle, the multiple text images are determined to belong to the filled image type. The target angle is N times 90°, and the value of N ranges from 1 to N and from 4 to 4.
4. The text image processing method according to claim 1, characterized in that, The process of filling the multiple target regions in the multiple text images according to the multiple text image parameters specifically includes: When the multiple text image parameters are the first text image parameters, zero-padding is performed on the multiple target regions; When the multiple text image parameters are the second text image parameters, pixel filling processing is performed on the multiple target regions; The text image parameters are used to represent the integrity of the text image edges.
5. A text image processing device, characterized in that, The text image processing apparatus employs the text image processing method as described in any one of claims 1 to 4, and the text image processing apparatus comprises: The acquisition unit is used to acquire a text image set; Processing unit, used to determine multiple text image sets; The processing unit is also used for: The plurality of text images are divided into the plurality of text image sets according to the rotation angle and the range of the rotation angle; Generate multiple labels based on the multiple text image sets; Add the label to multiple text images within the corresponding text image set; The training unit is used to train the preset model based on the multiple text images after adding multiple labels; The application unit is used to correct the rotation angle of the target field in the text image to be processed using the trained preset model.
6. A text image processing system, characterized in that, include: A memory, wherein the memory stores a program; A processor that, when executing the program, implements the text image processing method as described in any one of claims 1 to 4.
7. A readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the text image processing method as described in any one of claims 1 to 4.