Handwritten text detection method and device, server and storage medium
By performing text type recognition and marking of handwritten text regions, the problem of low text detection accuracy in existing technologies is solved, achieving efficient recognition of handwritten text and improving the accuracy and reliability of text detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN BANK CO LTD
- Filing Date
- 2022-11-04
- Publication Date
- 2026-05-12
AI Technical Summary
Existing deep learning-based text detection technologies have low accuracy when dealing with diverse digital documents containing various specialized printed fonts, handwritten fonts, and text slant.
By acquiring the image to be detected, text detection and text type recognition are performed to determine the text region. The handwritten text region is then labeled and the characters are recognized. The labeled image and the character recognition results are output. The accuracy of text detection is improved by using a classification model and edge detection operators.
It effectively identifies the text type in a text region, improves the accuracy of text detection, provides a foundation for text recognition, and enhances the reliability of text recognition.
Smart Images

Figure CN115620315B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of text detection and recognition, and specifically to a handwritten text detection method, apparatus, server, and storage medium. Background Technology
[0002] Commonly used deep learning-based text detection techniques can achieve good detection results for regular printed text, such as ID card, bank card recognition, license plate recognition, and PDF to Word conversion. However, when dealing with diverse digital document types that include various specialized printed fonts, handwritten fonts, and text slant, the accuracy of existing text detection techniques is low. Summary of the Invention
[0003] This invention provides a handwritten text detection method, apparatus, server, and storage medium to improve the accuracy of text detection.
[0004] On one hand, embodiments of the present invention provide a handwritten text detection method, the method comprising:
[0005] Acquire the image to be detected;
[0006] Text detection is performed on the image to be detected to determine the text regions in the image to be detected;
[0007] Perform text type detection on the text region to obtain the text type of the text in the text region;
[0008] If there is a target text region in the text region that is handwritten text, then the target text region is marked to obtain the marked image to be detected;
[0009] Perform character recognition on the handwritten text in the target text region to obtain the character recognition result of the target text region;
[0010] Output the labeled image to be detected and the text recognition results of the target text region in the labeled image to be detected.
[0011] On the other hand, embodiments of the present invention provide a handwritten text detection device, the device comprising:
[0012] The acquisition module is used to acquire the image to be detected;
[0013] The text region detection module is used to perform text detection on the image to be detected and determine the text regions in the image to be detected.
[0014] The text type detection module is used to perform text type detection on each of the text regions and determine the text type of the text in each of the text regions;
[0015] The marking module is used to mark the target text region if there is a target text region of handwritten text type in each of the text regions, so as to obtain the marked image to be detected;
[0016] The text detection module is used to perform character recognition on the handwritten text in the target text region to obtain the character recognition result of the target text region;
[0017] The output module is used to output the labeled image to be detected and the text recognition results of the target text region in the labeled image to be detected.
[0018] On the other hand, embodiments of the present invention provide a server, including a memory and a processor; the memory stores an application program, and the processor is used to run the application program in the memory to perform the operations in the above-described handwritten text detection method.
[0019] On the other hand, embodiments of the present invention provide a storage medium storing a plurality of instructions, which are adapted for loading by a processor to execute the steps in the above-described handwritten text detection method.
[0020] This invention involves acquiring an image to be detected; performing text detection on the image to determine text regions; performing text type detection on each text region to determine the text type of the text in each text region; if a target text region with handwritten text exists in each text region, the target text region is marked to obtain a marked image to be detected; performing character recognition on the handwritten text in the target text region to obtain the character recognition result of the target text region; and outputting the marked image to be detected and the character recognition result of the target text region in the marked image to be detected. This invention, through text detection, can determine text regions in an image to be detected, and through text type detection of text regions, can effectively determine whether the text in the text region is handwritten text, improving the accuracy of text detection and providing a foundation for text recognition. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram illustrating an application scenario of the handwritten text detection method provided in this embodiment of the invention;
[0023] Figure 2This is a flowchart illustrating the handwritten text detection method provided in an embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram of the structure of the text detection model provided in an embodiment of the present invention;
[0025] Figure 4 This is a schematic diagram of the handwritten text detection device provided in an embodiment of the present invention;
[0026] Figure 5 This is a schematic diagram of the server structure provided in an embodiment of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] As in the background technology, financial and banking operations often require text recognition of scanned digital document images. Since scanned digital document images often contain handwritten and printed fonts, the lack of text type recognition makes it impossible to effectively distinguish between handwritten and printed font areas in the digital document image, reducing the text information extracted from the target area and thus reducing the reliability of the text recognition results.
[0029] Based on this, in order to improve the accuracy of text detection and ensure the reliability of text recognition results, this invention provides a handwritten text detection method applicable to the financial technology field or other fields. This method can determine the text region in the image to be detected through text detection, and can effectively determine whether the text in the text region is a handwritten text line through text type detection, thereby improving the accuracy of text detection and providing a foundation for text recognition.
[0030] like Figure 1 As shown, Figure 1 This is a schematic diagram of an application scenario of the handwritten text detection method provided in the embodiments of the present invention. The application scenario shown includes a client 103, a server 101, and a network 102.
[0031] In this system, client 103 connects to server 101 via a network. Server 101 receives the image to be detected sent by client 103 via network 102, performs handwritten text detection on the image to be detected, determines the text regions in the image to be detected, performs text type detection on each text region to obtain the text type of the text in each text region, and if there is a target text region with the text type being handwritten text in each text region, the target text region is marked to obtain the marked image to be detected. The handwritten text in the target text region is then recognized to obtain the text recognition result of the target text region. The marked image to be detected and the text recognition result of the target text region in the marked image to be detected are output. Finally, the marked image to be detected and the text recognition result of the target text region in the marked image to be detected are returned to client 103 via network 102.
[0032] In some embodiments of the present invention, the client 103 includes, but is not limited to, various personal computers, laptops, smart devices, tablets, and portable wearable devices. The server 101 can be a standalone server, or a server network or server cluster composed of servers, such as computers, network hosts, a single network server, multiple network servers, or a cloud server composed of multiple servers. The cloud server is composed of a large number of computers or network servers based on cloud computing.
[0033] In some embodiments of the present invention, the network 102 may be a wired network or a wireless network. In some embodiments of the present invention, the wired network or wireless network uses standard communication technologies and / or protocols. The network may be the Internet or any network, including but not limited to wide area networks, metropolitan area networks, local area networks, 3rd Generation Partnership Project (3GPP), Long Term Evolution (LTE), Worldwide Interoperability for Microwave Access (WiMAX) mobile communications, or computer network communications based on the TCP / IP protocol suite, User Datagram Protocol (UDP), etc.
[0034] like Figure 2 As shown, Figure 2 This is a flowchart illustrating the handwritten text detection method provided in an embodiment of the present invention. The handwritten text detection method shown is applied to... Figure 1The server 101 shown specifically includes steps 201 to 206 in its handwritten text detection method:
[0035] 201, Obtain the image to be detected.
[0036] In some embodiments of the present invention, the image to be detected sent by the client 103 can be obtained when a text detection instruction is received from the client 103.
[0037] In some embodiments of the present invention, in response to a text detection request sent by client 103, a text image acquisition instruction can be sent to client 103, causing client 103 to return an image to be detected to server 101. Server 101 then acquires the image to be detected returned by client 103 based on the text image acquisition instruction. Client 103 can either return a stored image to be detected to server 101, or it can invoke an image acquisition device within client 103 to acquire the text to be detected, obtain the image to be detected, and return the acquired image to server 101. The image acquisition device can be a camera or an image sensor.
[0038] 202. Perform text detection on the image to be detected to determine the text regions in the image.
[0039] A text region refers to an image region in the image to be detected that contains text. In some embodiments, the text includes, but is not limited to, handwritten text, printed text, etc.
[0040] In some embodiments of the present invention, text detection of the image to be detected can be performed using a regression-based text detection method. These regression-based text detection methods include, but are not limited to, CTPN, Texbox, EAST, SedLink, MDST, CTD, LDMO, and PCR.
[0041] In some embodiments of the present invention, text detection can also be performed on the image to be detected using a segmentation-based text detection method. Specifically, a text region probability map can be obtained by detecting whether pixels in the image to be detected belong to text targets. A binarized image of the image to be detected is then obtained based on the text region probability map and a preset threshold. The text region is then determined based on the binarized image. Binarization is achieved by setting the pixel values of pixels with probability values greater than or equal to the preset threshold to a first preset value, and setting the pixel values of pixels with probability values less than the preset threshold to a second preset value, based on the text region probability map and the preset threshold. Segmentation-based text detection methods include, but are not limited to, Pixelink, MSR, PSENet, PAN, DBNet, and FCENet.
[0042] In some embodiments of the present invention, text detection can also be performed on the image to be detected using object detection methods to determine the text regions in the image. For example, text detection can be performed on the image to be detected using YOLOV, CNN, R-CNN, FastR-CNN, FasterR-CNN, or VGG object detection methods to determine the text regions in the image.
[0043] 203. Perform text type detection on the text region to obtain the text type of the text in the text region.
[0044] In some embodiments of the present invention, the text type includes handwritten text and printed text.
[0045] In some embodiments of the present invention, a preset classification model can be used to detect the text type of each text region to obtain the text type of the text in each text region.
[0046] In some embodiments, the classification model can be a machine learning-based classification model, such as a logistic regression-based classification model, a random forest-based classification model, a dictionary-based classification model, a clustering-based classification model, etc.
[0047] In other embodiments, the classification model can also be a neural network-based classification model, such as one based on Long Short-Term Memory (LSTM), Convolutional Neural Networks (CNN), De-Convolutional Networks (DN), Deep Neural Networks (DNN), Deep Convolutional Inverse Graphics Networks (DCIGN), Region-based Convolutional Networks (RCNN), Faster Region-based Convolutional Networks (Faster RCNN), and Bidirectional Encoder Representations from Transformers (BERT).
[0048] In some embodiments of the present invention, text type detection can be performed on text regions using template matching to obtain the text type of text in each text region. Specifically, pixel distribution data of the text region is obtained based on the grayscale value or pixel value of the pixels in the text region. The pixel distribution data of the text region is compared with preset pixel distribution template data to determine the similarity between the pixel distribution data of the text region and multiple pixel distribution template data. The pixel distribution template data with the highest similarity is set as the target pixel distribution template data, and the text type corresponding to the target pixel distribution template data is set as the text type of the text in the text region. The pixel distribution template data can be pixel distribution data of different text types stored in advance, or pixel distribution data of different text types generated based on a pre-stored template generation model. The template generation model can be a generative model based on a generative network or a generative adversarial network. The embodiments of the present invention do not specifically limit the template generation model; the pixel distribution data can be a grayscale histogram.
[0049] In some embodiments of the present invention, text type detection can also be performed based on the texture information of the text region to obtain the text type of the text in each text region. Specifically, an image region containing the text region is cropped, and texture detection is performed on the image region containing the text region using a texture detection method to obtain the texture features of the text region. The texture features are compared with multiple benchmark texture features in pre-stored benchmark texture feature data to determine the target benchmark texture feature with the highest similarity. The text type of the target benchmark texture feature is then set as the text type of the text in the text region. The benchmark texture feature data includes text images of various text types and the texture features of each text image.
[0050] 204. If there is a target text region with handwritten text in each text region, then the target text region is marked to obtain the marked image to be detected.
[0051] In some embodiments of the present invention, when there is a target text region in the text region where the text type is handwritten text, the outline of the target text region is determined based on the edge information of the target text region, and the target text region is marked based on the outline of the target text region to obtain the marked image to be detected.
[0052] In some embodiments, edge information of the target text region can be obtained by performing edge detection on the target text region using an edge detection operator, wherein the edge detection operator can be any one of the Sobel operator, Isotropic Sobel operator, Roberts operator, Prewitt operator, Laplacian operator, and Canny operator.
[0053] In other embodiments, in order to improve the accuracy of subsequent text recognition, when there is a target text region in the text region where the text type is handwritten text, the image to be detected can be sharpened, and the edge information of the target text region can be obtained by performing edge detection on the target text region in the sharpened image to be detected using the edge detection operator described above.
[0054] 205. Perform character recognition on the handwritten text in the target text region to obtain the character recognition results of the target text region.
[0055] In some embodiments of the present invention, an image can be captured in the target text region to obtain a handwritten text image in the target text region, and text recognition can be performed on the handwritten text image in the target text region to obtain the text recognition result of the target text region.
[0056] In some embodiments, features can be extracted from the handwritten text image in the target text region using a CTC (Conectionist Temporal Classification) algorithm. This yields text features in the handwritten text image within the target text region. These text features are then encoded and decoded to obtain the text recognition result of the handwritten text image in the target text region. The CTC (Conectionist Temporal Classification) algorithm can be a text recognition algorithm based on CRNN, ResNet, MobileNet, or VGG.
[0057] In other embodiments, the handwritten text image in the target text region can be input into the encoder based on Sequence2Sequence to obtain the semantic vector of the handwritten text image, and then the semantic vector can be decoded by the decoder to obtain the text recognition result.
[0058] In other embodiments, a correction-based text recognition method performs rule transformation on the handwritten text image in the target text region to obtain a transformed handwritten text image, and then performs character recognition on the transformed handwritten text image to obtain the character recognition result.
[0059] In other embodiments, handwritten text in the target text region can be recognized using a Transformer-based method to obtain the text recognition result for the target text region. The Transformer-based method includes, but is not limited to, SRN-based recognition methods, NRTR-based recognition methods, and SRACN-based recognition methods.
[0060] 206. Output the labeled image to be detected and the text recognition results of the target text region in the labeled image to be detected.
[0061] In some embodiments of the present invention, after obtaining the text recognition result, the marked image to be detected and the text recognition result of the target text region in the marked image to be detected are output, and the text recognition result is returned to the client 103.
[0062] The handwritten text detection method provided in this invention can determine the text region in the image to be detected through text detection, and can effectively determine whether the text in the text region is a handwritten text line through text type detection, thereby improving the accuracy of text detection and providing a foundation for text recognition.
[0063] In some embodiments of the present invention, considering that segmentation-based text detection algorithms have good performance in text detection methods, but segmentation-based text detection algorithms can only identify text regions in the image to be detected, and cannot effectively identify the text type of the text region. If another model or another algorithm is used to identify the text type of the text region after obtaining the text region, it may increase the detection time of the handwritten text detection method, and it is necessary to adapt to the data specifications between different models, which will increase the workload. Therefore, in order to improve the efficiency and accuracy of text detection, the embodiments of the present invention add text type detection to the segmentation-based text detection algorithm to obtain a new segmentation-based text detection algorithm, which realizes the detection of text regions in the image to be detected and the identification of text types in the text regions.
[0064] For example, the embodiment of the present invention takes the segmentation-based text detection algorithm based on the DBNet network as an example. On the basis of the region segmentation branch of the DBNet network, a classification branch is added to obtain a text detection model. According to the preset text detection model, text detection is performed on the image to be detected to determine the text regions in the image to be detected. Text type detection is performed on each text region to determine the text type of the text in each text region.
[0065] Specifically, such as Figure 3 As shown, Figure 3 This is a schematic diagram of the structure of the text detection model provided in the embodiment of the present invention. The text detection model shown includes an input layer, a feature extraction layer, a binarization layer, a segmentation layer, a classification layer, and an output layer.
[0066] The feature extraction layer, comprising MobileNetV3 and FPN pyramid feature network, is used to perform feature detection on the input image to be detected, obtaining a feature map of the image to be detected. The binarization layer is used to obtain a probability feature map and a probability feature map based on the feature map of the image to be detected, and then performs differentiable binarization based on the probability feature map and the probability feature map to obtain an approximate binarized feature map of the image to be detected. The segmentation layer determines the text regions in the image to be detected based on the approximate binarized feature map. The classification layer is used to predict the category based on the feature map of the image to be detected, obtain the probability that the text in the text region belongs to each text type, and determine the text type of the text in the text region based on the probability that the text in the text region belongs to each text type. The output layer is used to output the location information of the text regions in the image to be detected, and if there is a target text region in the text region with the text type of handwritten text, then the target text region is marked to obtain a marked image to be detected, and the marked image to be detected is output.
[0067] In some embodiments of the present invention, an initial text detection model can be trained using collected sample data including handwritten text to obtain a text detection model.
[0068] In some embodiments of the present invention, in order to improve the generalization ability and processing ability of the text detection model, the initial text detection model can be trained by collecting sample data that includes both handwritten and printed text to obtain the text detection model.
[0069] In some embodiments of the present invention, considering that the images to be detected differ in different scenarios—for example, the images to be detected differ in financial, banking, and educational scenarios—and that using a unified text detection model for text detection requires training the initial text detection model with sample data covering all financial or banking scenarios to make the text detection applicable to different scenarios, the sample data would be large and the training time would be long. Moreover, when a new scenario appears, the sample data needs to be reset and the training repeated, which is not conducive to the widespread application of handwritten text detection methods. Based on this, embodiments of the present invention pre-train the initial model to obtain a pre-trained model, deploy the pre-trained model to the target scenario, and adjust the model parameters of the pre-trained model by collecting sample data from the target scenario to obtain a sample detection model. The initial model, pre-trained model, and sample detection model have the same model structure, but their respective model parameters are different, including but not limited to network weights.
[0070] In some embodiments of the present invention, the method for obtaining the pre-trained model includes steps a1 to a6:
[0071] Step a1: Obtain a handwritten single-character dataset, and randomly combine the handwritten single-character images in the dataset to obtain multiple single-line handwritten text images.
[0072] In some embodiments of the present invention, the handwritten single-character dataset includes a variety of Chinese characters and handwritten single-character images of each Chinese character.
[0073] In some embodiments of the present invention, based on the handwritten single character images of each Chinese character in the handwritten single character dataset, a first preset number of handwritten single character images can be selected and spliced together each time to obtain an image containing a set of handwritten characters. The spliced image containing a set of handwritten characters is set as a single-line handwritten text image. The above steps for generating single-line handwritten text images are repeated to obtain single-line handwritten text images of preset sample data.
[0074] In some embodiments of the present invention, in order to improve the readability of characters in the generated single-line handwritten text image, based on the semantic information of news reports and the handwritten single-character images of each Chinese character in the handwritten single-character dataset, a first preset number of handwritten single-character images are selected and spliced together each time to obtain an image containing a set of handwritten characters. The spliced image containing a set of handwritten characters is set as a single-line handwritten text image. The above steps for generating single-line handwritten text images are repeated to obtain single-line handwritten text images of preset sample data.
[0075] Step a2: Select a preset number of target single-line handwritten text images from multiple single-line handwritten text images, place the selected target single-line handwritten text images in a preset canvas, and determine the handwritten text area of the target single-line handwritten text images in the canvas.
[0076] Step a3: Set the canvas after determining the handwritten text area as the original sample image.
[0077] Step a4: Place the preset printed text in the remaining image areas of the original sample image, excluding the handwritten text area, to obtain the second sample image; the second sample image includes both the printed text area and the handwritten text area.
[0078] In some embodiments of the present invention, considering that digital document images containing both printed and handwritten text have a significant characteristic—that handwritten content is often accompanied by lines of a certain length below it, with a blank area above for users to fill in information—the relative position of the written content to the lines is not fixed when different users write; sometimes the content overlaps with the lines, and sometimes it floats above them. Therefore, to improve the authenticity of the sample data, embodiments of the present invention generate a preset number of initial sample images. Each initial sample image is a blank image with a pixel value of 255. Each initial sample image is set as a blank canvas. For each blank canvas, a preset number of target single-line handwritten text images are selected from multiple single-line handwritten text images. The selected target single-line handwritten text images are placed into the blank canvas to obtain the original sample image. A text line is generated in the image area where the target single-line handwritten text image is located in the original sample image to obtain the initial second sample image. Embodiments of the present invention do not limit the length, color, size, or position of the added text line relative to the image area where the target single-line handwritten text image is located in the original sample image.
[0079] In some embodiments of the present invention, considering that in financial or banking business, the proportion of handwritten text content in digital document images that simultaneously contain printed and handwritten text is less than the proportion of printed text content, in order to increase the data authenticity of the second sample data and thus improve the accuracy of the pre-trained model, the embodiments of the present invention generate printed text content with random font and random length in the remaining image regions of each initial second sample image, except for the image region where the target single-line handwritten text image is located, to obtain a second sample image that simultaneously contains printed text regions and handwritten text regions.
[0080] In some embodiments of the present invention, in order to increase the complexity of the samples and the difficulty of the detection task, text lines can be added to the image area where the printed area is located in the second sample image.
[0081] Step a5: Label the handwritten text regions in each second sample image to obtain the second sample data.
[0082] In some embodiments of the present invention, for each second sample image, annotation information is obtained based on the size information of the second sample image, the actual position information of the handwritten text area in the second sample image, and the text in each handwritten text area. The second sample image is annotated based on the annotation information of each second sample image to obtain second sample data.
[0083] For example, for any second sample image, with the top left corner of the second sample image as the origin, x as the horizontal coordinate, y as the vertical coordinate, w as the length of the handwritten text region, h as the height of the handwritten text region, and points as the coordinates (x, y) of the four points of the handwritten text region, arranged clockwise starting from the top left corner of the handwritten text region; and transcription as the text in the handwritten text region, the annotation information of the second sample image is obtained as: train / 0001.jpg\t[{“transcription”:“xxxxxxx”,“points”:[[x1,y1],[x1+w1,y2],[x1+w1,y1+h1],[x1,y1+h1]]}、{“transcription”:“XXXXXX”,“points”:[[x2,y2],[x2+w2,y2],[x2+w2,y2+h2],[x2,y2+h2]]}。 The image identifier for the second sample image is 0001.jpg.
[0084] Step a6: Input the second sample data into the initial model and train the initial model to obtain the pre-trained model.
[0085] In some embodiments of the present invention, second sample data is input into an initial model for text region prediction to obtain predicted position information of text regions in second sample images from the second sample data. A position training loss is obtained based on the position difference between the actual position information of the handwritten text region in the annotation information of the second sample image and the predicted position information of the text region in the second sample image. A classification training loss is obtained based on the cross-entropy between the predicted position of the text region in the second sample image and the actual position information of the handwritten text region in the annotation information of the second sample image. A total training loss is obtained based on the classification training loss and the position training loss. The initial model is iteratively trained based on the total training loss. When the initial model meets a preset model convergence condition, iterative training stops, and a pre-trained model is obtained. The preset model convergence condition can be either the total training loss being less than or equal to a preset loss threshold, or the number of iterations being greater than or equal to a preset number of iterations threshold.
[0086] In some embodiments of the present invention, Figure 3 Taking the model structure shown as an example, to improve the accuracy of the text region predicted by the model, this embodiment of the invention obtains the training position information of the handwritten text region in each second sample image based on the approximate binarized feature map of the second sample image. Based on the training position information, the target training loss of the initial model is determined. The initial model is then iteratively trained according to the target training loss. When the initial model meets the preset model convergence condition, the iterative training stops, and a pre-trained model is obtained. Specifically, the training method of the initial model includes steps b1 to b5:
[0087] Step b1: Input the second sample image from the second sample data into the initial model for feature extraction to obtain an approximate binarized feature map of each second sample image.
[0088] In some embodiments of the present invention, in step b1, feature extraction can be performed on the second sample image using an initial model to obtain a probability feature map and a threshold feature map. The probability feature map and the threshold feature map are then subjected to differentiable binarization processing to obtain an approximate binarized feature map of the second sample image. Specifically, the method for determining the approximate binarized feature map includes:
[0089] (1) Input the second sample image in the second sample data into the initial model to extract features at different scales, and obtain feature maps at different scales for each second sample image.
[0090] (2) Combine the feature maps of each second sample image at different scales to obtain the combined feature map of each second sample image.
[0091] (3) Perform image convolution on the feature map of each second sample image to obtain the probability feature map of each second sample image, and perform upsampling operation on the feature map of each second sample image to obtain the threshold feature map of each second sample image.
[0092] (4) Differentiable binarization is performed based on the difference between the probability feature map and the threshold feature map of each second sample image to obtain the approximate binarized feature map of each second sample image.
[0093] In some embodiments, feature extraction at different scales can be performed on each second sample image based on MobileNetV3 in the initial model to obtain feature maps at different scales for each second sample image.
[0094] In other embodiments, features at different scales can be extracted from each second sample image based on the ResNet50 network in the initial model to obtain feature maps at different scales for each second sample image.
[0095] In some embodiments of the present invention, feature maps of different scales of each second sample image can be combined using the FPN pyramid feature network in the initial model to obtain the combined feature map of each second sample image.
[0096] In some embodiments of the present invention, the difference P between the probability feature map and the threshold feature map of each second sample image can be used as a basis. i,j -T i,j ,pass Differentiable binarization is performed to obtain an approximate binarized feature map. Here, K is the inflation factor, and P... i,jFor the pixels on the probability feature map of the second sample image, T i,j These are the pixels on the threshold feature map of the second sample image.
[0097] In some embodiments of the present invention, image convolution includes a convolution operation and a deconvolution operation.
[0098] Step b2: Perform contour recognition on the approximate binarized feature map of each second sample image to obtain the training location information of the handwritten text region in each second sample image.
[0099] In some embodiments of the present invention, contour recognition can be performed on the approximate binarized feature map of each second sample image by an edge detection operator to obtain the edge information of the handwritten text region in each second sample image. Based on the edge information of the handwritten text region in each second sample image, the position coordinates of each vertex of the handwritten text region in each second sample image are obtained. Based on the position coordinates of each vertex of the handwritten text region in each second sample image, the training position information of the handwritten text region in each second sample image is determined.
[0100] Step b3: Based on the true binarized image and the approximate binarized feature map of each second sample image, the first training loss of the initial model is obtained. The true binarized image of the second sample image is obtained by binarizing the handwritten text region in the second sample image.
[0101] In some embodiments of the present invention, the first training loss of the initial model can be obtained based on the losses corresponding to the probability feature map, the threshold feature map, and the approximately binarized feature map, respectively. Specifically, the method for determining the first training loss includes:
[0102] (1) The binarization loss is obtained based on the cross-entropy between the real binarized map and the approximate binarized feature map of each second sample image.
[0103] (2) Based on the threshold feature map of each second sample image, determine the predicted handwritten text region of each second sample image, and obtain the threshold loss based on the distance between the predicted handwritten text region of each second sample image and the handwritten text region of each second sample image.
[0104] (3) Determine the first training loss of the initial model based on the binarization loss and the threshold loss.
[0105] In some embodiments of the present invention, the binarization loss L can be obtained. a and threshold loss L b After that, through L b +α*L a +βL b The first training loss of the initial model is obtained.
[0106] In some embodiments of the present invention, the predicted handwritten text region of each second sample image can be obtained by post-processing based on the threshold feature map of each second sample image.
[0107] Step b4: Based on the cross-entropy between the training location information and the actual location information of the handwritten text region in each second sample image, the second training loss of the initial model is obtained. Here, the actual location information refers to the location information of the image region where the target single-line handwritten text image is located in the second sample image.
[0108] Step b5: Obtain the target training loss of the initial model based on the first training loss and the second training loss. Iterate the initial model based on the target training loss of the initial model. Stop the iterative training when the initial model meets the preset model convergence condition, and obtain the pre-trained model.
[0109] In some embodiments, after obtaining the first training loss and the second training loss, the sum of the first training loss and the second training loss can be set as the target training loss of the initial model.
[0110] In other embodiments, after obtaining the first training loss and the second training loss, the mean of the first training loss and the second training loss can be set as the target training loss of the initial model. The mean can be an arithmetic square or a weighted average.
[0111] In other embodiments, the weights corresponding to the first training loss and the second training loss can be obtained, and the sum of the weights of the first training loss and the second training loss can be obtained based on the first training loss and the second training loss and the weights corresponding to the first training loss and the second training loss, and the sum of the weights of the first training loss and the second training loss can be set as the target training loss of the initial model.
[0112] In some embodiments of the present invention, after obtaining the pre-trained model, the pre-trained model is deployed to the target scene, and digital document images containing both printed and handwritten text in the target scene are collected to obtain first sample data. The model parameters of the pre-trained model are then adjusted based on the first sample data to obtain the text detection model. Specifically, the method for determining the text detection model includes steps c1 to c3:
[0113] Step c1: Obtain the first sample data. The first sample data includes multiple first sample images, the handwritten text in each first sample image, and the location information of the handwritten text region where the handwritten text is located.
[0114] In some embodiments of the present invention, each first sample image can be labeled according to the location information of the handwritten text area where the handwritten text is located, following step a5 above.
[0115] Step c2: Input the first sample data into the pre-trained model for text detection to obtain the predicted location information of the handwritten text region in each first sample image.
[0116] Step c3: Based on the predicted location information of the handwritten text region in each first sample image and the location information of the handwritten text region in each first sample image, adjust the model parameters of the pre-trained model to obtain the text detection model.
[0117] In some embodiments of the present invention, the predicted position information of the handwritten text region in the first sample image and the position difference between the position information of the handwritten text region in each first sample image can be compared with a preset position difference threshold. If each position difference is less than or equal to the preset position difference threshold, it indicates that the prediction accuracy of the pre-trained model meets the requirements, and the pre-trained model is set as a text detection model. If there is a position difference greater than the preset position difference threshold, it indicates that the prediction accuracy of the pre-trained model needs to be further optimized. Then, based on the predicted position information of the handwritten text region in each first sample image and the position difference between the position information of the handwritten text region in each first sample image, the model parameters of the pre-trained model are iteratively adjusted. When the number of adjustments is greater than or equal to a preset number of adjustments threshold, or when the position difference is less than or equal to the preset position difference threshold, the model parameter adjustment is stopped, and the current pre-trained model is set as a text detection model.
[0118] In some embodiments of the present invention, after obtaining the text detection model, the image to be detected is input into the text detection model, and the text detection model performs text detection on the input image to be detected to determine the text regions in the image to be detected. Text type detection is performed on each text region to determine the text type of the text in each text region. If there is a target text region in the text region whose text type is handwritten text, the target text region is marked to obtain the marked image to be detected.
[0119] In some embodiments of the present invention, the text detection model extracts features from the input image to be detected to obtain a feature map of the image to be detected, performs feature calculation on the feature map to obtain a probability matrix of the image to be detected, performs binarization processing on the probability matrix to obtain a binarization matrix of the image to be detected, selects target pixels with preset values in the binarization matrix of the image to be detected based on the binarization matrix, determines the text regions in the image to be detected based on the target pixels, and extracts the position information of each text region.
[0120] In some embodiments, the probability matrix includes all pixels in the image to be detected and the probability of each pixel in the text region.
[0121] In some embodiments, binarization processing based on the probability matrix to obtain a binarized matrix of the image to be detected includes: setting the pixel values of pixels in the image to be detected whose probability values are greater than or equal to the preset probability threshold to a first preset value, and setting the pixel values of pixels in the image to be detected whose probability values are less than the preset probability threshold to a second preset value, thereby obtaining a binarized matrix of the image to be detected. The first and second preset values are different; the first preset value can be 1, 0, or 255, and the second preset value can be 0 or 1. For example, when the first preset value is 255, the second preset value can be 0.
[0122] In some embodiments of the present invention, if there is no target text region of handwritten text type in the text region, the image to be detected is output, a prompt message is output, and the prompt message is returned to the client 103. For example, the prompt message "Handwritten text not detected" can be output.
[0123] In some embodiments of the present invention, after marking the image to be detected, the target text region marked in the marked image to be detected is cropped to obtain an image region containing the handwritten text in the target text region, and the image region containing the handwritten text in the target text region is subjected to text recognition to obtain the text recognition result of the target text region.
[0124] In some embodiments, the handwritten text in the target text region can be recognized according to the character recognition method in step 205 above, and the character recognition result of the target text region can be obtained.
[0125] In other embodiments, image segmentation can be performed based on the target text region in the image to be detected to obtain the target image region where the target text region is located in the image to be detected; the target image region is input into a preset character recognition model to perform character recognition on the handwritten text in the target image region to obtain the character recognition result of the target text region.
[0126] Among them, the text recognition model can be a neural network-based text recognition model, such as YOLOV, Long Short-Term Memory (LSTM), Convolutional Neural Networks (CNN), De-Convolutional Networks (DN), Deep Neural Networks (DNN), Deep Convolutional Inverse Graphics Networks (DCIGN), Region-based Convolutional Networks (RCNN), Faster Region-based Convolutional Networks (Faster RCNN), and Bidirectional Encoder Representations from Transformers (BERT) models.
[0127] In some embodiments of the present invention, in order to ensure data security, when obtaining the text recognition result, it is determined whether there are pre-stored risk characters in the text recognition result. If there are risk characters, the handwritten text corresponding to the text recognition result with risk characters is desensitized to obtain the desensitized text recognition result, and the marked image to be detected and the desensitized text recognition result of the target text region in the marked image to be detected are output.
[0128] In some embodiments of the present invention, if risky characters exist, the handwritten text corresponding to the text recognition result containing risky characters is desensitized to obtain the desensitized text recognition result. Based on the position information of the target text region where the desensitized handwritten text is located, the target text region with the same position information in the marked image to be detected is desensitized to obtain the desensitized image to be detected. The desensitized image to be detected and the desensitized text recognition result of the target text region in the marked image to be detected are output.
[0129] In some embodiments, desensitization can be performed by pre-defined character replacement, masking, or other methods.
[0130] The handwritten text detection method provided in this invention can determine the text region in the image to be detected through text detection, and can effectively determine whether the text in the text region is a handwritten text line through text type detection, thereby improving the accuracy of text detection and providing a foundation for text recognition.
[0131] To better implement the handwritten font detection method provided in the embodiments of the present invention, based on the handwritten font detection method, the embodiments of the present invention provide a handwritten text detection device, such as... Figure 4 As shown, Figure 4 This is a schematic diagram of the handwritten text detection device provided in an embodiment of the present invention. The handwritten text detection device shown includes:
[0132] The acquisition module is used to acquire the image to be detected;
[0133] The text region detection module is used to perform text detection on the image to be detected and to determine the text regions in the image.
[0134] The text type detection module is used to perform text type detection on each text region and determine the text type of the text in each text region;
[0135] The marking module is used to mark the target text regions if there are target text regions of handwritten text type in each text region, so as to obtain the marked image to be detected;
[0136] The text detection module is used to perform character recognition on handwritten text in the target text area and obtain the character recognition results of the target text area;
[0137] The output module is used to output the labeled image to be detected and the text recognition results of the target text region in the labeled image to be detected.
[0138] In some embodiments of the present invention, the handwritten text detection device includes:
[0139] The model detection module is used to: perform text detection on the image to be detected according to the preset text detection model, determine the text regions in the image to be detected, perform text type detection on each text region, and determine the text type of the text in each text region; the text detection model is obtained by adjusting the model parameters of the pre-trained model based on the first sample data containing handwritten text regions.
[0140] In some embodiments of the present invention, the handwritten text detection device further includes:
[0141] The training module is used to acquire first sample data. The first sample data includes multiple first sample images, each of which includes handwritten text and the location information of the handwritten text region. The first sample data is input into a pre-trained model for text detection to obtain the predicted location information of the handwritten text region in each first sample image. Based on the predicted location information and the location information of the handwritten text region in each first sample image, the model parameters of the pre-trained model are adjusted to obtain the text detection model.
[0142] In some embodiments of the present invention, the handwritten text detection device further includes:
[0143] The pre-training module is used to acquire a handwritten single-character dataset, randomly combine handwritten single-character images from the dataset to obtain multiple single-line handwritten text images; select a preset number of target single-line handwritten text images from these images, place them on a preset canvas, and determine the handwritten text region of each target single-line handwritten text image on the canvas; set the canvas after determining the handwritten text region as the original sample image; place preset printed text in the remaining image regions of the original sample image excluding the handwritten text region to obtain second sample images; the second sample images include both printed and handwritten text regions; annotate the handwritten text regions in each second sample image to obtain second sample data; input the second sample data into the initial model to train the initial model, obtaining the pre-trained model.
[0144] In some embodiments of the present invention, the pre-training module is used for:
[0145] The second sample image from the second sample data is input into the initial model for feature extraction, resulting in an approximate binarized feature map of each second sample image;
[0146] Contour recognition is performed on the approximate binarized feature maps of each second sample image to obtain the training location information of the handwritten text region in each second sample image;
[0147] The first training loss of the initial model is obtained based on the true binarized map and the approximate binarized feature map of each second sample image; the true binarized map of the second sample image is obtained by binarizing the handwritten text region in the second sample image.
[0148] The second training loss of the initial model is obtained by the cross-entropy between the training position information of the handwritten text region in each second sample image and the real position information of the handwritten text region in each second sample image; the real position information is the position information of the image region where the target single-line handwritten text image is located in the second sample image.
[0149] The target training loss of the initial model is obtained based on the first training loss and the second training loss. The initial model is then iteratively trained based on the target training loss. When the initial model meets the preset model convergence condition, the iterative training stops, and the pre-trained model is obtained.
[0150] In some embodiments of the present invention, the pre-training module is used for:
[0151] The second sample image from the second sample data is input into the initial model to extract features at different scales, resulting in feature maps of each second sample image at different scales.
[0152] The feature maps of each second sample image at different scales are combined to obtain the combined feature map of each second sample image.
[0153] Image convolution is performed on the combined feature maps of each second sample image to obtain the probability feature maps of each second sample image. Upsampling is then performed on the combined feature maps of each second sample image to obtain the threshold feature maps of each second sample image.
[0154] Differentiable binarization is performed on the difference image between the probability feature map and the threshold feature map of each second sample image to obtain the approximate binarized feature map of each second sample image.
[0155] In some embodiments of the present invention, the pre-training module is used for:
[0156] The binarization loss is obtained based on the cross-entropy between the true binarized image and the approximate binarized feature image of each second sample image.
[0157] Based on the threshold feature map of each second sample image, the predicted handwritten text region of each second sample image is determined, and the threshold loss is obtained based on the distance between the predicted handwritten text region of each second sample image and the handwritten text region of each second sample image.
[0158] The first training loss of the initial model is determined based on the binarization loss and the threshold loss.
[0159] In some embodiments of the present invention, the model detection module is used for:
[0160] The image to be detected is input into a preset text detection model to obtain the probability matrix of the image to be detected;
[0161] Binarize the probability matrix of the image to be detected to obtain the binarized matrix of the image to be detected.
[0162] Select target pixels with preset values in the binarization matrix of the image to be detected, and determine the text region in the image to be detected based on the selected target pixels.
[0163] In some embodiments of the present invention, the output module is used for:
[0164] Image segmentation is performed based on the target text region in the image to be detected to obtain the target image region where the target text region is located in the image to be detected.
[0165] The target image region is input into a preset text recognition model to perform text recognition on the handwritten text in the target image region, and the text recognition result of the target text region is obtained.
[0166] The handwritten text detection device provided in this embodiment of the invention can determine the text region in the image to be detected through text detection, and can effectively determine whether the text in the text region is a handwritten text line through text type detection, thereby improving the accuracy of text detection and providing a foundation for text recognition.
[0167] This invention also provides a server, such as... Figure 5 As shown, it illustrates a schematic diagram of the server structure involved in an embodiment of the present invention, specifically:
[0168] The server may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 5 The server structure shown does not constitute a limitation on the server and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0169] in:
[0170] The processor 401 is the control center of the server, connecting various parts of the server through various interfaces and lines. It performs various server functions and processes data by running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, thereby providing overall monitoring of the server. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.
[0171] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the server, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0172] The server also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0173] The server may also include an input unit 404, which can be used to receive input numeric or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0174] Although not shown, the server may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the server loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:
[0175] Acquire the image to be detected;
[0176] Perform text detection on the image to be detected to identify the text regions in the image;
[0177] Perform text type detection on the text region to obtain the text type of the text in the text region;
[0178] If a target text region of handwritten text exists in the text region, the target text region is marked to obtain the marked image to be detected;
[0179] Perform character recognition on the handwritten text in the target text region and obtain the character recognition results of the target text region;
[0180] Output the labeled image to be detected and the text recognition results of the target text region in the labeled image to be detected.
[0181] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0182] To this end, embodiments of the present invention provide a storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the handwritten text detection methods provided in the embodiments of the present invention. For example, the instructions can execute the following steps:
[0183] Acquire the image to be detected;
[0184] Perform text detection on the image to be detected to identify the text regions in the image;
[0185] Perform text type detection on the text region to obtain the text type of the text in the text region;
[0186] If a target text region of handwritten text exists in the text region, the target text region is marked to obtain the marked image to be detected;
[0187] Perform character recognition on the handwritten text in the target text region and obtain the character recognition results of the target text region;
[0188] Output the labeled image to be detected and the text recognition results of the target text region in the labeled image to be detected.
[0189] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0190] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0191] Since the instructions stored in the storage medium can execute the steps in any of the handwritten text detection methods provided in the embodiments of the present invention, the beneficial effects that any of the handwritten text detection methods provided in the embodiments of the present invention can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0192] The present invention provides a detailed description of a handwritten text detection method, apparatus, server, and storage medium. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for detecting handwritten text, characterized in that, The method includes: Acquire the image to be detected; Text detection is performed on the image to be detected to determine the text regions in the image to be detected; Perform text type detection on the text region to obtain the text type of the text in the text region; If there is a target text region in the text region that is handwritten text, then the target text region is marked to obtain the marked image to be detected; Perform character recognition on the handwritten text in the target text region to obtain the character recognition result of the target text region; Output the labeled image to be detected and the text recognition results of the target text region in the labeled image to be detected; The step of performing text detection on the image to be detected, determining the text regions in the image to be detected, and performing text type detection on the text regions to obtain the text type of the text in the text regions includes: The text detection model is used to detect text in the image to be detected according to a preset text detection model, thereby identifying text regions in the image to be detected. The text type of the text in the text regions is then detected to determine the text type of the text in the text regions. The text detection model is obtained by adjusting the model parameters of a pre-trained model based on first sample data containing handwritten text regions. Before performing text detection on the image to be detected according to the preset text detection model, first sample data is acquired; the first sample data includes multiple first sample images, each first sample image including handwritten text and the location information of the handwritten text region where the handwritten text is located; the first sample data is input into a pre-trained model for text detection to obtain the predicted location information of the handwritten text region in each first sample image; based on the predicted location information of the handwritten text region in each first sample image and the location information of the handwritten text region in each first sample image, the model parameters of the pre-trained model are adjusted to obtain the text detection model; Before inputting the first sample data into the pre-trained model for text detection, a handwritten single-character dataset is obtained. The handwritten single-character images in the dataset are randomly combined to obtain multiple single-line handwritten text images. A predetermined number of target single-line handwritten text images are selected from these images. The selected target single-line handwritten text images are placed in a predetermined canvas, and the handwritten text region of the target single-line handwritten text image in the canvas is determined. The canvas after determining the handwritten text region is set as the original sample image. Preset printed text is placed in the remaining image regions of the original sample image, excluding the handwritten text region, to obtain a second sample image. The handwritten text regions in each of the second sample images are labeled to obtain second sample data. The second sample data is input into the initial model to train the initial model, thus obtaining a pre-trained model.
2. The handwritten text detection method as described in claim 1, characterized in that, The step of inputting the second sample data into the initial model and training the initial model to obtain a pre-trained model includes: The second sample image from the second sample data is input into the initial model for feature extraction to obtain an approximate binarized feature map of each second sample image; Contour recognition is performed on the approximate binarized feature maps of each of the second sample images to obtain the training location information of the handwritten text region in each of the second sample images; The first training loss of the initial model is obtained based on the true binarized map and the approximate binarized feature map of each second sample image; the true binarized map of the second sample image is obtained by binarizing the handwritten text region in the second sample image. The second training loss of the initial model is obtained based on the cross-entropy between the training position information of the handwritten text region in each of the second sample images and the real position information of the handwritten text region in each of the second sample images; the real position information is the position information of the image region where the target single-line handwritten text image is located in the second sample image. The target training loss of the initial model is obtained based on the first training loss and the second training loss. The initial model is then iteratively trained based on the target training loss. When the initial model meets the preset model convergence condition, the iterative training is stopped, and a pre-trained model is obtained.
3. The handwritten text detection method as described in claim 2, characterized in that, The step of inputting the second sample image from the second sample data into the initial model for feature extraction to obtain an approximate binarized feature map for each second sample image includes: The second sample image from the second sample data is input into the initial model to extract features at different scales, resulting in feature maps of each second sample image at different scales. The feature maps of each second sample image at different scales are combined to obtain the combined feature map of each second sample image. Image convolution is performed on the feature maps of the combined second sample images to obtain the probability feature maps of the second sample images. Upsampling is then performed on the feature maps of the combined second sample images to obtain the threshold feature maps of the second sample images. Differentiable binarization is performed on the difference image between the probability feature map and the threshold feature map of each second sample image to obtain the approximate binarized feature map of each second sample image.
4. The handwritten text detection method as described in claim 3, characterized in that, The step of obtaining the first training loss of the initial model based on the true binarized map and the approximate binarized feature map of each of the second sample images includes: The binarization loss is obtained based on the cross-entropy between the true binarized map of each second sample image and the approximate binarized feature map of each second sample image; Based on the threshold feature map of each second sample image, the predicted handwritten text region of each second sample image is determined, and the threshold loss is obtained based on the distance between the predicted handwritten text region of each second sample image and the handwritten text region of each second sample image. The first training loss of the initial model is determined based on the binarization loss and the threshold loss.
5. The handwritten text detection method as described in claim 1, characterized in that, The step of performing text detection on the image to be detected according to a preset text detection model to determine the text region in the image to be detected includes: The image to be detected is input into a preset text detection model to obtain the probability matrix of the image to be detected; Binarize the probability matrix of the image to be detected to obtain the binarized matrix of the image to be detected; Select target pixels with preset values in the binarization matrix of the image to be detected, and determine the text region in the image to be detected based on the position of the target pixels.
6. The handwritten text detection method according to any one of claims 1 to 5, characterized in that, The step of performing character recognition on the handwritten text in the target text region to obtain the character recognition result of the target text region includes: Image segmentation is performed on the target text region in the image to be detected to obtain the target image region where the target text region is located in the image to be detected. The target image region is input into a preset character recognition model to perform character recognition on the handwritten text in the target image region, thereby obtaining the character recognition result of the target text region.
7. A handwritten text detection device, characterized in that, The device includes: The acquisition module is used to acquire the image to be detected; The text region detection module is used to perform text detection on the image to be detected and determine the text regions in the image to be detected. The text type detection module is used to perform text type detection on each of the text regions and determine the text type of the text in each of the text regions; The marking module is used to mark the target text region if there is a target text region of handwritten text type in each of the text regions, so as to obtain the marked image to be detected; The text detection module is used to perform character recognition on the handwritten text in the target text region to obtain the character recognition result of the target text region; The output module is used to output the labeled image to be detected and the text recognition results of the target text region in the labeled image to be detected; The model detection module is used to perform text detection on the image to be detected according to a preset text detection model, determine the text region in the image to be detected, perform text type detection on the text region, and determine the text type of the text in the text region; the text detection model is obtained by adjusting the model parameters of a pre-trained model based on first sample data containing handwritten text regions; A training module is used to acquire first sample data; the first sample data includes multiple first sample images, each first sample image including handwritten text and the location information of the handwritten text region where the handwritten text is located; the first sample data is input into a pre-trained model for text detection to obtain the predicted location information of the handwritten text region in each first sample image; based on the predicted location information of the handwritten text region in each first sample image and the location information of the handwritten text region in each first sample image, the model parameters of the pre-trained model are adjusted to obtain a text detection model; The pre-training module is used to acquire a handwritten single-character dataset, randomly combine handwritten single-character images from the dataset to obtain multiple single-line handwritten text images; select a preset number of target single-line handwritten text images from the multiple single-line handwritten text images, place the selected target single-line handwritten text images in a preset canvas, and determine the handwritten text region of the target single-line handwritten text image in the canvas; set the canvas after determining the handwritten text region as the original sample image; place preset printed text in the remaining image regions of the original sample image except for the handwritten text region to obtain a second sample image; label the handwritten text region in each of the second sample images to obtain second sample data; input the second sample data into the initial model, train the initial model, and obtain a pre-trained model.
8. A server, characterized in that, It includes a memory and a processor; the memory stores an application program, and the processor runs the application program within the memory to perform the operations in the handwritten text detection method according to any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium stores a plurality of instructions adapted for loading by a processor to execute the steps of the handwritten text detection method according to any one of claims 1 to 6.