A method, device, and storage medium for detecting the integrity of text data
By performing instance segmentation and text detection on the target image, we can determine whether there are partially overlapping text boxes in the target image area, which solves the problem of insufficient universality and reliability of text data integrity detection in the prior art, and realizes flexible and efficient text data integrity detection.
Patent Information
- Application Number
- CN202210064587.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-19
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-01-19
AI Technical Summary
The prior art is difficult to flexibly detect the integrity of text data in complex and diverse scenarios, resulting in insufficient detection universality and reliability.
By performing instance segmentation processing on the target image, the target image area is determined, and text detection is performed on the area to determine whether there are partially overlapping original text box areas to determine the integrity of the text data.
It realizes flexible text data integrity detection in different complex scenarios, improving the universality and reliability of detection.
Smart Images

Figure CN114419652B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to a method, apparatus, and storage medium for detecting the integrity of text data. Background Art
[0002] With the development of computer technology, when people handle various affairs, the picture materials uploaded no longer need to be manually checked by relevant staff. Instead, technologies such as Optical Character Recognition (OCR) are used to extract the text in the pictures, and a pre-set template is used to determine which positions in the pictures should have text, so as to verify whether the text in the pictures is complete. Although the method of verifying text through a pre-set template can achieve the verification of the integrity of the text in the pictures, the pre-set template is specific and can only judge the integrity of the text in the pictures in a single scenario, and it is difficult to determine the integrity of the text in the pictures in complex and diverse scenarios. Therefore, how to flexibly detect the integrity of text data, so as to improve the generality and reliability of the integrity detection of text data, is a technical problem that needs to be solved urgently at present. Summary of the Invention
[0003] Embodiments of the present application provide a method, apparatus, and storage medium for detecting the integrity of text data, which can flexibly detect the integrity of text data, thereby improving the generality and reliability of the integrity detection of text data.
[0004] On the one hand, embodiments of the present application provide a method for detecting the integrity of text data, including:
[0005] Performing instance segmentation processing on a target image to determine a target image region of a target instance in the target image;
[0006] Performing text detection processing on the target image region to obtain at least one original text box, and determining an original text box region of each original text box in the target image;
[0007] If there is a target original text box region that partially overlaps with the target image region, it is determined that the text data in the target image region is incomplete.
[0008] On the one hand, embodiments of the present application provide an apparatus for detecting the integrity of text data. The apparatus for detecting the integrity of text data includes a processing unit, and the processing unit is used to execute the above method for detecting the integrity of text data.
[0009] On the one hand, embodiments of the present application provide an electronic device. The electronic device includes an input interface and an output interface, and further includes:
[0010] A processor, adapted to implement one or more instructions; and,
[0011] A computer storage medium storing one or more instructions, the one or more instructions being adapted to be loaded and executed by the processor to perform the integrity detection method of the above text data.
[0012] On the one hand, an embodiment of the present application provides a computer storage medium storing computer program instructions, which are used to perform the integrity detection method of the above text data when executed by a processor.
[0013] On the one hand, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, which are used to perform the integrity detection method of the above text data when executed by the processor.
[0014] In an embodiment of the present application, instance segmentation processing is performed on a target image to determine a target image region of a target instance in the target image; text detection processing is performed on the target image region to obtain at least one original text box, and an original text box region of each original text box in the target image is determined; if there is a target original text box region that partially overlaps the target image region, it is determined that the text data in the target image region is incomplete. In an embodiment of the present application, by determining whether the target image region and the original text box region of the original text box obtained by performing text detection processing on the target image region partially overlap, it is possible to determine whether the text data in the target image region is complete; at the same time, since no templates or other external conditions are used in the process of detecting the integrity of the text data, this method can flexibly detect the integrity of the text data, thereby improving the generality and reliability of the integrity detection of the text data. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a schematic structural diagram of an integrity detection system for text data provided by an embodiment of the present application;
[0017] Figure 2 It is a schematic flowchart of a method for detecting the integrity of text data provided by an embodiment of the present application;
[0018] Figure 3 It is a schematic diagram of a target image area provided by an embodiment of the present application;
[0019] Figure 4 It is another schematic diagram of a target image area provided by an embodiment of the present application;
[0020] Figure 5 It is a schematic diagram of partial overlap of areas provided by an embodiment of the present application;
[0021] Figure 6 It is a schematic flowchart of another method for detecting the integrity of text data provided by an embodiment of the present application;
[0022] Figure 7 It is yet another schematic diagram of a target image area provided by an embodiment of the present application;
[0023] Figure 8a It is a schematic diagram of a filled image provided by an embodiment of the present application;
[0024] Figure 8b It is another schematic diagram of partial overlap of areas provided by an embodiment of the present application;
[0025] Figure 8c It is a schematic diagram of factors causing incomplete text data provided by an embodiment of the present application;
[0026] Figure 9 It is yet another schematic flowchart of a method for detecting the integrity of text data provided by an embodiment of the present application;
[0027] Figure 10 It is a schematic diagram of an original text box that does not belong to the target image area provided by an embodiment of the present application;
[0028] Figure 11 It is a schematic structural diagram of a device for detecting the integrity of text data provided by an embodiment of the present application;
[0029] Figure 12 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.
[0031] In order to flexibly detect the integrity of text data and thus improve the generality of the detection method, the embodiments of the present application provide a scheme for detecting the integrity of text data. First, instance segmentation processing is performed on a target image to determine a target image region of a target instance in the target image; then, text detection processing is performed on the target image region to obtain at least one original text box, and the original text box region of each original text box in the target image is determined; finally, if there is a target original text box region that partially overlaps with the target image region, it is determined that the text data in the target image region is incomplete.
[0032] In one embodiment, the above scheme for detecting the integrity of text data can be executed by a terminal device. Among them, the terminal device can include any one or more of a smart phone, a tablet computer, a laptop computer, a desktop computer, an intelligent vehicle, and a smart wearable device; the terminal device can also be an independent server, a cloud server, a server cluster, or a distributed system, etc., which is not limited here. The terminal device can use the image captured by the user as the target image, or select an image from the image database in the terminal device as the target image, and then the terminal device performs instance segmentation processing on the target image to determine the target image region of the target instance in the target image; then the terminal device performs text detection processing on the target image region to obtain at least one original text box, and determines the original text box region of each original text box in the target image; finally, if there is a target original text box region that partially overlaps with the target image region, the terminal device determines that the text data in the target image region is incomplete.
[0033] Based on the above scheme for detecting the integrity of text data, the embodiments of the present application provide a system for detecting the integrity of text data. See Figure 1 , which is a schematic structural diagram of a system for detecting the integrity of text data provided by the embodiments of the present application. Figure 1The integrity detection system for the text data shown may include a terminal device 101 and a server 102. Among them, the terminal device 101 may include any one or more of a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart vehicle, and a smart wearable device. The server 102 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device 101 and the server 102 may be directly or indirectly communicatively connected by a wired or wireless communication method, and this application does not limit this here.
[0034] In one embodiment, a user of the terminal device 101 may use a captured image as a target image, or select an image from an image database in the terminal device as a target image, and then upload the target image to the server 102. The server 102 first performs instance segmentation processing on the target image to determine a target image region of the target instance in the target image, then performs text detection processing on the target image region to obtain at least one original text box, and determines an original text box region of each original text box in the target image; finally, the server 102 determines whether there is a target original text box region that partially overlaps with the target image region; if there is a target original text box region that partially overlaps with the target image region, the server 102 determines that the text data in the target image region is incomplete, and may send a prompt message about the incomplete text data to the terminal device 101 to prompt the user of the terminal device 101 to re-upload the target image; if there is no target original text box region that partially overlaps with the target image region, the server 102 determines that the text data in the target image region is complete, and may send a prompt message about successful upload to the terminal device 101 to prompt the user that the upload of the target image has been completed.
[0035] Based on the above integrity detection solution for text data and the integrity detection system for text data, an embodiment of this application provides a method for detecting the integrity of text data. Refer to Figure 2 , which is a schematic flowchart of a method for detecting the integrity of text data provided by an embodiment of this application. Figure 2 The integrity detection method for the text data shown may be executed by Figure 1 the server or the terminal device shown. Figure 2 The integrity detection method for the text data shown may include the following steps:
[0036] S201. Perform instance segmentation on the target image to determine the target image region of the target instance in the target image.
[0037] In an embodiment of the present application, the target image may be an image captured or an image selected from multiple images in an image database. Among them, the manner in which the terminal device or the server selects the image may be that the user selects one or more images from multiple images as the target image, or the terminal device or the server randomly selects one or more images from multiple images as the target image, which is not limited herein.
[0038] In addition, the target image region may be the image region of the object recognized in the target image or the image region of the object recognized as a preset type in the target image. Among them, the preset type may be set by the user or the system, which is not limited herein.
[0039] For example, insurance user A needs to claim insurance reimbursement. Therefore, user A takes a picture of the hospital invoice with a mobile phone to obtain image B. Due to the cluttered background, in addition to the invoice, some other bills are also captured in image B. Then, insurance user A uploads image B as the target image to the insurance reimbursement platform. At this time, the insurance reimbursement platform performs instance segmentation on image B. Since the insurance reimbursement platform is mainly for insurance reimbursement and needs to detect whether the text data in the invoice for reimbursement is complete, the insurance reimbursement platform will use the image region recognized as the invoice as the target image region, and the image regions of other bills not recognized as invoices will not be used as the target image regions.
[0040] In an embodiment of the present application, the manner of performing instance segmentation on the target image may be to perform instance segmentation on the target image through a trained instance segmentation model. Exemplarily, the instance segmentation model may be a deep learning model such as MaskR-CNN (Mask Region CNN), Fast R-CNN (Fast Region CNN, which has a faster detection speed than the Region Convolutional Neural Network), YOLACT (a real-time instance segmentation model published in ICCV in 2019, which mainly realizes instance segmentation through two parallel convolutional neural sub-networks), etc. Different models can be flexibly selected based on different requirements, which is not limited herein. Among them, the process of training the instance segmentation model is a commonly used technical means for those skilled in the art and will not be elaborated herein.
[0041] In a specific implementation, multiple target images can be used as samples to input into the instance segmentation model Mask R-CNN. Then, the instance segmentation model Mask R-CNN performs instance segmentation processing on the multiple target images. Finally, the accuracy of the instance segmentation model Mask R-CNN is evaluated through the instance segmentation results. Among them, each target image may contain 0, 1, or multiple invoices or expense lists. The finally obtained instance segmentation results are shown in Table 1 and Table 2. Table 1 is used to indicate the possible number of invoices in each target image, and Table 2 is used to indicate the possible number of expense lists in each target image.
[0042] Table 1
[0043]
[0044] Table 2
[0045]
[0046] Among them, precision refers to the proportion of samples classified as positive examples that are actually positive examples. For example, there are 100 target images classified as having 1 invoice, but only 80 of these 100 are truly target images with 1 invoice, and the other 20 target images have no or multiple invoices. Then the precision is 80%; recall refers to the proportion of samples predicted as positive examples and actually being positive examples among the total samples that are actually positive examples. For example, there are 90 target images classified as having 1 invoice, but only 85 of these 90 are truly target images with 1 invoice, and the number of target images actually having 1 invoice in the input samples is 100. Then the recall is 85%. The F1 score is a measure for classification problems and is the harmonic mean of precision and recall, with a maximum of 1 and a minimum of 0; the calculation formula for the F1 score can be: (2 * recall * precision) / (recall + precision).
[0047] In addition, macro-average refers to taking the average after adding the F1 scores of each category; weighted average is an improvement of the macro-average. During the calculation process, it considers the proportion of the number of samples in each category in the total samples, and then uses this proportion as the weight when adding the F1 scores of each category, and finally calculates the average. Specifically, both macro-average and weighted average are indicators for measuring the quality of classification results, which are common technical means for those skilled in the art and will not be elaborated here.
[0048] It can be found from the instance segmentation results shown in Table 1 and Table 2 that by using the instance segmentation model Mask R-CNN to determine the target image region of the invoice or expense list in the target image, the misjudgment rate of the target image region of the target image is very low, that is, the accuracy is relatively high, which is beneficial to the subsequent detection of the text data in the target image region, and thus can effectively enhance the robustness of the entire text data integrity detection method.
[0049] In one embodiment, the method for performing instance segmentation processing on the target image to determine the target image region of the target instance in the target image may be as follows: 1) Perform instance segmentation processing on the target image to determine the image regions of each instance in at least one instance in the target image; 2) Obtain the position information of each image region in the target image; 3) If it is determined based on the position information of each image region that each image region is located in the central region of the target image, then each image region is used as the target image region.
[0050] Specifically, generally, when a user takes a photo and uploads a bill such as an invoice or an expense list, the user will try to take the invoice or expense list to be verified in the middle position of the image. Therefore, it is possible to determine whether the image region obtained after performing instance segmentation processing on the target image belongs to the central region of the target image, so as to determine whether the image region is the image region of the text data that the user wants to verify.
[0051] For example, please refer to the appendix Figure 3 , Figure 3 shows a schematic diagram of a target image region. By performing instance segmentation processing on the target image 301, the image regions of two instances in the target image 301 can be determined, which are the image region 302 and the image region 303 respectively. Among them, the central region 304 of the target image 301 can be set as the region of 1 / 4 to 3 / 4 of the length of the target image 301 and 1 / 4 to 3 / 4 of the width. Then, the position information of the image region 302 and the image region 303 in the target image 301 can be obtained. By calculating the position information of the image region 302 and the position information of the image region 303 respectively, it can be determined that the center point of the image region 302 is the center point 306, and the center point of the image region 303 is the center point 305. By comparing the positions of the center point 306, the center point 305 and the central region 304, as Figure 3 shown, it can be determined that the center point 306 falls into the central region 304, while the center point 305 does not fall into the central region 304. Therefore, it can be determined that the image region 302 is the target image region.
[0052] Optionally, it is also possible to determine whether each image region is located in the central region of the target image by judging the size of the overlapping region between each image region and the central region. If the region size of the overlapping region is greater than the preset region size, it is determined that the image region is located in the central region; if it is less than, it is determined that the image region is not located in the central region.
[0053] Optionally, it is also possible to obtain the position of the center point of the target image and the center point positions of each image region; then compare the distance between the center point of the target image and the center points of each image region. If the distance between the center point of any image region and the center point of the target image is less than the third preset distance threshold, it is determined that the any image region is the target image region. Optionally, the third preset distance threshold can be a length, the number of pixels, etc., which is not limited herein. Exemplarily, the third preset distance threshold can be 0.8 mm, 0.9 - 1.1 cm, 2 pixel widths, or 2 - 4 pixels, etc., which is not limited herein. Optionally, it is also possible to determine whether each image region is located in the central region of the target image by other means, which is not limited herein.
[0054] In one embodiment, the method for performing instance segmentation on the target image to determine the target image region of the target instance in the target image can be as follows: 1) Perform instance segmentation on the target image to determine the image regions of each instance in at least one instance in the target image; 2) Determine the first image region in the determined at least one image region, where the region size of the first image region is greater than the region sizes of each second image region, and the second image region is the image region other than the first image region in the at least one image region; 3) Obtain the ratio of the region size of each second image region to the region size of the first image region; 4) Use the second image regions whose ratio is greater than the second preset ratio threshold and the first image region as the target image region. Optionally, the second preset ratio threshold can be a numerical value such as a fraction, a percentage, etc., such as 1 / 2, 80%, 50% - 80%, etc.
[0055] For example, please refer to the appendix Figure 4 , Figure 4A schematic diagram of another target image region is shown. The second preset ratio threshold is set to 1 / 2. The target image 401 is input into the instance segmentation model Mask R-CNN for instance segmentation processing, and the image regions (i.e., MASK regions) of four instances in the target image can be determined as image region 402, image region 403, image region 404, and image region 405 respectively. By calculating the regional areas of image region 402, image region 403, image region 404, and image region 405 in the target image 401, it can be determined that the regional area of image region 402 is 640,000 pixels, the regional area of image region 403 is 450,000 pixels, the regional area of image region 404 is 80,000 pixels, and the regional area of image region 405 is 38,000 pixels; thus, it can be determined that image region 402 is the first image region, and image region 403, image region 404, and image region 405 are the second image regions. Since the ratio of image region 403 to image region 402 is greater than 1 / 2, while the ratios of image region 404 and image region 405 to image region 402 are both less than 1 / 2, it can be determined that the target image region is image region 402 and image region 403.
[0056] Optionally, in addition to the regional area, the regional size may also be data such as the length of the region, the width of the region, the perimeter of the region, etc. for measuring the regional size of the image region, which is not limited herein. Exemplarily, the regional size of image region C may be the length of image region C, i.e., 1000 pixels.
[0057] In one embodiment, the method for performing instance segmentation processing on the target image to determine the target image region of the target instance in the target image may also be: 1) performing instance segmentation processing on the target image to determine the image regions of each instance in at least one instance in the target image; 2) determining a first image region among the determined at least one image region, where the regional size of the first image region is greater than the regional sizes of each second image region, and the second image region is the image region other than the first image region among the at least one image region; 3) obtaining the ratio of the regional size of each second image region to the regional size of the first image region; 4) taking the second image region whose ratio is greater than the second preset ratio threshold and the first image region as the preselected target image regions; 5) obtaining the position information of each preselected target image region in the target image; 6) if it is determined based on the position information of each preselected target image region that each preselected target image region is located in the central region of the target image, then taking each preselected target image region as the target image region.
[0058] Optionally, the image region can also be preselected by the region size first, and then the target image region can be determined by the central region; in addition, the region size method and the central region method can also be used simultaneously to determine the target image region. When the image region is determined as the target image region by both methods, the image region is determined as the target image region, which is not limited here.
[0059] S202. Perform text detection processing on the target image region to obtain at least one original text box, and determine the original text box region of each original text box in the target image.
[0060] In the embodiment of the present application, the method for performing text detection processing on the target image region can be to perform text detection processing on the target image through a trained text detection model. Exemplarily, the text detection model can be a deep learning model such as DBNet (a differentiable binary segmentation network for text detection), CTPN (a network for text detection that includes a convolutional neural network and a recurrent neural network), SegLink (a convolutional neural network that can detect text with a rotation angle), etc. Different models can be flexibly selected based on different requirements, which is not limited here. The process of training the text detection model is a commonly used technical means for those skilled in the art and will not be elaborated here.
[0061] In the embodiment of the present application, in order to facilitate subsequent judgment of partial overlap between the target image region and the original text box region, the size of the text box region of the original text box obtained after the text detection processing is greater than or slightly greater than the position of the text data corresponding to the original text box in the target image region.
[0062] S203. If there is a target original text box region that partially overlaps with the target image region, it is determined that the text data in the target image region is incomplete.
[0063] In the embodiment of the present application, the method for determining whether the target image region and the original text box region partially overlap can be: determining the region position of the target image region in the target image and the position information of the original text box region in the target image; through position analysis processing of the position information, determining the positions of the four corner points of the original text box region in the target image. If the positions of some of the four corner points fall within the region position of the target image region while the positions of some corner points do not fall within the region position of the target image region, it is determined that the target image region and the original text box region partially overlap. Among them, the corner points are used to indicate the intersection points obtained by the intersection of the two side boundaries of the original text box.
[0064] For example, please refer to the appendix Figure 5 , Figure 5A schematic diagram showing partial overlap of regions is presented. After instance segmentation processing of the target image 501, the image regions 502 and 503 are obtained, and the image region 502 is determined as the target image region. Then, after text detection processing of the target image 501, 9 original text boxes as shown in Figure 5 are obtained, as well as the original text box regions of each original text box; at the same time, the positions of the four corner points of each original text box region in the target image are obtained. Among them, the position of the corner point 505 of the original text box 504 does not fall into the target image region 502, while the positions of the other corner points of the original text box 504 all fall into the target image region 502; therefore, it can be determined that the original text box 504 partially overlaps with the target image region, and the original text box 504 is the target original text box.
[0065] Optionally, it is also possible to determine whether there is an original text box with a part of the region falling into the target image region and another part not falling into the target image region through the position information of the original text box and the position information of the target image region, and thus whether there is an original text box region that partially overlaps with the target image region. Optionally, it is also possible to determine whether there is an original text box region that partially overlaps with the target image region by other means, which is not limited here.
[0066] In one embodiment, it can also be that after instance segmentation processing of the target image to determine the image regions of at least one instance in the target image, text detection processing is performed on at least one image region, and it is determined whether there is a text box region that partially overlaps with the image region. If so, it is determined that the text data in the image region is incomplete; then, it is determined whether the image region belongs to the target image region by the method of region size or central region mentioned in step S201. If the image region does not belong to the target image region, it is determined that the text data in the target image is complete; if the image region belongs to the target image region, it is determined that the text data in the target image is incomplete.
[0067] In the embodiments of the present application, instance segmentation processing is performed on a target image to determine a target image region of a target instance in the target image; text detection processing is performed on the target image region to obtain at least one original text box, and an original text box region of each original text box in the target image is determined; if there is a target original text box region that partially overlaps with the target image region, it is determined that the text data in the target image region is incomplete. In the embodiments of the present application, by determining whether the target image region and the original text box region of the original text box obtained by performing text detection processing on the target image region partially overlap, it is possible to determine whether the text data in the target image region is complete; in addition, since no template is used or other external conditions are relied on during the integrity detection process, this method can be applied to the integrity detection of text data in different complex scenarios, has strong versatility, and thus can flexibly detect the integrity of text data, thereby improving the versatility and reliability of the integrity detection of text data.
[0068] Based on the above text data integrity detection solution and the text data integrity detection system, the embodiments of the present application provide another text data integrity detection method. Refer to Figure 6 , which is a schematic flowchart of another text data integrity detection method provided by the embodiments of the present application. Figure 6 The text data integrity detection method shown can be executed by Figure 1 the server or terminal device shown. Figure 6 The text data integrity detection method shown may include the following steps:
[0069] S601, perform instance segmentation processing on the target image to determine a target image region of the target instance in the target image.
[0070] In the embodiments of the present application, the method of performing instance segmentation processing on the target image to determine a target image region of the target instance in the target image may be: perform instance segmentation processing on the target image to determine the image regions of each instance in at least one instance in the target image; obtain the distances between the respective region boundaries of any one of the determined at least one image region and the corresponding image boundaries in the target image; if there is at least one region boundary with a distance less than a second preset distance threshold, then use any one of the image regions as the target image region.
[0071] Optionally, the second preset distance threshold may be the number of pixels, or other parameters for measuring the distance size in the image, etc., and is not limited herein. Exemplarily, the second preset distance threshold may be 0.8 mm, 0.9 - 1.1 cm, or may also be 2 pixels, or 2 - 4 pixels, and is not limited herein.
[0072] Optionally, the second preset distance threshold may be determined based on the number of pixel points filled during the mirror padding of the image edge of the target image in step 602. For example, if 3 pixels are filled during the mirror padding of the image edge of the target image, then the second preset distance threshold may be 3 pixels. If the width of each pixel is 0.109 mm, then the second preset distance threshold may also be expressed as 0.327 mm.
[0073] Specifically, it is possible to determine the possibility that important information such as invoices and expense lists corresponding to each image region is not fully captured by judging the distance between the region boundaries of each image region and the corresponding image boundary in the target image. When the distance between the region boundary of the image region and its corresponding image boundary is small, it indicates that the possibility that the image region is not fully captured is relatively large. Therefore, it is necessary to determine this image region as the target image region, and then further judge whether the text data in this image region is complete.
[0074] For example, please refer to Figure 7 , Figure 7 FIG. shows a schematic diagram of another target image region. Set the second preset distance threshold to 20 pixels. After performing instance segmentation processing on the target image 701, the image regions 702 and 703 are obtained. The distances between the respective region boundaries of the image region 702 and the corresponding image boundaries in the target image are distance a, distance b, distance c, and distance d respectively; the distances between the respective region boundaries of the image region 703 and the corresponding image boundaries in the target image are distance e, distance f, distance g, and distance h respectively. Since the distances a, b, c, and d of the image region 702 are all greater than 20 pixels, while the distance h of the image region 703 is less than 20 pixels, it can be determined that the image region 703 is the target image region.
[0075] In one embodiment, the method for performing instance segmentation processing on the target image to determine the target image region of the target instance in the target image may also be: performing instance segmentation processing on the target image to determine the image regions of each instance in at least one instance in the target image; determining the region sizes of each image region, and if there is an image region whose region size is greater than the preset size threshold, then determine this image region as the target image region.
[0076] Optionally, the region size may be the area of the image region, or the length, width, etc. of the image region, which are data for measuring the size. Among them, the area, length, width, etc. of the image region may be represented by the number of pixels, or by length units such as millimeters and centimeters, which is not limited here. Optionally, the preset size threshold may be the area, or the length, width, etc. of the image region for measuring the image size, such as the area of the region is 30,000 pixels, the length of the image region is 500 pixels, and the width of the image region is 3,000 - 10,000 pixels, etc.
[0077] Exemplarily, after performing instance segmentation processing on the target image Y, image regions y1, y2, and y3 are obtained. Then, by calculating the area of each image region based on the position information of each image region in the target image, it can be determined that the area of image region y1 is 180,000 pixels, the area of image region y2 is 120,000 pixels, and the area of image region y3 is 30,000 pixels. Since the preset size threshold is 100,000 pixels, it can be determined that the target image regions are image region y1 and image region y2.
[0078] In one embodiment, the method for performing instance segmentation processing on the target image and determining the target image region of the target instance in the target image may also be: first, determining the preselected target image region by whether the distance between the region boundary of the image region and its corresponding image boundary is greater than the second preset distance threshold, and then determining the final target image region by whether the region size of the preselected image region is greater than the preset size threshold; or first, determining the preselected target image region by whether the region size of the preselected image region is greater than the preset size threshold, and then determining the final target image region by whether the distance between the region boundary of the image region and its corresponding image boundary is greater than the second preset distance threshold.
[0079] Optionally, one or more of the various methods for performing instance segmentation processing on the target image and determining the target image region of the target instance in the target image mentioned in step S301 and step S201 may also be selected to determine the target image region of the target instance in the target image, which is not limited here.
[0080] Specifically, during the process of determining the target image region, some unnecessary or unimportant image regions obtained after instance segmentation processing can be screened out, which is beneficial to improving the efficiency of the integrity detection of text data and also beneficial to improving the accuracy of the subsequent integrity detection of text data.
[0081] S602, perform mirror padding processing on the image edge of the target image to obtain a padded image.
[0082] S603. Determine the filled image area corresponding to the target image area in the filled image.
[0083] In the embodiments of the present application, the mirror filling process for the image edge of the target image refers to taking the pixel values of the pixel points symmetric to the four image boundaries of the upper, lower, left, and right of the target image, as well as the pixel values of the pixel points at the four vertices of the target image, and performing corresponding filling on the target image to obtain the filled image. The method for determining the filled image area corresponding to the target image area in the filled image can be: performing instance segmentation processing on the filled image to determine the filled image area corresponding to the target instance in the filled image. When the filled image area and the target instance corresponding to the target image area are the same, it is determined that this filled image area is the filled image area corresponding to the target image area.
[0084] For example, please refer to the attached Figure 8a , which shows a schematic diagram of a filled image. After performing instance segmentation processing on the target image 801, it can be determined that the image areas of the 3 instances in the target image include image area 802, image area 803, and image area 804. Among them, the target image areas are image area 802 and image area 803. Then, perform mirror filling processing on the image edge of the target image 801 to obtain the filled image 805; then perform instance segmentation processing on the filled image 805, and it can be determined that the filled image area corresponding to image area 802 in the filled image 805 is image area 806, the filled image area corresponding to image area 803 is image area 807, and the filled image area corresponding to image area 804 is image area 808.
[0085] S604. Perform text detection processing on the filled image area to obtain at least one filled text box and the filled text box area of each filled text box in the filled image.
[0086] In the embodiments of the present application, the method for performing text detection processing on the filled image area to obtain the filled text box and the filled text box area can refer to the method for performing text detection processing on the target image area in step S202 to obtain the original text box and the original text box area, which will not be elaborated here.
[0087] S605. If there is a target filled text box area that partially overlaps with the target image area, it is determined that the text data in the target image area is incomplete.
[0088] In an embodiment of the present application, the method for determining whether the target filled text box area partially overlaps with the target image area may specifically be: obtaining first position information of the target image area in the filled image, and second position information of each filled text box area in the filled image; and determining whether each filled text box area partially overlaps with the target image area according to the first position information and the second position information of each filled text box area.
[0089] Exemplarily, please refer to the attached Figure 8b , which shows another schematic diagram of partial area overlap. Text detection processing is performed on the filled image area 806 and the filled image area 807 corresponding to the target image in the filled image 805, and filled text box areas such as the filled text box area 809, the filled text box area 810, the filled text box area 811, and the filled text box area 812 are obtained. The positions of the filled image 805 and the target image 801 are overlapped, as shown in the image 813. The positions of the image areas 802 and 803 that can be determined as the target image area in the filled image 805 can be obtained. Then, through the filled text box areas of the filled text boxes such as the filled text box 809, the filled text box 810, the filled text box 811, and the filled text box 812 in the filled image, the target filled text box areas that partially overlap with the target image area can be determined as the filled text box areas of the filled text boxes 809, 810, 811, and 812 in the filled image.
[0090] In one embodiment, after determining that there is a target filled text box area that partially overlaps with the target image area, it is also possible to:
[0091] 1) Determine a target intersection point among the intersection points where the area boundary of the target filled text box area intersects with the area boundary of the target image area;
[0092] 2) Obtain the distance between the target intersection point and the target image boundary of the filled image, where the target image boundary matches the text box sub-area in the target filled text box area that does not overlap with the target image area;
[0093] 3) If the distance is less than the first preset distance threshold, determine that the factor causing the text data of the target image area to be incomplete is the first factor; if the distance is greater than or equal to the first preset distance threshold, determine that the factor causing the text data of the target image area to be incomplete is the second factor. Optionally, the first preset distance threshold may be a length, or the number of pixels, etc., which is not limited herein. Exemplarily, the first preset distance threshold may be 0.8 mm, 0.9 - 1.1 cm, or 2 pixels, or 2 - 4 pixels, etc., which is not limited herein.
[0094] Optionally, the matching between the target image boundary and the text box sub-region in the target filling text box region that does not overlap with the target image region means that the distance between the first filling text box sub-region in the target filling text box region and the target image boundary is less than the distance between the second filling text box sub-region in the target filling text box region and the target image boundary. The first filling text box sub-region refers to the text sub-region in the target filling text box region that does not overlap with the target image region, and the second filling text box sub-region refers to the text sub-region in the target filling text box that overlaps with the target image region.
[0095] Optionally, the first factor refers to the missing text data in the text data of the target image region due to reasons such as not fully capturing the target instance during the shooting process; the second factor refers to the occlusion of the text data in the target image region caused by other image regions or by the folding of the target instance corresponding to the target image region itself.
[0096] Optionally, the target intersection point can be the intersection point closest to the target image boundary. Optionally, the target intersection point does not need to be determined. As long as there is an intersection point among the intersection points where the region boundary of the target filling text box region intersects with the region boundary of the target image region and the distance between this intersection point and the target image boundary of the filled image is less than the first preset distance threshold, then it is determined that the factor causing the text data in the target image region to be incomplete is the first factor; if not, the factor causing the text data in the target image region to be incomplete is the second factor.
[0097] For example, please refer to the appendix Figure 8c , which shows a schematic diagram of the factors for incomplete text data. Set the first preset distance threshold to 15 pixels. For the image region 802 as the target image region, as shown in image 814, it can be determined that the intersection points where the region boundary of the filling text box region of the filling text box 812 intersects with the region boundary of the image region 802 include intersection point 814 and intersection point 815. At the same time, it can also be determined that the image boundary of the filled image that matches the text box sub-region in the filling text box region of the filling text box 812 and does not overlap with the image region 802 is image boundary M1. Then, the distance A between intersection point 814 and image boundary M1 is less than the distance B between intersection point 815 and image boundary M1. Therefore, it can be determined that intersection point 814 is the target intersection point; further, through the position information of the filling text box region of the filling text box 812 and the position information of the image region 802 in the filled image, it is determined that the distance A is 38 pixels, which is greater than 15 pixels. Therefore, finally, it can be determined that the factor for the incomplete text data in the image region 802 is the second factor, that is, the text data is occluded. The processing steps for the filling text box 811 in the image region 802 are the same as those of the filling text box 812 above, and thus it can also be determined that the factor for the incomplete text data in the image region 802 is the second factor, which will not be elaborated here.
[0098] For the image region 803 as the target image region, as Figure 8c shown, it can be determined that the intersection points where the region boundary of the filled text box region of the filled text box 809 intersects the region boundary of the image region 803 include the intersection point 816 and the intersection point 817. At the same time, it can also be determined that the image boundary of the filled image that matches the text box sub-region in the filled text box region of the filled text box 809 and not overlapping with the image region 803 is the image boundary M2. Then, based on the position information of the filled text box region of the filled text box 809 and the position information of the image region 803 in the filled image, it can be determined that the distance C between the intersection point 816 and the image boundary M2 is the same as the distance D between the intersection point 817 and the image boundary M2, both being 10 pixels. Therefore, both the intersection point 816 and the intersection point 817 can be used as target intersection points. Since 10 pixels is less than 15 pixels, ultimately, it can be determined that the factor for the incomplete text data in the image region 803 is the first factor, that is, the missing text data caused by the text data not being fully captured. The processing steps for the filled text box 810 in the image region 803 are the same as those for the filled text box 809 above, and thus it can also be determined that the factor for the incomplete text data in the image region 802 is the first factor, which will not be elaborated here.
[0099] In specific implementation, after performing the above steps S601 to S605 with multiple target images as samples through the integrity detection algorithm of text data, the integrity detection results as shown in Table 3 can be obtained. From Table 3, it can be found that among the 3 types of detection results, the misjudgment rate of the overall integrity detection result is very low. Therefore, this method for the integrity detection of text data has a high accuracy rate and strong generalization ability during the detection process. Thus, this method is suitable for wide and general use.
[0100] Table 3
[0101]
[0102] In the embodiments of the present application, first, instance segmentation processing is performed on a target image to determine a target image region of a target instance in the target image; then, mirror padding processing is performed on the image edge of the target image to obtain a padded image, and a padded image region corresponding to the target image region is determined in the padded image; then, text detection processing is performed on the padded image region to obtain at least one padded text box and a padded text box region of each padded text box in the padded image; finally, if there is a target padded text box region that partially overlaps with the target image region, it is determined that the text data in the target image region is incomplete. In the embodiments of the present application, mirror padding processing is performed on the image edge of the target image to obtain a padded image; then, it is determined whether the padded text box regions of the respective padded text boxes obtained in the padded image region corresponding to the target image region in the padded image partially overlap with the target image region, which can effectively determine whether the text data in some target image regions close to the image edge of the target image is complete. Since no templates or other external conditions are used in the integrity detection process, the integrity of the text data can be flexibly detected, thereby improving the generality and reliability of the integrity detection of the text data. In addition, it is also possible to determine the factor that causes the text data in the target image region to be incomplete by obtaining the distance between the target intersection point and the target image boundary of the padded image and determining whether the distance is greater than a first preset distance threshold, which can facilitate subsequent users to make corresponding modifications to the target image and is beneficial to improving the user experience.
[0103] Based on the above text data integrity detection scheme and text data integrity detection system, embodiments of the present application provide a method for detecting the integrity of text data. Refer to Figure 9 , which is a flowchart of a method for detecting the integrity of text data provided by an embodiment of the present application. Figure 9 The text data integrity detection method shown can be executed by Figure 1 the server and terminal device shown. Figure 9 The text data integrity detection method shown may include the following steps:
[0104] S901, the terminal device sends the target image to the server.
[0105] In the embodiments of the present application, the terminal device sending the target image to the server can be performed by wireless communication or wired communication, or can be sent after encryption, or can be sent by other means, which is not limited herein.
[0106] S902, the server performs instance segmentation processing on the target image to determine the target image region of the target instance in the target image.
[0107] S903. The server performs text detection processing on the target image region to obtain at least one original text box and determines the original text box region of each original text box in the target image.
[0108] Among them, for the specific implementation manners of steps S902 to S903, reference can be made to the specific implementation manners of steps S201 to S202, which will not be elaborated here.
[0109] S904. The server determines the original text box region that partially overlaps with the target image region as the target original text box region.
[0110] Among them, for the specific implementation manner of determining the target original text box region determined in step S904, reference can be made to the specific implementation manner of the method for determining whether the target image region and the original text box region partially overlap in step S203, which will not be elaborated here.
[0111] S905. The server determines the first original text box sub-region and the second original text box sub-region in the target original text box region.
[0112] In the embodiment of the present application, the first original text box sub-region refers to the text box sub-region of the target original text box region that does not overlap with the target image region, and the second original text box sub-region refers to the text box sub-region of the target original text box region that overlaps with the target image region.
[0113] In the embodiment of the present application, the determination of the first original text box sub-region and the second original text box sub-region in the target original text box region may be: obtaining the third position information of the target original text box region in the target image and the fourth position information of the target image region in the target image; then determining the first original text box sub-region where the target original text box region overlaps with the target image region and the second original text box sub-region where the target original text box region does not overlap with the target image region based on the third position information and the fourth position information. Specifically, a two-dimensional coordinate system can be established based on the target image, and the coordinates of the target original text box region and the coordinates of the target image region can be determined; then, through corresponding mathematical calculations on the coordinates of the target original text box region and the coordinates of the target image region, the coordinates of the first original text box sub-region and the coordinates of the second original text box sub-region are obtained, thereby completing the determination of the first original text box sub-region and the second original text box sub-region in the target original text box region.
[0114] S906. The server obtains the ratio of the region size of the first original text box sub-region to the region size of the second original text box sub-region.
[0115] S907. If the ratio is less than the first preset proportional threshold, the server determines that the text data of the target image area is incomplete.
[0116] In an embodiment of the present application, the area size may be any data for measuring the size of the area, such as area, length, width, perimeter, number of pixels, etc., which is not limited herein. The method for obtaining the ratio of the area size of the first original text box sub-area and the area size of the second original text box sub-area may be: obtaining the third position information of the target original text box area in the target image and the fourth position information of the target image area in the target image; then determining the area size of the first original text box sub-area and the area size of the second original text box sub-area in the target original text box area based on the third position information and the fourth position information, and then determining the ratio based on the area size of the first original text box sub-area and the area size of the second original text box sub-area.
[0117] In one embodiment, the method for obtaining the ratio of the area size of the first original text box sub-area and the area size of the second original text box sub-area may also be: obtaining the third position information of the target original text box area in the target image and the fourth position information of the target image area in the target image; then determining the first original text box sub-area and the second original text box sub-area of the target original text box area based on the third position information and the fourth position information; then determining the area size of the first original text box sub-area based on the size of the pixel points in the first original text box sub-area, and determining the area size of the second original text box sub-area based on the size of the pixel points in the second original text box sub-area.
[0118] In an embodiment of the present application, if the ratio is less than the first preset proportional threshold, it indicates that the target original text box belongs to the target image area, so it can be determined that the text data of the target image area is incomplete; if the ratio is greater than the first preset proportional threshold, it indicates that the target original text box does not belong to the target image area, and it may be that the text box of other image areas protrudes into the target image area, so it can be determined that the text data of the target image area is complete. Optionally, the first preset proportional threshold may be a decimal, a fraction, a percentage, etc., such as 0.7, 1 / 2, 80%, 50% - 80%, etc.
[0119] For example, please refer to the appendix Figure 10 , Figure 10A schematic diagram of an original text box that does not belong to the target image area is shown. Set the first preset ratio threshold to 0.4. After performing instance segmentation processing on the target image 1001, it can be determined that the image areas corresponding to 2 instances are image area 1002 and image area 1003 respectively. Among them, by means of the area size, it can be determined that image area 1002 is the target image area. Then, text detection processing is performed on image area 1002. Since the position of the original text box 1004 in image area 1003 is very close to the position of image area 1002, the original text box 1004 will be detected during the detection process and will be detected as a target original text box that partially overlaps with image area 1002. Therefore, after determining the target original text box of image area 1002, it is also necessary to determine whether the target original text box belongs to image area 1002. At this time, through the position information of the original text box 1004, the area of the first original text box sub-region 1006 of the original text box 1004 can be calculated to be 8000 pixels, and the area of the second original text box sub-region 1005 is 1000 pixels, so that the ratio can be determined to be 8. Since 8 is greater than 0.5, it can be determined that the original text box 1004 does not belong to image area 1002, and there are no other original text boxes in image area 1002 that partially overlap with image area 1002. Therefore, it can be determined that the text data of image area 1002 is complete.
[0120] In one embodiment, it is also possible to: 1) determine an original text box area that completely overlaps with the target image area in at least one original text box area; 2) obtain the angle difference between the inclination angle of the determined original text box area and the inclination angle of the target original text box area; 3) if the angle difference is less than the preset angle threshold, determine that the text data of the target image area is incomplete. Specifically, if the target original text box area belongs to the target image area, then the inclination angle of the target original text box area in the filled image should be approximately the same as the inclination angle of the original text box area that completely overlaps with the target image area, so as to determine whether the text data of the target image area is complete, effectively avoiding misjudgment and being beneficial to improving the accuracy of the integrity detection of the text data.
[0121] Exemplarily, a two-dimensional coordinate system can be established based on the target image, and a preset angle threshold is set to 5 degrees. Then, the first coordinate information of the original text box area M that completely overlaps with the target image area is obtained, as well as the second coordinate information of the target original text box area N that partially overlaps with the target image area; furthermore, the first tilt angle is calculated to be 10 degrees based on the first coordinate information, and the second tilt angle is calculated to be 68 degrees based on the second coordinate information. Therefore, the angle difference is 58 degrees, so it can be determined that the target original text box area N does not belong to the target image area, and further it can be determined that the text data of the target image area is complete.
[0122] In one embodiment, there may be multiple original text box areas that completely overlap with the target image area. At this time, the average tilt angle of the tilt angles of all the original text box areas that completely overlap with the target image area can be calculated, and then the angle difference is obtained by comparing the average tilt angle with the tilt angle of the target original text box area. Finally, the angle difference is compared with the preset angle threshold.
[0123] In one embodiment, when there are multiple original text box areas that completely overlap with the target image area, the angle difference between the tilt angle of the target original text box area and the tilt angle of each original text box area that completely overlaps with the target image area can also be determined. If each angle difference is less than the preset angle threshold, it is determined that the target original text box area belongs to the target image area.
[0124] Optionally, when there are multiple original text box areas that completely overlap with the target image area, the tilt angles can also be compared in other ways to determine whether the target original text box area belongs to the target image area, which is not limited herein.
[0125] S908, the server generates detection result information.
[0126] In the embodiments of the present application, the detection result information can be used to indicate whether the text data of the target image is complete; when the text data is incomplete, the detection result information can also indicate which target image area in the target image corresponds to the incomplete text data in the target instance, and what the factors for the incomplete text data are. Optionally, the detection result information can also include other information for determining whether the text data is complete, which is not limited herein.
[0127] Exemplarily, after performing integrity detection on the text data in the target image region M2 corresponding to the invoice and the target image region M3 corresponding to the expense list in the target image M1, it can be determined that the text data in the target image region M2 is incomplete, and the factor causing the incomplete text data is "not fully captured"; while the text data in the target image region M3 is complete. Therefore, the generated detection result information can be "The text in the target image M1 is incomplete, please re-upload", or "The text of the invoice in the target image M1 is incomplete, please re-upload", or "The text of the invoice in the target image M1 is not fully captured, please re-upload".
[0128] S909, the server sends the detection result information to the terminal device.
[0129] Among them, the manner of sending the detection result information to the terminal device in step S909 is a commonly used technical means by those skilled in the art and will not be elaborated here.
[0130] In the embodiment of the present application, first, the terminal device sends the target image to the server, and then the server performs instance segmentation processing on the target image to determine the target image region of the target instance in the target image; and the server performs text detection processing on the target image region to obtain at least one original text box and determines the original text box region of each original text box in the target image; then, the server determines the original text box region that partially overlaps with the target image region as the target original text box region, and determines the first original text box sub-region and the second original text box sub-region in the target original text box region; after that, the server obtains the ratio of the region size of the first original text box sub-region to the region size of the second original text box sub-region. If the ratio is less than the first preset ratio threshold, the server determines that the text data in the target image region is incomplete; finally, the server generates the detection result information and sends it to the terminal device. In the embodiment of the present application, by the ratio of the region size of the first original text box sub-region to the region size of the second original text box sub-region, it can be determined whether the target original text box region that partially overlaps with the target image region belongs to the target image region. For the target original text box region that does not belong to the target image region, it is not considered that the text data in the target image region is incomplete, effectively avoiding misjudgment and being beneficial to improving the accuracy of the integrity detection of the text data. In addition, since no template or other external conditions are used in the process of the integrity detection of the text data, it has strong versatility. Therefore, the integrity of the text data can be flexibly detected, thereby improving the versatility and reliability of the integrity detection of the text data.
[0131] Based on the above embodiment of the integrity detection method for text data, the embodiment of the present application provides an integrity detection device for text data. See Figure 11, which is a schematic structural diagram of an integrity detection device for text data provided by an embodiment of the present application. The integrity detection device for text data may include a processing unit 1101. Figure 11 The integrity detection device for text data shown may operate the following units:
[0132] The processing unit 1101 is configured to perform instance segmentation processing on a target image to determine a target image region of a target instance in the target image;
[0133] The processing unit 1101 is further configured to perform text detection processing on the target image region to obtain at least one original text box and determine an original text box region of each original text box in the target image;
[0134] The processing unit 1101 is further configured to determine that the text data in the target image region is incomplete if there is a target original text box region that partially overlaps with the target image region.
[0135] In one embodiment, the processing unit 1101 is further configured to perform mirror padding processing on the image edge of the target image to obtain a padded image; determine a padded image region corresponding to the target image region in the padded image; perform text detection processing on the padded image region to obtain at least one padded text box and a padded text box region of each padded text box in the padded image; determine that the text data in the target image region is incomplete if there is a target padded text box region that partially overlaps with the target image region.
[0136] In one embodiment, the processing unit 1101 is further configured to determine a target intersection among the intersections where the region boundary of the target padded text box region intersects with the region boundary of the target image region; obtain the distance between the target intersection and the target image boundary of the padded image, and the target image boundary matches the text box sub-region that does not overlap with the target image region in the target padded text box region; if the distance is less than a first preset distance threshold, determine that the factor causing the text data in the target image region to be incomplete is the first factor; if the distance is greater than or equal to the first preset distance threshold, determine that the factor causing the text data in the target image region to be incomplete is the second factor.
[0137] In one embodiment, the processing unit 1101 is further configured to perform instance segmentation processing on the target image to determine an image region of each instance in at least one instance in the target image; obtain the distance between each region boundary of any image region among the determined at least one image region and the corresponding image boundary in the target image; if there is at least one region boundary with a distance less than a second preset distance threshold, use any image region as the target image region.
[0138] In one embodiment, the processing unit 1101 is further configured to determine, in at least one original text box area, an original text box area that completely overlaps with the target image area; obtain an angle difference between the inclination angle of the determined original text box area and the inclination angle of the target original text box area; if the angle difference is less than a preset angle threshold, determine that the text data of the target image area is incomplete.
[0139] In one embodiment, the processing unit 1101 is further configured to determine a first original text box sub-area and a second original text box sub-area in the target original text box area, where the first original text box sub-area refers to the text box sub-area in the target original text box area that does not overlap with the target image area, and the second original text box sub-area refers to the text box sub-area in the target original text box area that overlaps with the target image area; obtain a ratio of the area size of the first original text box sub-area to the area size of the second original text box sub-area; if the ratio is less than a first preset ratio threshold, determine that the text data of the target image area is incomplete.
[0140] In one embodiment, the processing unit 1101 is further configured to perform instance segmentation processing on the target image to determine the image areas of each instance in at least one instance in the target image; obtain the position information of each image area in the target image; if it is determined, based on the position information of each image area, that each image area is located in the central area of the target image, then use each image area as the target image area.
[0141] In one embodiment, the processing unit 1101 is further configured to perform instance segmentation processing on the target image to determine the image areas of each instance in at least one instance in the target image; determine a first image area among the determined at least one image area, where the area size of the first image area is larger than the area sizes of each second image area, and the second image area is an image area other than the first image area among the at least one image area; obtain a ratio of the area size of each second image area to the area size of the first image area; use the second image areas with a ratio greater than a second preset ratio threshold and the first image area as the target image area.
[0142] According to an embodiment of the present application, Figure 2 、 Figure 6 and Figure 9 each step involved in the method for detecting the integrity of the text data shown can be performed by Figure 11 each unit in the device for detecting the integrity of the text data shown.
[0143] According to another embodiment of the present application, Figure 11Each unit in the integrity detection device for the text data shown can be separately or all combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units in terms of function to form, which can achieve the same operations without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In actual applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the integrity detection device for the text data divided based on logical functions can also include other units. In actual applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.
[0144] According to another embodiment of this application, it can be achieved by running a computer program (including program code) capable of executing the respective steps involved in the corresponding methods shown in Figure 2 , Figure 6 and Figure 9 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct the integrity detection device for the text data shown in Figure 11 , and to implement the integrity detection method for the text data in the embodiments of this application. The computer program can be recorded on, for example, a computer-readable storage medium, loaded into the above computing device through the computer-readable storage medium, and run therein.
[0145] In the embodiments of this application, instance segmentation processing is performed on the target image to determine the target image area of the target instance in the target image; text detection processing is performed on the target image area to obtain at least one original text box, and the original text box area of each original text box in the target image is determined; if there is a target original text box area that partially overlaps with the target image area, it is determined that the text data in the target image area is incomplete. In the embodiments of this application, by judging whether the target image area and the original text box area of the original text box obtained by performing text detection processing on the target image area partially overlap, it can be determined whether the text data in the target image area is complete; at the same time, since no templates or other external conditions are used during the integrity detection process of the text data, this method can flexibly detect the integrity of the text data, thereby improving the generality and reliability of the integrity detection of the text data.
[0146] Based on the above method embodiments and device embodiments, this application also provides an electronic device. Refer to Figure 12 , which is a schematic structural diagram of an electronic device provided in the embodiments of this application. Figure 12The electronic device shown may at least include a processor 1201, an input interface 1202, an output interface 1203, and a computer storage medium 1204. Among them, the processor 1201, the input interface 1202, the output interface 1203, and the computer storage medium 1204 may be connected by a bus or other means.
[0147] The computer storage medium 1204 can be stored in the memory of the electronic device. The computer storage medium 1204 is used to store a computer program, and the computer program includes program instructions. The processor 1201 is used to execute the program instructions stored in the computer storage medium 1204. The processor 1201 (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device, and is adapted to implement one or more instructions. Specifically, it is adapted to load and execute one or more instructions to implement the integrity detection method process or corresponding functions of the above text data.
[0148] The embodiment of the present application also provides a computer storage medium (Memory). The computer storage medium is a memory device in the electronic device and is used to store programs and data. It can be understood that the computer storage medium here can include both the built-in storage medium in the terminal and, of course, the extended storage medium supported by the terminal. The computer storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor 1201 are also stored in this storage space. These instructions can be one or more computer programs (including program code). It should be noted that the computer storage medium here can be a high-speed random access memory (RAM) memory, or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer storage medium located far from the aforementioned processor.
[0149] In one embodiment, one or more instructions stored in the computer storage medium can be loaded and executed by the processor 1201 to implement the above-mentioned Figure 2 , Figure 6 and Figure 9 corresponding steps of the method in the embodiment of the integrity detection method of text data. In a specific implementation, one or more instructions in the computer storage medium are loaded and executed by the processor 1201 as follows:
[0150] The processor 1201 performs instance segmentation processing on the target image to determine the target image region of the target instance in the target image;
[0151] The processor 1201 performs text detection processing on the target image region, obtains at least one original text box, and determines the original text box region of each original text box in the target image;
[0152] The processor 1201 determines that the text data of the target image region is incomplete if there is a target original text box region that partially overlaps the target image region.
[0153] In one embodiment, the processor 1201 performs mirror padding processing on the image edge of the target image to obtain a padded image;
[0154] The processor 1201 determines the padded image region corresponding to the target image region in the padded image;
[0155] The processor 1201 performs text detection processing on the padded image region, obtains at least one padded text box, and the padded text box region of each padded text box in the padded image;
[0156] The processor 1201 determines that the text data in the target image region is incomplete if there is a target padded text box region that partially overlaps the target image region.
[0157] In one embodiment, the processor 1201 determines a target intersection point among the intersection points where the region boundary of the target padded text box region intersects the region boundary of the target image region;
[0158] The processor 1201 obtains the distance between the target intersection point and the target image boundary of the padded image, and the target image boundary matches the text box sub-region in the target padded text box region that does not overlap with the target image region;
[0159] The processor 1201 determines that the factor causing the text data of the target image region to be incomplete is the first factor if the distance is less than the first preset distance threshold;
[0160] The processor 1201 determines that the factor causing the text data of the target image region to be incomplete is the second factor if the distance is greater than or equal to the first preset distance threshold.
[0161] In one embodiment, the processor 1201 performs instance segmentation processing on the target image to determine the target image region of the target instance in the target image, including:
[0162] The processor 1201 performs instance segmentation processing on the target image to determine the image region of each instance in at least one instance in the target image;
[0163] The processor 1201 obtains the distance between each region boundary of any one of the determined at least one image region and the corresponding image boundary in the target image;
[0164] The processor 1201 determines that if the distance of at least one region boundary is less than the second preset distance threshold, any image region is used as the target image region.
[0165] In one embodiment, the processor 1201 determines the original text box regions that completely overlap with the target image region in at least one original text box region;
[0166] The processor 1201 obtains the angle difference between the inclination angle of the determined original text box region and the inclination angle of the target original text box region;
[0167] If the angle difference is less than the preset angle threshold, the processor 1201 determines that the text data of the target image region is incomplete.
[0168] In one embodiment, the processor 1201 determines a first original text box sub-region and a second original text box sub-region in the target original text box region. The first original text box sub-region refers to the text box sub-region of the target original text box region that does not overlap with the target image region, and the second original text box sub-region refers to the text box sub-region of the target original text box region that overlaps with the target image region;
[0169] The processor 1201 obtains the ratio of the region size of the first original text box sub-region to the region size of the second original text box sub-region;
[0170] If the ratio is less than the first preset ratio threshold, the processor 1201 determines that the text data of the target image region is incomplete.
[0171] In one embodiment, the processor 1201 performs instance segmentation processing on the target image to determine the target image region of the target instance in the target image, including:
[0172] The processor 1201 performs instance segmentation processing on the target image to determine the image regions of each instance in at least one instance in the target image;
[0173] The processor 1201 obtains the position information of each image region in the target image;
[0174] If it is determined based on the position information of each image region that each image region is located in the central region of the target image, the processor 1201 uses each image region as the target image region.
[0175] In one embodiment, the processor 1201 performs instance segmentation processing on the target image to determine the target image region of the target instance in the target image, including:
[0176] The processor 1201 performs instance segmentation processing on the target image to determine the image regions of each instance in at least one instance within the target image;
[0177] The processor 1201 determines a first image region among the determined at least one image region, where the region size of the first image region is larger than the region sizes of each second image region, and the second image regions are the image regions other than the first image region among the at least one image region;
[0178] The processor 1201 obtains the ratio of the region size of each second image region to the region size of the first image region;
[0179] The processor 1201 uses the second image regions and the first image region whose ratio is greater than the second preset ratio threshold as the target image regions.
[0180] The embodiments of the present application provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the method embodiments as described above Figure 2 、 Figure 6 and Figure 9 shown. Among them, the computer-readable storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0181] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for detecting the integrity of text data, characterized in that, Including: Performing instance segmentation processing on the target image to determine the target image region of the target instance in the target image; Performing text detection processing on the target image region to obtain at least one original text box, and determining the original text box region of each original text box in the target image to obtain at least one original text box region; Determining, in the at least one original text box region, the original text box region that completely overlaps with the target image region; Obtaining the angle difference between the inclination angle of the determined original text box region and the inclination angle of the target original text box region, where the target original text box region includes the original text box regions that partially overlap with the target image region in the at least one original text box region; If the angle difference is less than a preset angle threshold, determining that the text data of the target image region is incomplete.
2. The method according to claim 1, wherein The method further includes: Performing mirror padding processing on the image edge of the target image to obtain a padded image; Determining, in the padded image, the padded image region corresponding to the target image region; Performing text detection processing on the padded image region to obtain at least one padded text box and the padded text box region of each padded text box in the padded image; If there is a target padded text box region that partially overlaps with the target image region, determining that the text data of the target image region is incomplete.
3. The method according to claim 2, characterized in that, The method further includes: Determining target intersection points among the intersection points where the region boundary of the target padded text box region intersects with the region boundary of the target image region; Obtaining the distance between the target intersection points and the target image boundary of the padded image, where the target image boundary matches the text box sub-region in the target padded text box region that does not overlap with the target image region; If the distance is less than a first preset distance threshold, determining that the factor causing the text data of the target image region to be incomplete is the first factor; If the distance is greater than or equal to the first preset distance threshold, determining that the factor causing the text data of the target image region to be incomplete is the second factor.
4. The method according to claim 2, wherein The performing instance segmentation processing on the target image to determine the target image region of the target instance in the target image includes: Performing instance segmentation processing on the target image to determine the image regions of each instance in at least one instance in the target image; Obtaining the distance between each region boundary of any one of the determined at least one image region and the corresponding image boundary in the target image; If there is at least one region boundary with a distance less than a second preset distance threshold, using the any one image region as the target image region.
5. The method according to claim 1, wherein The method further includes: Determining a first original text box sub-region and a second original text box sub-region in the target original text box region, where the first original text box sub-region refers to the text box sub-region in the target original text box region that does not overlap with the target image region, and the second original text box sub-region refers to the text box sub-region in the target original text box region that overlaps with the target image region; Obtain the ratio of the regional size of the first original text box sub-region to the regional size of the second original text box sub-region; If the ratio is less than the first preset ratio threshold, it is determined that the text data in the target image region is incomplete.
6. The method according to claim 1, wherein The instance segmentation processing of the target image to determine the target image region of the target instance in the target image includes: Perform instance segmentation processing on the target image to determine the image regions of each instance in at least one instance in the target image; Obtain the position information of each of the image regions in the target image; If it is determined based on the position information of each of the image regions that each of the image regions is located in the central region of the target image, then each of the image regions is used as the target image region.
7. The method according to claim 1, characterized in that The instance segmentation processing of the target image to determine the target image region of the target instance in the target image includes: Perform instance segmentation processing on the target image to determine the image regions of each instance in at least one instance in the target image; Determine a first image region among the determined at least one image region, where the regional size of the first image region is greater than the regional sizes of each second image region, and the second image region is an image region other than the first image region among the at least one image region; Obtain the ratio of the regional size of each of the second image regions to the regional size of the first image region; Use the second image regions with the ratio greater than the second preset ratio threshold and the first image region as the target image region.
8. An integrity detection device for text data, characterized in that, The integrity detection device for the text data includes a processing unit, and the processing unit is configured to execute the integrity detection method for the text data according to any one of claims 1-7.
9. A computer storage medium, characterized in that, Computer program instructions are stored in the computer storage medium, and when the computer program instructions are executed by a processor, they are used to execute the integrity detection method for the text data according to any one of claims 1-7.
Citation Information
Patent Citations
Boundary-aware object removal and content filling
CN111242852A
Mixed-pasting bill image processing method, device, computer equipment and storage medium
CN111931664A