Image recognition method and device
By analyzing the layout field characteristics of paper certificate pictures, confirming the picture type and configuring segmentation anchor points, the problem of high layout requirements and low recognition accuracy in the existing technology of picture segmentation methods is solved, and more flexible and accurate picture recognition is achieved.
Patent Information
- Application Number
- CN202011266787.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-11-12
AI Technical Summary
When identifying pictures of page-type paper documents, the segmentation method requires high layout requirements and cannot flexibly configure the segmentation method, resulting in low recognition accuracy, especially when the picture is tilted or folded.
By analyzing the layout field characteristics in the picture to be segmented, confirming the image type, and configuring the segmentation anchor points based on the shared fields, flexible segmentation and identification of the picture can be achieved.
It improves the accuracy of image segmentation and recognition, and can automatically configure segmentation anchor points when the image layout is not fixed to ensure the accuracy of recognition results.
Smart Images

Figure CN112381096B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent recognition, and specifically to a method and device for image recognition. Background Art
[0002] At present, a very high accuracy rate can be achieved for simple card-type document recognition. For image recognition of page-type paper documents, the images are usually segmented according to the layout or text format in the image. This segmentation method has high requirements on the layout of the image. The recognition accuracy rate is low for images with non-fixed layout fields. It also cannot flexibly configure the segmentation method based on problems such as the tilt of the image in the page-type document.
[0003] Therefore, how to accurately identify image information in paper images has become a problem that needs to be solved urgently. Summary of the invention
[0004] The present application provides an image recognition method and device, which can accurately recognize image information in paper images.
[0005] In a first aspect, some embodiments of the present application provide a method for image segmentation, comprising: confirming the type of image to be segmented, wherein the type of image to be segmented is at least obtained by analyzing the characteristics of a layout field in the image to be segmented; configuring a segmentation anchor point according to the type of image to be segmented, wherein the segmentation anchor point is determined based on a common field included in the image to be segmented; and segmenting the image to be segmented according to the segmentation anchor point.
[0006] Therefore, the embodiment of the present application configures corresponding segmentation anchor points according to the type of layout field of the image to be segmented, so that the corresponding segmentation anchor points can be automatically configured according to the actual layout of the image, thereby improving the flexibility of segmenting the image and thus improving the accuracy of segmenting and identifying the image.
[0007] In combination with the first aspect, in at least one embodiment of the present application, the layout field includes multiple reference fields and target fields corresponding to each reference field in the multiple reference fields; the confirmation of the type of the picture to be segmented includes: confirming that the picture to be segmented belongs to a first category of pictures, wherein the first category of pictures includes: the position between the target field and the corresponding reference field is fixed; according to the type of the picture to be segmented, configuring the segmentation anchor point includes: determining the segmentation anchor point of the first category of pictures based on the reference field.
[0008] Therefore, the embodiment of the present application confirms that the picture to be segmented belongs to the first category of pictures, determines the segmentation anchor point based on the reference field in the first category of pictures, and can find the segmentation anchor point of the picture to be segmented with a fixed position between the target field and the corresponding reference field.
[0009] In combination with the first aspect, in another embodiment, the layout field includes multiple reference fields and target fields corresponding to each reference field in the multiple reference fields; the confirmation of the type of the picture to be segmented includes: confirming that the picture to be segmented belongs to a second category of pictures, wherein the second category of pictures includes: the position between the target field and the corresponding reference field is not fixed; the configuration of the segmentation anchor point according to the type of the picture to be segmented includes: obtaining the segmentation anchor point of the second category of pictures based on the determined target field.
[0010] Therefore, the embodiment of the present application can identify the second type of picture in which the position between the target field and the reference field is not fixed by obtaining the segmentation anchor point of the second type of picture, and use the target field as a reference for obtaining the segmentation anchor point, so that when the target field in the picture is skewed, the segmentation anchor point can be accurately found, thereby accurately segmenting and identifying the picture.
[0011] In combination with the first aspect, in another implementation, obtaining the segmentation anchor point of the second category of pictures based on the target field includes: obtaining the common target field from multiple target fields included in multiple pictures; and determining the segmentation anchor point of the second category of pictures according to the common target field.
[0012] Therefore, the embodiment of the present application can find an indispensable target field from multiple target fields in the image as a segmentation anchor point by obtaining a determined target field, thereby eliminating inaccurate segmentation anchor points and incomplete image recognition caused by inconsistency in the target field.
[0013] In combination with the first aspect, in another embodiment, confirming the type of the picture to be segmented includes: confirming that the picture to be segmented belongs to a third category of pictures, wherein the third category of pictures is determined by whether there is a broken line in the picture to be segmented; configuring the segmentation anchor point according to the type of the picture to be segmented includes: determining the broken line segmentation anchor point of the third category of pictures according to the target field closest to the broken line; segmenting the picture to be segmented according to the segmentation anchor point includes: segmenting the third category of pictures into at least two sub-pictures according to the broken line segmentation anchor point.
[0014] Therefore, the embodiment of the present application can split the entire page of the fold line image by determining the fold line segmentation anchor point of the third type of image according to the target field closest to the fold line, thereby improving the segmentation and recognition accuracy when there is a fold line in the image.
[0015] In combination with the first aspect, in another embodiment, the confirming the type of the image to be segmented includes: confirming the type of the sub-image according to the characteristics of the layout field; configuring the segmentation anchor point according to the type of the image to be segmented includes: configuring the layout segmentation anchor point of the sub-image according to the type of the sub-image; and segmenting the image to be segmented according to the segmentation anchor point includes: segmenting the sub-image according to the layout segmentation anchor point.
[0016] Therefore, the embodiment of the present application configures the layout segmentation anchor points of the sub-image according to the type of the sub-image. When the segmentation of the third type of image is completed, the type of the sub-image can be confirmed again, and the segmentation anchor points can be configured again according to the type, thereby improving the accuracy of segmentation and recognition.
[0017] In combination with the first aspect, in another implementation, the picture to be segmented is a paper picture.
[0018] In a second aspect, a method for extracting image information includes: using the image segmentation method as described in the first aspect and any one of its embodiments to segment an image to be segmented to obtain a segmented image; and extracting image information from the segmented image.
[0019] In a third aspect, a method for image recognition includes:
[0020] The image to be segmented is segmented using the image segmentation method as described in the first aspect and any one of its embodiments to obtain segmented images; image information in the segmented images is extracted; and the image information is identified.
[0021] In a fourth aspect, a picture segmentation device is provided, comprising: a classification module, configured to confirm the type of a picture to be segmented, wherein the type of the picture to be segmented is at least obtained by analyzing the characteristics of a layout field in the picture to be segmented; an anchor point setting module, configured to configure a segmentation anchor point according to the type of the picture to be segmented, wherein the segmentation anchor point is determined based on a common field included in the picture to be segmented; and a segmentation module, configured to segment the picture to be segmented according to the segmentation anchor point.
[0022] In combination with the fourth aspect, in one embodiment, the layout field includes multiple reference fields and target fields corresponding to each reference field in the multiple reference fields; the classification module is specifically configured to: confirm that the picture to be segmented belongs to a first category of pictures, wherein the first category of pictures includes: the position between the target field and the corresponding reference field is fixed; the anchor point setting module is specifically configured to: determine the segmentation anchor point of the first category of pictures based on the reference field.
[0023] In combination with the fourth aspect, in another embodiment, the layout field includes multiple reference fields and target fields corresponding to each reference field in the multiple reference fields; the classification module is specifically configured to: confirm that the picture to be segmented belongs to the second category of pictures, wherein the second category of pictures includes: the position between the target field and the corresponding reference field is not fixed; the anchor point setting module is specifically configured to: based on the determined target field, obtain the segmentation anchor point of the second category of pictures.
[0024] In combination with the fourth aspect, in another embodiment, the anchor point setting module is specifically configured to: obtain the segmentation anchor point of the second type of picture based on the target field, including: obtaining a common target field from multiple target fields included in multiple pictures; and determining the segmentation anchor point of the second type of picture based on the common target field.
[0025] In combination with the fourth aspect, in another embodiment, the classification module is specifically configured to: confirm that the picture to be segmented belongs to the third category of pictures, wherein the third category of pictures is determined by analyzing the characteristics of the layout field in the picture to be segmented and whether there is a fold line in the picture to be segmented; the anchor point setting module is specifically configured to: determine the fold line segmentation anchor point of the third category of pictures according to the target field closest to the fold line; the segmentation module is specifically configured to: segment the third category of pictures into at least two sub-pictures according to the fold line segmentation anchor point.
[0026] In combination with the fourth aspect, in another embodiment, the classification module is specifically configured to: confirm the type of the sub-image according to the characteristics of the layout field; the anchor point setting module is specifically configured to: configure the layout segmentation anchor point of the sub-image according to the type of the sub-image, wherein the segmentation of the image to be segmented according to the segmentation anchor point includes: segmenting the sub-image according to the layout segmentation anchor point.
[0027] In combination with the fourth aspect, in another implementation, the picture to be segmented is a paper picture.
[0028] In a fifth aspect, a device for extracting image information includes: using the image segmentation device as described in the fourth aspect and any one of its embodiments to segment the image to be segmented to obtain segmented images; and an extraction module configured to extract image information from the segmented images.
[0029] In the sixth aspect, an image recognition device comprises: using the image segmentation device as described in the fourth aspect and any one of its embodiments to segment the image to be segmented to obtain segmented images; using the image information extraction device as described in the fifth aspect to extract image information from the segmented images; and a recognition module configured to recognize the image information.
[0030] In the seventh aspect, an electronic device includes: a processor, a memory and a bus, wherein the processor is connected to the memory via the bus, and the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, they are used to implement the methods described in the first aspect, the second aspect, the third aspect and any one of the embodiments.
[0031] In an eighth aspect, a computer-readable storage medium stores a computer program, and when the computer program is executed by a server, the method described in any one of the first aspect, the second aspect, the third aspect and all the embodiments is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a scene graph shown in an embodiment of the present application;
[0033] Figure 2 This is an implementation process of an image recognition method shown in an embodiment of the present application;
[0034] Figure 3 This is a schematic diagram of the first type of picture shown in the embodiment of the present application;
[0035] Figure 4 It is a schematic diagram of the second type of picture shown in the embodiment of the present application;
[0036] Figure 5 It is a schematic diagram of the third type of picture shown in the embodiment of the present application;
[0037] Figure 6 This is a specific embodiment diagram of the second type of picture shown in the embodiment of the present application;
[0038] Figure 7 It is an internal module diagram of an image recognition device shown in an embodiment of the present application;
[0039] Figure 8 It is a diagram of an internal module of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0041] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0042] In the embodiments of the present application, the embodiments of the present application can be applied to various scenarios. Figure 1 As shown in the example, the scene includes a camera 110, a paper document 120, a picture to be segmented 130 and a server 140. The camera takes a picture of the paper document to obtain the picture to be segmented. The server obtains the picture to be segmented and uses the picture recognition method described in the embodiment of the present application to identify it. The above only selects one application scenario of the embodiment of the present application for illustration, and does not constitute a limitation on the application scenario. The embodiment of the present application is not limited to this.
[0043] In the related solutions, a high accuracy rate can be achieved for simple card-type document recognition. For page-type paper document image recognition, the image is usually segmented according to the layout or text format in the image. This segmentation method has high requirements for the image layout. The recognition accuracy rate is low for images with unfixed layout fields. It is also impossible to flexibly configure the segmentation method according to the fold, angle, tilt and other issues in the page-type document image. Therefore, how to accurately identify the image information in the paper image has become an urgent problem to be solved.
[0044] In view of the above situation, an embodiment of the present application provides an image recognition method and device, which confirms the type of the image to be segmented, wherein the type of the image to be segmented is at least obtained by analyzing the characteristics of the layout field in the image to be segmented; configures segmentation anchor points according to the type of the image to be segmented, wherein the segmentation anchor points are selected based on the common fields of the image to be segmented; segments the image to be segmented according to the segmentation anchor points; extracts image information in the segmented image; and identifies the image information, so as to accurately identify image information in a paper image.
[0045] Combine the following Figure 2 , describe in detail a specific embodiment of the image recognition method, such as Figure 2 The steps shown include:
[0046] 210, confirm the type of the image to be segmented.
[0047] The server confirms the type of the image to be segmented, wherein the type of the image to be segmented is at least obtained by analyzing the characteristics of the layout field in the image to be segmented.
[0048] The characteristics of the layout field of the embodiment of the present application are the layout characteristics of the fields existing in the image to be segmented, and the layout field includes multiple reference fields and target fields corresponding to each of the multiple reference fields. As an example, the reference field is a printed field fixed in the image to be segmented by printing, and the target field is a printed field printed in the image to be segmented corresponding to the content of each reference field. For example, Figure 3 The "name, gender and major" on the photographed picture are the printed reference fields, while the "Zhang San" printed corresponding to the name reference field, the "Male" printed corresponding to the gender reference field, and the "Automation" printed corresponding to the major reference field are the printed target fields.
[0049] In some embodiments of the present application, the type of the image to be segmented is obtained by analyzing the characteristics of the layout field in the image to be segmented. In other embodiments of the present application, the type of the image to be segmented is confirmed by analyzing the characteristics of the layout field in the image to be segmented and whether there are broken lines in the image to be segmented. The specific embodiments are described as follows.
[0050] In one embodiment, the server confirms that the picture to be segmented belongs to a first category of pictures, wherein the first category of pictures includes: a position between the target field and the corresponding reference field is fixed.
[0051] After the server obtains the image to be segmented, it confirms that the position between the target field and the corresponding reference field is fixed (that is, the target field corresponds to the corresponding reference field without misalignment or skewness). This type of image belongs to the first type of image. Figure 3 As shown, the reference fields Name, Gender and Major correspond neatly to the target fields Zhang San, Male and Automation, respectively, without any misalignment or skewness.
[0052] In one embodiment, the server confirms that the picture to be segmented belongs to a second type of picture, wherein the second type of picture includes: a position between the target field and the corresponding reference field is not fixed.
[0053] After the server obtains the image to be segmented, it confirms that the position between the target field and the corresponding reference field is not fixed (that is, the target field is misaligned or skewed with the reference field). This type of image belongs to the second type of image. Figure 4As shown, the reference fields include name, gender and major, and the target fields corresponding to the three reference fields are "Zhang San, Male and Automation". Figure 4 There is a situation where the position of each target field and its corresponding reference field is not fixed (that is, the font of the target field is skewed relative to the corresponding reference field), and they are not neatly corresponding. This type of target field and the corresponding reference field have unfixed positions and belong to the second type of pictures.
[0054] In one embodiment, the server confirms that the image to be segmented belongs to a third category of images, wherein the third category of images is determined by analyzing features of a layout field in the image to be segmented and whether there is a broken line in the image to be segmented.
[0055] After obtaining the image to be segmented, the server confirms that the position between the target field and the corresponding reference field of the image to be segmented is not fixed and there is a broken line. This type of image belongs to the third type of image. Figure 5 As shown, the overall layout has tilted fields (for example, tilted field 1, field 2, field 3...field 10), and folds as shown in L1, L2, L3 and L4 appear. This type of picture is confirmed as the third type of picture.
[0056] It should be noted that the reference field in the embodiments of the present application is a field with relatively regular layout (for example, a template field on an inherent template), while the target field is irregularly laid out compared to the corresponding reference field (for example, a printed field or handwritten text added after the template field on an inherent template). Some embodiments of the present application do not specifically limit the way in which the reference field and the target field are fixed on a paper carrier. In some embodiments of the present application, the reference field is a printed field on a printed image layout, and the target field is a printed field printed corresponding to each reference field. For example Figure 3 Name, gender, and major shown; the target field is the field on the image layout printout, for example Figure 3 As shown in FIG. 1 , FIG. 2 , FIG. 3 , and FIG. 4 , the positional relationship between the reference field and the target field can be fixed without misalignment or skewness, or can be non-fixed with misalignment or skewness.
[0057] The above describes the process of the server confirming the type of the image to be segmented. The following describes the process of configuring the segmentation anchor point according to the type of the image to be segmented.
[0058] 220, configuring segmentation anchor points according to the type of the image to be segmented.
[0059] A segmentation anchor point is configured according to the type of the picture to be segmented, wherein the segmentation anchor point is selected based on a common field of the picture to be segmented.
[0060] In the process of configuring the segmentation anchor points on the server, no matter which type of the above-mentioned image to be segmented is determined, the segmentation anchor points are selected based on the common fields in the image to be segmented.
[0061] In an embodiment of the present application, a common field is a field that appears in every image to be segmented of the same type, for example, the ID number field and name field in a paper image of a household registration book. If the selected segmentation anchor point does not belong to a common field, some images will be unable to be segmented.
[0062] In one embodiment, confirming the type of the picture to be segmented includes: confirming that the picture to be segmented belongs to a first category of pictures, wherein the first category of pictures includes: the position between the target field and the corresponding reference field is fixed; and configuring a segmentation anchor point according to the type of the picture to be segmented, including: determining the segmentation anchor point of the first category of pictures based on the reference field.
[0063] When the server confirms that the image to be segmented belongs to the first category of images, the reference field is directly determined as the segmentation anchor point of the first category of images, or the horizontal line marked under the reference field is used as the segmentation anchor point for segmenting the first category of images, for example: Figure 3 In the layout shown, the horizontal line below the reference field "Gender" or "Sex" is used as the split anchor point.
[0064] Therefore, the embodiment of the present application confirms that the picture to be segmented belongs to the first category of pictures, determines the segmentation anchor point based on the reference field in the first category of pictures, and can find the segmentation anchor point of the picture to be segmented with a fixed position between the target field and the corresponding reference field.
[0065] In one embodiment, confirming the type of the picture to be segmented includes: confirming that the picture to be segmented belongs to a second category of pictures, wherein the second category of pictures includes: the position between the target field and the corresponding reference field is not fixed; configuring the segmentation anchor point according to the type of the picture to be segmented includes: obtaining the segmentation anchor point of the second category of pictures based on the determined target field.
[0066] In one embodiment, the determined target field is determined by the following method: obtaining the determined target field from a plurality of target fields included in a plurality of pictures, wherein each of the plurality of pictures includes the same reference field, the determined target field exists in each of the plurality of pictures, and the plurality of pictures belong to the second category of pictures.
[0067] In the case where the server confirms that the image to be segmented belongs to the second type of image, since the positions between the target fields and the corresponding reference fields are not fixed, the reference fields cannot be directly used as segmentation anchors. In this embodiment, first, from multiple images containing the same reference fields, the target fields that exist in each image are filtered out, and then the segmentation anchors are determined from the filtered target fields. For example: as Figure 6 shown, Figure 6 shows the information in the household register registration page. Among them, multiple reference fields such as name, gender, ethnicity, date of birth, etc. belong to the typesetting rules, while "Zhang X, male, Han ethnicity, X year X month X day" belongs to the target fields that may have irregular typesetting corresponding to each reference field (the reason for being the target fields is that after the image is segmented, the content of the target fields is required for information extraction and recognition). To ensure that there is no omission in the extraction of target information, the segmentation anchors for such images in the embodiments of the present application are statistically determined from multiple target fields. In addition, due to the randomness and irregularity of the input of the target fields, there may be cases where some reference fields have no corresponding target fields. Therefore, the embodiments of the present application also determine the common target fields from multiple similar images (for example, permanent resident registration card type images) as the basis for anchor point selection. Figure 6 In [reference document], "130******" and "X year X month X day" are respectively used as the segmentation anchors in the horizontal and vertical directions. Regardless of the printing situation of the entire layout, no information will be omitted. If a non-determined target field, such as "no religious belief", is selected as the segmentation anchor (for example: there are 5 permanent resident registration card type images, and 3 of them do not have the field of no religious belief, so this field cannot be used as the segmentation anchor), information will be omitted in some images that do not have the target information in this column. After the server finds the confirmed target fields, the confirmed target fields close to the middle position are selected from these target fields as the segmentation anchors.
[0068] Therefore, by confirming that the image to be segmented belongs to the second type of image and obtaining the segmentation anchor of the second type of image based on the determined target fields, the embodiments of the present application can identify the second type of image with non-fixed positions between the target fields and the reference fields, use the target fields as the reference for obtaining the segmentation anchor, so as to accurately find the segmentation anchor in the case where the target fields in the image are skewed, and thus accurately segment and recognize the image. By obtaining the determined target fields, the indispensable target fields can be found from various target fields in the image as the segmentation anchors, so as to eliminate the inaccuracy of the segmentation anchor caused by the non-uniformity of the target fields and the incomplete recognition of the image.
[0069] In one embodiment, confirming the type of the image to be segmented includes: confirming that the image to be segmented belongs to a third category of images, wherein the third category of images is determined by analyzing the characteristics of the layout fields in the image to be segmented and whether there is a fold line in the image to be segmented; configuring the segmentation anchor point according to the type of the image to be segmented includes: determining the fold line segmentation anchor point of the third category of images according to the target field that is closest to the fold line.
[0070] When the server confirms that the type of the image to be segmented belongs to the third type of image, it determines the target field near the broken line as the segmentation anchor point. Figure 5 As shown, L1, L2, L3 and L4 are the fold lines of the entire page, and the entire page is divided into the first part, the second part, the third part and the fourth part using field 3, field 4 and field 10 on the fold lines as segmentation anchor points.
[0071] Therefore, the embodiment of the present application confirms that the image to be segmented belongs to the third category of images, determines the fold line segmentation anchor point of the third category of images according to the target field closest to the fold line, and can split the entire page of the fold line image, thereby improving the segmentation and recognition accuracy when there are fold lines in the image.
[0072] The above describes an embodiment of configuring segmentation anchor points according to the type of the picture to be segmented. The following describes an embodiment of segmenting the picture to be segmented according to the segmentation anchor points.
[0073] 230, the image to be segmented is segmented according to the segmentation anchor points.
[0074] After confirming the type of the image to be segmented and configuring the segmentation anchor points, the server segments the image to be segmented according to the segmentation anchor points.
[0075] In one embodiment, segmenting the to-be-segmented picture according to the segmentation anchor point includes: when the to-be-segmented picture belongs to a third category of pictures, segmenting the third category of pictures into at least two sub-pictures according to the broken line segmentation anchor point;
[0076] The confirming the type of the image to be segmented includes: confirming the type of the sub-image according to the characteristics of the layout field; configuring the layout segmentation anchor points of the sub-image according to the type of the sub-image, wherein the segmenting of the image to be segmented according to the segmentation anchor points also includes: segmenting the sub-image according to the layout segmentation anchor points.
[0077] When the image to be segmented belongs to the third category of images, the third category of images is segmented into at least two sub-images according to the broken line segmentation anchor points. When it is determined whether the sub-image belongs to the first category of images, the second category of images, or the third category of images, the layout segmentation anchor points of the sub-images are configured according to the above method according to the type of the sub-image, and the sub-images are further segmented according to the layout segmentation anchor points.
[0078] Therefore, in the embodiment of the present application, when the image to be segmented belongs to the third category of images, the third category of images is segmented into at least two sub-images according to the broken line segmentation anchor point, and the layout segmentation anchor points of the sub-images are configured according to the types of the sub-images. When the segmentation of the third category of images is completed, the types of the sub-images can be confirmed again, and the segmentation anchor points can be configured again according to the types, thereby improving the accuracy of segmentation and recognition.
[0079] The above describes the process of the server segmenting the to-be-segmented picture according to the segmentation anchor points. The following describes the steps of extracting picture information from the segmented picture and identifying picture information.
[0080] 240, extracting image information from the segmented image.
[0081] The image to be segmented is segmented using the method of steps 210 to 230, and after the segmented images are obtained, the image information in the segmented images is extracted.
[0082] 250, identify image information.
[0083] The image to be segmented is segmented using the method of steps 210 to 230. After the segmented image is obtained, step 240 is performed to extract the image information in the segmented image, and then step 250 is performed to identify the image information to obtain the image content.
[0084] The above describes in detail the image segmentation method, image information extraction method and image recognition method. Figure 7 Describe a picture segmentation device, a picture information extraction device and a picture recognition device.
[0085] like Figure 7 As shown, a picture segmentation device includes: a classification module 710, an anchor point setting module 720, and a segmentation module 730.
[0086] In one embodiment, a picture segmentation device includes: a classification module, configured to confirm the type of the picture to be segmented, wherein the type of the picture to be segmented is at least obtained by analyzing the characteristics of the layout field in the picture to be segmented; an anchor point setting module, configured to configure a segmentation anchor point according to the type of the picture to be segmented, wherein the segmentation anchor point is determined based on the common fields included in the picture to be segmented; and a segmentation module, configured to segment the picture to be segmented according to the segmentation anchor point.
[0087] In one embodiment, the layout field includes multiple reference fields and target fields corresponding to each reference field in the multiple reference fields; the classification module is specifically configured to: confirm that the image to be segmented belongs to a first category of images, wherein the first category of images includes: the position between the target field and the corresponding reference field is fixed; the anchor point setting module is specifically configured to: determine the segmentation anchor point of the first category of images based on the reference field.
[0088] In another embodiment, the layout field includes multiple reference fields and target fields corresponding to each reference field in the multiple reference fields; the classification module is specifically configured to: confirm that the picture to be segmented belongs to the second category of pictures, wherein the second category of pictures includes: the position between the target field and the corresponding reference field is not fixed; the anchor point setting module is specifically configured to: based on the determined target field, obtain the segmentation anchor point of the second category of pictures.
[0089] In another embodiment, the anchor point setting module is specifically configured to: obtain the segmentation anchor point of the second category of pictures based on the target field, including: obtaining a common target field from multiple target fields included in multiple pictures; and determining the segmentation anchor point of the second category of pictures based on the common target field.
[0090] In another embodiment, the classification module is specifically configured to: confirm that the picture to be segmented belongs to the third category of pictures, wherein the third category of pictures is determined by analyzing the characteristics of the layout fields in the picture to be segmented and whether there is a fold line in the picture to be segmented; the anchor point setting module is specifically configured to: determine the fold line segmentation anchor point of the third category of pictures according to the target field closest to the fold line; the segmentation module is specifically configured to: segment the third category of pictures into at least two sub-pictures according to the fold line segmentation anchor point.
[0091] In another embodiment, the classification module is specifically configured to: confirm the type of the sub-image according to the characteristics of the layout field; the anchor point setting module is specifically configured to: configure the layout segmentation anchor point of the sub-image according to the type of the sub-image, wherein the segmentation of the image to be segmented according to the segmentation anchor point includes: segmenting the sub-image according to the layout segmentation anchor point.
[0092] In another embodiment, the image to be segmented is a paper image. Figure 7 The classification module 710, the anchor point setting module 720, and the segmentation module 730 shown in FIG. Figures 1 to 6 The method embodiment involves various processes in the image segmentation method. Figure 7 The operations and / or functions of each module in Figures 1 to 6 For details, please refer to the description in the above method embodiment. To avoid repetition, detailed description is appropriately omitted here.
[0093] like Figure 7 As shown, a picture information extraction device includes: a picture segmentation device and an extraction module 740.
[0094] In one embodiment, a picture segmentation device is used to segment a picture to be segmented to obtain segmented pictures; and an extraction module is configured to extract picture information from the segmented pictures.
[0095] In the embodiments of the present application, Figure 7 The classification module 710, the anchor point setting module 720, the segmentation module 730, and the extraction module 740 shown in FIG. Figures 1 to 6 The method embodiment involves various processes in the image information extraction method. Figure 7 The operations and / or functions of each module in Figures 1 to 6 For details, please refer to the description in the above method embodiment. To avoid repetition, detailed description is appropriately omitted here.
[0096] like Figure 7 As shown, a picture recognition device includes: a picture information extraction device and a recognition module 750.
[0097] In one embodiment, a picture segmentation device is used to segment a picture to be segmented to obtain segmented pictures; a picture information extraction device is used to extract picture information from the segmented pictures; and the recognition module 750 is configured to recognize the picture information.
[0098] In the embodiments of the present application, Figure 7The classification module 710, the anchor point setting module 720, the segmentation module 730, the extraction module 740 and the recognition module 750 shown in the figure can realize Figures 1 to 6 Each process in the picture recognition method in the method embodiment. Figure 7 The operations and / or functions of each module in Figures 1 to 6 For details, please refer to the description in the above method embodiment. To avoid repetition, detailed description is appropriately omitted here.
[0099] like Figure 8 As shown, an embodiment of the present application further proposes an electronic device, comprising: a processor 810, a memory 820 and a bus 830, wherein the processor is connected to the memory via the bus, and the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, they are used to implement any of the methods described in all the above-mentioned implementation modes. For details, please refer to the description in the above-mentioned method embodiments. To avoid repetition, the detailed description is appropriately omitted here.
[0100] Among them, the bus is used to realize the direct connection and communication of these components. Among them, the processor in the embodiment of the present application can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0101] The memory may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc. The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the method described in the above embodiment may be executed.
[0102] Understandably, Figure 8 The structure shown is for illustration only and may also include Figure 8 More or fewer components as shown, or with Figure 8 Different configurations are shown. Figure 8 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0103] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a server, any method described in any of the above-mentioned embodiments is implemented. For details, please refer to the description in the above-mentioned method embodiments. To avoid repetition, the detailed description is appropriately omitted here.
[0104] The above describes in detail the implementation steps of an image recognition method and the internal modules of an image recognition device. Figures 3 to 6 A specific embodiment of a picture recognition method is described.
[0105] In the image segmentation, extraction and recognition method, a positioning model is used to locate the segmentation anchor points in the image to be segmented. For any image, the layout can be segmented through the segmentation anchor points. This application utilizes a target detection algorithm to accurately locate the anchor point information in the image.
[0106] by Figure 6 Taking the second type of image in the resident registration card as an example, due to the complex printing method of the target fields in the second type of image, it is very difficult to obtain accurate layout analysis results without segmenting the entire layout. However, by segmenting the permanent population registration card and analyzing each segmented layout to obtain the layout analysis result of the entire image, it becomes very easy. Therefore, for the layout analysis of the permanent population registration card, the most important thing is to obtain the segmentation anchor point of any permanent population registration card image.
[0107] Faster-RCNN is a classic target detection algorithm that can accurately classify targets and locate target location information. The implementation method of this application is based on the Faster-RCNN algorithm. The reference fields in the permanent resident registration card are marked as three categories. The date of birth and ID number are marked as birthday and identification respectively, and other reference fields are marked as text. By training a large amount of labeled data, the corresponding positioning model can be obtained. For any permanent resident registration card image, the trained positioning model can accurately locate the segmentation anchor point in the permanent resident registration card image.
[0108] In a specific embodiment, Figure 3As shown, after the model is trained in advance, the server obtains the image to be segmented, confirms that the position between the target field and the corresponding reference field is fixed, and there is no misalignment or skewness, and determines that the image to be segmented belongs to the first category of images. The first category of images are input into the trained positioning model. In order to obtain the required target field "Zhang San, male", "gender" is set as the segmentation anchor point, a vertical critical value is set as the segmentation, and all target field sets that are less than the vertical critical value are identified to obtain the required target field.
[0109] The advantage of the above method for identifying the first type of images is that the image reference field is generally fixed and there is no situation where the segmentation anchor point cannot be detected due to the non-existence of the reference field.
[0110] In a specific embodiment, Figure 4 As shown, the server obtains the image to be segmented, confirms that the position between the target field and the corresponding reference field is not fixed, and there are situations such as wrong lines and skewness. It is determined that the image to be segmented belongs to the second category of images, and the second category of images are input into the trained positioning model. In order to obtain the required target field "Zhang San, Male", "Automation" is set as the segmentation anchor point, a vertical critical value is set as the segmentation, and all target field sets that are less than the vertical critical value are identified to obtain the required target field.
[0111] The above method for identifying the second type of images has the advantage that the selected segmentation anchor point can be changed according to different printing conditions. No matter how the printing conditions change, the required target field can be located through the set segmentation anchor point.
[0112] Optionally, in the above analysis of the paper image recognition method, if the image belongs to the first category of images, the recognition method of the first category of images is preferentially selected.
[0113] In a specific embodiment, Figure 6 As shown, the server obtains the image to be segmented, confirms that the position between the target field and the corresponding reference field is not fixed, and there are cases of misalignment, skewness, etc., and determines that the image to be segmented belongs to Figure 6 For the second type of pictures shown, due to the complicated printing of the target fields, it is necessary to select appropriate target fields as segmentation anchor points to segment the layout. In order to obtain the target field information corresponding to "Relationship with the head of household", "Gender", "Nationality" and "Date of Birth" in the upper right corner of the picture, if the reference field "Date of Birth" is used as the anchor point, it is likely to cause the target field "X year X month X day" to be missed, and thus the expected target field cannot be obtained. If the target field "X year X month X day" (determine the vertical threshold) and the target field "130******" (determine the horizontal threshold) are used as anchor points, no matter how the target field is printed, no information will be missed. Figure 6For example, to determine the vertical segmentation anchor point, it seems that you can choose "no religious belief", but through analyzing a large amount of data, it is found that the religious information column of many pictures does not exist, but the "date of birth" column exists in every household registration book, so the information in the upper right corner cannot be obtained. For the second type of pictures, choosing the right anchor point information to segment the layout is very important for layout analysis, and a large amount of picture data information needs to be trained to determine the corresponding segmentation anchor point. After the second type of pictures are segmented, extracted, and recognized, they are reassembled into a complete picture to obtain the corresponding picture information.
[0114] In a specific embodiment, Figure 5 As shown, the server obtains the image to be segmented, confirms that the position between the target field and the corresponding reference field is not fixed, there are misaligned lines, skewness, and broken lines, and determines that the image to be segmented belongs to Figure 5 The third category of pictures shown.
[0115] When the server confirms that the type of the image to be segmented belongs to the third type of image, it determines the target field near the broken line as the segmentation anchor point. Figure 5 As shown, L1, L2, L3 and L4 are the fold lines of the entire page. Based on L1 and L2, field 3 and field 10 are used as segmentation anchor points to divide the entire page into the upper and lower parts. Based on L3 and L4, field 4 is used as the segmentation anchor point to divide the entire page into the left and right parts. The entire page is divided into the first part, the second part, the third part and the fourth part. Among them, L1, L2, L3 and L4 change with the inclination angle of the left and right halves of the picture, which greatly increases the accuracy of page segmentation. If the dividing line is horizontal or vertical and unchanged, if the upper and lower dividing lines of the left half directly select the horizontal straight line L5 passing through field 3 as the upper dividing line, field 10 will be mistakenly divided into the upper half of the picture. This is not the segmentation result we expect. Therefore, it is very important to obtain an accurate page dividing line for folding pictures. The server will continue to determine the type of each part, for example: the first and fourth parts belong to the second type of pictures, and the second and third parts belong to the first type of pictures, and then identify each part according to the above method. After the identification is completed, it is combined into the original layout. The layout segmentation of the third type of pictures above is illustrated by dividing the picture into 4 parts. In actual situations, the picture can be divided into n (n>=2) parts according to needs.
[0116] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application. It should be noted that similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.
[0117] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for image segmentation, characterized in that: The picture segmentation method comprises: Confirming the type of the image to be segmented, wherein the type of the image to be segmented is at least obtained by analyzing the characteristics of the layout field in the image to be segmented; According to the type of the picture to be segmented, configuring a segmentation anchor point, wherein the segmentation anchor point is determined based on a common field included in the picture to be segmented; Segmenting the image to be segmented according to the segmentation anchor points; The layout field includes: a plurality of reference fields, and target fields corresponding to each of the reference fields; the reference fields are printed fields fixed in the image to be segmented by printing, and the target fields are printed fields corresponding to the content of each reference field in the image to be segmented; The confirming the type of the image to be segmented includes: If it is confirmed that the picture to be segmented belongs to the second type of picture, wherein the second type of picture is: the position between the target field and the corresponding reference field in the picture to be segmented is not fixed; The configuring the segmentation anchor point according to the type of the image to be segmented includes: Based on the determined target field, obtaining the segmentation anchor point of the second type of picture includes: Acquire the common field from the multiple target fields included in the multiple pictures; Determine a segmentation anchor point of the second type of picture according to the common field; If it is confirmed that the picture to be segmented belongs to the third category of pictures, wherein the third category of pictures is: the position between the target field and the corresponding reference field in the picture to be segmented is not fixed, and there is a broken line in the picture to be segmented; The configuring the segmentation anchor point according to the type of the image to be segmented includes: Determine the polyline segmentation anchor point of the third type of image according to the target field closest to the polyline; The step of segmenting the image to be segmented according to the segmentation anchor points includes: The third type of image is divided into at least two sub-images according to the broken line segmentation anchor point.
2. The image segmentation method according to claim 1, characterized in that: The step of confirming the type of the image to be segmented further includes: If it is confirmed that the picture to be segmented belongs to the first type of picture, wherein the first type of picture is: the position between the target field and the corresponding reference field in the picture to be segmented is fixed; The configuring the segmentation anchor point according to the type of the image to be segmented includes: Determine a segmentation anchor point of the first type of picture based on the reference field.
3. The image segmentation method according to claim 1, characterized in that: When it is confirmed that the picture to be segmented belongs to the third type of picture, the confirming the type of the picture to be segmented includes: Determining the type of the sub-image according to the characteristics of the layout field; The configuring the segmentation anchor point according to the type of the image to be segmented includes: According to the type of the sub-image, configuring the layout segmentation anchor point of the sub-image; The step of segmenting the image to be segmented according to the segmentation anchor points includes: The sub-image is segmented according to the layout segmentation anchor points.
4. The image segmentation method according to claim 1, characterized in that: The picture to be segmented is a paper picture.
5. A method for extracting image information, characterized in that: include: Segmenting the image to be segmented using the image segmentation method according to any one of claims 1 to 4 to obtain a segmented image; Extracting picture information from the segmented picture.
6. A method for image recognition, characterized in that: include: Segmenting the image to be segmented using the image segmentation method according to any one of claims 1 to 4 to obtain a segmented image; Extracting picture information from the segmented picture; The picture information is identified.
7. A picture segmentation device, characterized in that: The image segmentation device is based on the image segmentation method according to any one of claims 1 to 4, comprising: A classification module is configured to determine the type of the image to be segmented, wherein the type of the image to be segmented is at least obtained by analyzing the characteristics of the layout field in the image to be segmented; An anchor point setting module, configured to configure a segmentation anchor point according to the type of the picture to be segmented, wherein the segmentation anchor point is determined based on a common field included in the picture to be segmented; The segmentation module is configured to segment the image to be segmented according to the segmentation anchor points.
Citation Information
Patent Citations
Text layout analysis method and device, computer equipment and storage medium
CN111340037A