Image recognition method, device, electronic device and storage medium
By detecting the relative position relationship between the marking symbols and text information in the image, the problems of large computational complexity and poor robustness in the existing technology are solved, and an efficient image recognition method is realized.
Patent Information
- Application Number
- CN202111323401.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-09
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-11-09
AI Technical Summary
Existing technologies require multiple model trainings when identifying whether an option in an image is checked, resulting in large computational complexity and poor robustness.
By detecting the relative position relationship between the marker symbols and text information in the image to be recognized, the marking result of the text information is determined without considering the content and amount of the text information. The marker symbols and text positions are detected using a single-view detector and a text recognition model.
It reduces the problems of large computational complexity and poor robustness caused by multiple model training, and improves the efficiency and accuracy of image recognition.
Smart Images

Figure CN114049633B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of electronic technology, and more particularly to an image recognition method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] When conducting user surveys or conducting business, users are often required to fill out information on a document. The text content in the document typically includes several questions, each with a set of options. When filling out the document, the user selects at least one option based on their needs. After the user completes the questionnaire, it is necessary to determine whether the option in the document is checked. Summary of the Invention
[0003] The present disclosure provides an image recognition method, apparatus, electronic device and storage medium, computer-readable storage medium, and computer program product, which can alleviate the problem of high labor workload caused by multiple model training.
[0004] According to the first aspect of the present disclosure, an image recognition method is provided, including: detecting a marking symbol in an image to be recognized, and obtaining first position information of the marking symbol when the marking symbol exists in the image to be recognized; detecting text information in the image to be recognized, and obtaining second position information of the text information; and determining a marking result of the text information based on the relative position relationship between the first position information and the second position information.
[0005] According to an embodiment of the present disclosure, determining the marking result of the text information based on the relative position relationship between the first position information and the second position information includes: expanding the text area indicated by the second position information in the image to be identified to adjacent pixels to obtain an extended area; and when the symbol position indicated by the first position information is located in the extended area, determining that the marking result of the text information is marked.
[0006] According to an embodiment of the present disclosure, the text information includes at least two texts, and the text area indicated by the second position information includes at least two text sub-areas corresponding to the at least two texts respectively; for each sub-area in the at least two text sub-areas, the adjacent sub-areas of each sub-area in the at least two text sub-areas are determined; each sub-area is extended by a preset distance in a direction close to the adjacent sub-area to obtain an extended sub-area for each sub-area, wherein the preset distance is less than or equal to half of the spacing between each sub-area and the adjacent sub-area.
[0007] According to an embodiment of the present disclosure, when the symbol position indicated by the first position information is located in the extended area, determining that the marking result of the text information is marked includes: when the symbol position is located in any extended sub-area in the extended area, determining that the marking result of the text corresponding to any extended sub-area is marked.
[0008] According to an embodiment of the present disclosure, the method further includes: determining that the marker symbol is abnormal when the symbol position is outside the extended area; and outputting prompt information indicating that the marker symbol is abnormal.
[0009] According to an embodiment of the present disclosure, the method further includes: detecting text information using a text recognition model to obtain characters included in the text information; and outputting the characters when the marking result of the text information is marked.
[0010] According to an embodiment of the present disclosure, detecting the marker symbol in the image to be recognized includes: using a single-view detector to detect the marker symbol in the image to be recognized.
[0011] A second aspect of the present disclosure provides an image recognition device, comprising: a first position information determination module, a second position information determination module, and a marking result determination module. The first position information determination module is used to detect a marking symbol in an image to be recognized, and if a marking symbol exists in the image to be recognized, obtain first position information of the marking symbol. The second position information determination module is used to detect text information in the image to be recognized and obtain second position information of the text information. The marking result determination module is used to determine a marking result for the text information based on the relative positional relationship between the first position information and the second position information.
[0012] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned image recognition method.
[0013] A fourth aspect of the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the above-mentioned image recognition method.
[0014] The fifth aspect of the present disclosure further provides a computer program product, comprising a computer program, which implements the above-mentioned image recognition method when executed by a processor.
[0015] According to the embodiments of the present disclosure, the relative positional relationship between the marking symbol and the text information is used to determine the marking result of the text information. This eliminates the need to consider the content of the text information or the number of text items in the text information. Therefore, there is no need to train a separate model for each piece of text information, thereby alleviating the problems of high computational complexity and poor model robustness caused by multiple model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0017] Figure 1 Schematically illustrates an application scenario diagram of the image recognition method, apparatus, device, medium, and program product according to an embodiment of the present disclosure;
[0018] Figure 2 The following schematically shows a flow chart of an image recognition method according to an embodiment of the present disclosure;
[0019] Figure 3 Schematically shows a flow chart of an image recognition method according to another embodiment of the present disclosure;
[0020] Figure 4 Schematically shows a structural block diagram of an image recognition device according to an embodiment of the present disclosure; and
[0021] Figure 5 The block diagram schematically shows an electronic device suitable for implementing the image recognition method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0023] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0024] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0025] When expressions such as "at least one of A, B and C, etc." are used, they should generally be interpreted in accordance with the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0026] In order to identify whether options in a file are checked, in one technical solution, it is necessary to pre-train a model based on a set of options with specified option values and a specified number of options, which can identify whether each option in the set is checked. The generation of the model requires processes such as data preparation, model training, and model tuning.
[0027] When it is necessary to identify whether each option in an image is checked, all options in the image can be divided into multiple groups of options based on the question stem, and then the sub-image including a group of options can be cut out from the entire image. The cut-out sub-image is then input into the model trained for this group of options to obtain the check results of each option in the group of options.
[0028] With the above solution, the trained model can only be applied to a set of options with a specified value and a specified number of options. If at least one of the value or number of options in a set of options changes, a new model must be retrained to recognize the set of options. Therefore, the above technical solution suffers from high computational complexity and poor robustness.
[0029] An embodiment of the present disclosure provides an image recognition method, which includes: detecting a marking symbol in an image to be recognized, and obtaining first position information of the marking symbol when the marking symbol exists in the image to be recognized; detecting text information in the image to be recognized, and obtaining second position information of the text information; and determining a marking result of the text information based on the relative position relationship between the first position information and the second position information.
[0030] Figure 1 Schematically illustrates an application scenario diagram of the image recognition method, apparatus, device, medium, and program product according to an embodiment of the present disclosure;
[0031] like Figure 1 As shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0032] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0033] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0034] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process received data such as the image to be recognized, and feed back the processing results (such as the marking results and characters of each text in the text information generated based on the image to be recognized) to the terminal device.
[0035] It should be noted that the image recognition method provided in the embodiments of the present disclosure can generally be executed by the server 105. Accordingly, the image recognition device provided in the embodiments of the present disclosure can generally be set in the server 105. The image recognition method provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the image recognition device provided in the embodiments of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0036] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0037] The following will be based on Figure 1 The scene described by Figures 2 and 3 The image recognition method of the disclosed embodiment is described in detail.
[0038] Figure 2 The flowchart of the image recognition method according to the embodiment of the present disclosure is schematically shown.
[0039] like Figure 2 As shown, the image recognition method of this embodiment includes operations S210 to S230.
[0040] In operation S210 , a marking symbol in an image to be recognized is detected. If a marking symbol exists in the image to be recognized, first position information of the marking symbol is obtained.
[0041] The image to be identified can be the original image or a sub-image of the original image obtained by segmenting the original image. The original image is an image obtained by inputting the document to be identified using an input device. The document to be identified can be a questionnaire, a business processing form (such as a bank account opening form), etc. The input device can be a scanner, a camera, etc. The input process can be an operation such as scanning or taking a photo. For example, the original image is an image obtained by scanning a questionnaire using a scanner.
[0042] The image to be recognized includes text information. If the text information is marked, the image to be recognized also includes at least one marking symbol. The marking symbol represents a mark applied to the text information, and may be, for example, any of the symbols such as "√," "×," "○," "■," and "△." It should be understood that if the text information is not marked, the image to be recognized may not include a marking symbol.
[0043] The first position information may be the coordinates of a target point in the region where the marking symbol in the image to be recognized is located. The target point may be the center point of the region. The first position information may also be regional position information representing the region where the marking symbol in the image to be recognized is located. For example, the first position information may be a set of coordinates of at least some points in the region.
[0044] A pre-trained first model may be used to detect a marker symbol in the image to be recognized and determine first position information of the marker symbol. The pre-trained first model may be a You Only Look Once (YOLO) detector.
[0045] When training the first model, a certain number of samples must be prepared in advance, and the locations of the markers in each sample must be annotated. The samples can include different types of markers in different contexts, or markers of the same type. The markers can also include both typed and handwritten markers, thereby increasing sample diversity and improving the recognition capabilities of the first model.
[0046] In operation S220 , text information in the image to be recognized is detected to obtain second position information of the text information.
[0047] The text information can be a set of options below certain questions in the original document. The text information can include at least one text, which can be any option in the set. The text can include characters describing the specific content of the option, or characters indicating the check mark position. Characters indicating the check mark position can include characters such as a square check box and brackets.
[0048] The second position information may be the coordinate value of a target point in the area where each text in the image to be identified is located, and the target point may be the center point of the area. In practical applications, considering that the character lengths of each text in the text information may be different, in order to make the second position information more accurately describe the position of the text, the second position information may be the regional position information representing the area where each text in the image to be identified is located, for example, the second position information is a set of at least some coordinate points in the area. In addition, when the text includes at least two characters, each character corresponds to a character area in the image to be identified, and the regional position information of the text may include the character area of each character in the text. If the character areas of each character in a text are spaced apart, the position information of the continuous area from the first character area to the last character area can be determined as the regional position information of the text.
[0049] The second model can be used to detect text information in the image to be recognized and determine the first position information of the text information. The second model can be an Optical Character Recognition (ORC) model, etc., which is not limited in this disclosure.
[0050] In operation S230 , a marking result of the text information is determined according to the relative position relationship between the first position information and the second position information.
[0051] The marking results include marked and unmarked. Based on the relative positional relationship between the first position information and the second position information, the distance between the marking symbol and the text information can be determined. When the distance between any text in the text information and the marking symbol is close, the text is marked. When the distance between any text in the text information and all the marking symbols is far, the text is unmarked.
[0052] Those skilled in the art will understand that Figure 2 The sequence numbers of the steps are only used to define different operations and should not be regarded as a restriction on the order of execution of the method.
[0053] The disclosed embodiments separately determine the positions of markers and text within the image being processed, and then use the relative positional relationship between the markers and text to determine the labeling result for the text. Because the content of the text and the number of text items within it are not considered during the labeling process, there's no need to train a separate model for each piece of text. This alleviates the computational overhead associated with multiple model training cycles and improves the robustness of the disclosed image recognition method.
[0054] According to another embodiment of the present disclosure, Figure 2Operation S230 may include the following operations: first, expanding the text area indicated by the second position information in the image to be recognized toward adjacent pixels to obtain an extended area. Then, if the symbol position indicated by the first position information is located in the extended area, determining that the marking result of the text information is marked.
[0055] It should be noted that, considering that in actual applications, there may be a problem that the marking symbol is offset relative to the text information. For example, the marking symbol does not appear in the check box in the text, but appears around the text (for example, above, below, etc.). In order to adapt to the situation where the marking symbol is slightly offset, this embodiment expands the text area indicated by the second position information. The expansion means spreading from the text area to the outside of the text area, thereby obtaining an expanded area with a larger area. Therefore, when the marking symbol is slightly offset, the text near the marking symbol can still be identified as being marked.
[0056] In one example, the process of expanding the text area may include extending the edge of the text area by a predetermined distance outside the text area. The predetermined distance may be a preset fixed value, such as 1 centimeter, 2 centimeters, or the like. When the text information includes at least two texts, and the text area indicated by the second position information includes at least two text sub-areas corresponding to the at least two texts, the predetermined distance may be a preset fixed value or may be determined based on the distance between two adjacent text sub-areas, such as one-third of the distance between the two adjacent text sub-areas.
[0057] It should be understood that in other embodiments, if the problem of marking symbol offset does not need to be considered, the text area may not be expanded, and when the symbol position indicated by the first position information is located in the text area, the marking result of the text information is determined to be marked.
[0058] According to another embodiment of the present disclosure, the text information includes at least two texts, and the text area indicated by the second position information includes at least two text sub-areas corresponding to the at least two texts respectively. Accordingly, the operation of expanding the text area indicated by the second position information in the image to be identified to adjacent pixels to obtain an extended area in the above embodiment may include: first, for each sub-area of the at least two text sub-areas, determining the adjacent sub-areas of each sub-area of the at least two text sub-areas. Then, each sub-area is expanded by a preset distance in a direction close to the adjacent sub-area to obtain an extended sub-area for each sub-area. The preset distance is less than or equal to half of the spacing between each sub-area and the adjacent sub-area.
[0059] It should be noted that, when the text information includes at least two texts, the at least two texts may be arranged in a certain manner. For example, the at least two texts may be arranged horizontally, i.e., the at least two texts are spaced apart from each other from left to right. Alternatively, the at least two texts may be arranged vertically, i.e., the at least two texts are spaced apart from each other from top to bottom. Alternatively, the at least two texts may be arranged in a combination of horizontal and vertical arrangements, i.e., the at least two texts are arranged in rows and columns. In addition, in order to facilitate the distinction between different texts, the text sub-areas may be spaced apart from each other by a certain distance.
[0060] This embodiment determines the expansion direction of each text sub-region based on the relative positional relationship between two adjacent text sub-regions, and also determines the expansion distance of each text sub-region based on the distance between the two adjacent text sub-regions.
[0061] For example, for a rectangular text sub-region (for the convenience of distinction, the text sub-region is called a target sub-region), in the four directions above, below, left, and right, if there is a text sub-region adjacent to the target sub-region (i.e., an adjacent sub-region) in the first preset direction (any one of the four directions), and the distance between the target sub-region and the adjacent sub-region is the first preset distance. Then the extension distance of the target sub-region in the first preset direction is less than or equal to half of the first preset distance. In addition, for the above-mentioned target sub-region, in the four directions above, below, left, and right, if there is no adjacent sub-region in the second preset direction (any one of the four directions), then the extension distance of the target sub-region in the second preset direction can be the second preset distance. This embodiment does not limit the size relationship between the first preset distance and the second preset distance.
[0062] For example, the text information includes three texts, which are distributed left and right. The text sub-region corresponding to each text is a rectangle. The distance between the text sub-region on the left and the text sub-region in the middle is 6 cm, and the distance between the text sub-region in the middle and the text sub-region on the right is 8 cm.
[0063] For the middle text sub-region, the distance its left border extends to the left is less than or equal to 3 cm, and the distance its right border extends to the right is less than or equal to 4 cm. Furthermore, there are no adjacent sub-regions above or below the middle text sub-region, so the extension distances of the upper and lower borders of the middle text sub-region can use the second preset distance.
[0064] For the text sub-region on the left, the right boundary of the sub-region is extended to the right by a distance less than or equal to 3 cm. At the same time, there are no adjacent sub-regions above, below, or to the left of the text sub-region on the left. Therefore, the extension distances of the upper boundary, lower boundary, and left boundary of the middle text sub-region can adopt the second preset distance.
[0065] For the text sub-region on the right, the distance that its left border extends to the left is less than or equal to 4 cm. At the same time, there are no adjacent sub-regions above, below, or to the right of the text sub-region on the right. Therefore, the extension distances of the upper, lower, and right borders of the middle text sub-region can use the second preset distance.
[0066] This embodiment expands the text sub-region based on the connection direction and spacing of the text sub-regions, thereby accurately locating the text sub-region where the marking symbol is located.
[0067] According to another embodiment of the present disclosure, in the above embodiment, when the symbol position indicated by the first position information is located in the extended area, the operation of determining that the marking result of the text information is marked may include the following operations: when the symbol position is located in any extended sub-area in the extended area, determining that the marking result of the text corresponding to any extended sub-area is marked.
[0068] It should be noted that the text information includes at least two texts, each text has a text sub-region, and each text sub-region is expanded to obtain an extended sub-region of the text sub-region. The extended region includes an extended sub-region corresponding to each text in the text information.
[0069] The correspondence between an extended subregion and a text indicates that, in the image to be processed, the text is located within the extended subregion. When the text information includes at least two texts, it is necessary to determine the marking status of each text in the text information. For example, if the image to be recognized includes two marking symbols and five texts, and the first marking symbol is located in the extended subregion of the first text, and the second marking symbol is located in the extended subregion of the fourth text, then the first and fourth texts are marked, while the second, third, and fifth texts are not marked.
[0070] According to another embodiment of the present disclosure, the image recognition method further includes the following operations: first, when the symbol position is outside the extended area, determining that the marking symbol is abnormal, and then outputting prompt information indicating that the marking symbol is abnormal.
[0071] It should be noted that if the text information includes at least two texts, the extended area includes at least two extended sub-areas. If the symbol position is outside the extended area, it means that the symbol position is outside all extended sub-areas. If the symbol position does not fall within the extended area, it means that the annotation symbol exists in the image to be recognized, but the text annotated by the annotation symbol cannot be determined. In this case, the annotation symbol can be determined to be in an abnormal state.
[0072] Outputting a prompt indicating that a marking symbol is abnormal may involve circling the abnormal marking symbol in the image to be recognized, or may involve outputting first position information of the abnormal marking symbol. This prompts the user that the text marked by the abnormal marking symbol cannot be determined, and the user may manually verify the abnormal marking symbol as needed.
[0073] According to another embodiment of the present disclosure, the image recognition method further includes the following operations: first, detecting text information using a text recognition model to obtain characters included in the text information, and then outputting the characters if the marking result of the text information is marked.
[0074] It should be noted that the characters included in the text information represent the specific content of the text information. For example, if an option in the image to be recognized is "□20 to 30 years old", the characters of the option are "□20 to 30 years old". Outputting "□20 to 30 years old" can allow users to intuitively understand the content of the marked text.
[0075] In practical applications, the second model can be used to detect the second position information of the text information, and then the character recognition model can be used to detect the characters of the text information. The character recognition model can also be used to detect the second position information and the characters.
[0076] According to another embodiment of the present disclosure, Figure 2 Detecting the marker symbol in the image to be recognized in operation S230 may include the following operation: detecting the marker symbol in the image to be recognized using a single-view detector.
[0077] The single-look detector can use YOLO V4 or YOLO V5. The single-look detector has a faster detection speed and higher detection accuracy, thus ensuring the accuracy of the labeling results.
[0078] like Figure 3 As shown, the image recognition method of the present disclosure is described below using a specific embodiment. The method includes operations S310 to S360. Those skilled in the art will understand that the following embodiments are merely examples and the present disclosure is not limited thereto.
[0079] In operation S310 , a single-look detector is used to detect a marker symbol in an image to be recognized, and if the marker symbol exists in the image to be recognized, first position information of the marker symbol is obtained.
[0080] In operation S320, the text information in the image to be recognized is detected using a text recognition model to obtain second position information of the text information and characters included in the text information. The text information includes at least two texts, and the text area indicated by the second position information includes at least two text sub-areas corresponding to the at least two texts respectively.
[0081] In operation S330 , for each of the at least two text sub-regions, an adjacent sub-region of each of the at least two text sub-regions is determined.
[0082] In operation S340, each sub-region is expanded by a preset distance toward an adjacent sub-region to obtain an expanded sub-region for each sub-region. The preset distance is less than or equal to half of the distance between each sub-region and the adjacent sub-region.
[0083] In operation S350 , when the symbol position is located in any extended sub-region in the extended region, a marking result of the text corresponding to any extended sub-region is determined to be marked.
[0084] In operation S360 , in a case where the marking result of the text information is marked, the characters are output.
[0085] Based on the above image recognition method, the present disclosure also provides an image recognition device. Figure 4 The device is described in detail.
[0086] Figure 4 The structural block diagram of the image recognition device according to an embodiment of the present disclosure is schematically shown.
[0087] like Figure 4 As shown, the image recognition device 400 of this embodiment includes a first position information determination module 410 , a second position information determination module 420 and a marking result determination module 430 .
[0088] The first position information determination module 410 is used to detect the marking symbol in the image to be recognized, and obtain the first position information of the marking symbol when the marking symbol exists in the image to be recognized. In one embodiment, the first position information determination module 410 can be used to perform the operation S210 described above, which will not be repeated here.
[0089] The second position information determination module 420 is used to detect text information in the image to be recognized and obtain the second position information of the text information. In one embodiment, the second position information determination module 420 can be used to perform the operation S220 described above, which will not be repeated here.
[0090] The marking result determination module 430 is used to determine the marking result of the text information according to the relative position relationship between the first position information and the second position information. In one embodiment, the marking result determination module 430 can be used to perform the operation S230 described above, which will not be repeated here.
[0091] According to an embodiment of the present disclosure, the marking result determination module 430 includes an expansion submodule and a marking result determination submodule. The expansion submodule is configured to expand the text region indicated by the second position information in the image to be recognized toward adjacent pixels to obtain an expanded region. The marking result determination submodule is configured to determine that the marking result of the text information is marked if the symbol position indicated by the first position information is within the expanded region.
[0092] According to an embodiment of the present disclosure, the text information includes at least two texts, and the text region indicated by the second position information includes at least two text subregions corresponding to the at least two texts, respectively. The expansion submodule is configured to expand each of the at least two text subregions toward a predetermined number of adjacent pixels along a line connecting the at least two text subregions, thereby obtaining an expanded subregion for each text subregion; wherein the predetermined number is less than or equal to half the number of pixels between any two adjacent text subregions within the at least two text subregions.
[0093] According to an embodiment of the present disclosure, the marking result determining submodule is configured to determine that the marking result of the text corresponding to any extended subregion is marked when the symbol position is located in any extended subregion in the extended region.
[0094] According to an embodiment of the present disclosure, the image recognition device 400 further includes an anomaly determination module and a prompt information output module. The anomaly determination module is configured to determine that the marking symbol is abnormal when the symbol position is outside the extended area. The prompt information output module is configured to output prompt information indicating that the marking symbol is abnormal.
[0095] According to an embodiment of the present disclosure, the image recognition device 400 further includes a character determination module and a character output module. The character determination module is configured to detect text information using a text recognition model and obtain characters included in the text information. The character output module is configured to output characters when the text information is marked as marked.
[0096] According to an embodiment of the present disclosure, the first position information determining module 410 is configured to detect a marking symbol in an image to be recognized using a single-view detector.
[0097] According to an embodiment of the present disclosure, any multiple modules of the first position information determination module 410, the second position information determination module 420, and the marking result determination module 430 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the first position information determination module 410, the second position information determination module 420, and the marking result determination module 430 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in an appropriate combination of any of them. Alternatively, at least one of the first position information determining module 410 , the second position information determining module 420 , and the marking result determining module 430 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.
[0098] Figure 5 The block diagram schematically shows an electronic device suitable for implementing the image recognition method according to an embodiment of the present disclosure.
[0099] like Figure 5 As shown, the electronic device 500 according to an embodiment of the present disclosure includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage part 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include an onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0100] Various programs and data required for the operation of the electronic device 500 are stored in the RAM 503. The processor 501, ROM 502, and RAM 503 are connected to each other via a bus 504. The processor 501 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 502 and / or RAM 503. It should be noted that the programs may also be stored in one or more memories other than the ROM 502 and RAM 503. The processor 501 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0101] According to an embodiment of the present disclosure, the electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to the bus 504. The electronic device 500 may further include one or more of the following components connected to the I / O interface 505: an input portion 506 including a keyboard, a mouse, etc.; an output portion 507 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage portion 508 including a hard disk; and a communication portion 509 including a network interface card such as a LAN card or a modem. The communication portion 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in the drive 510 as needed, so that a computer program read therefrom can be installed into the storage portion 508 as needed.
[0102] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.
[0103] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 502 and / or RAM 503 described above and / or one or more memories other than ROM 502 and RAM 503.
[0104] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is executed in a computer system, the program code is used to cause the computer system to implement the item recommendation method provided by the embodiments of the present disclosure.
[0105] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the computer program is executed by the processor 501. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0106] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 505, and / or installed from the removable medium 511. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0107] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 505, and / or installed from the removable medium 511. When the computer program is executed by the processor 501, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0108] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0110] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways, even if such combinations and / or couplings are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or couplings are intended to fall within the scope of this disclosure.
[0111] The embodiments of the present disclosure are described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. An image recognition method, comprising: detecting a marking symbol in an image to be identified, and obtaining first position information of the marking symbol if the marking symbol exists in the image to be identified; Detecting text information in the image to be recognized and obtaining second position information of the text information; as well as determining a marking result of the text information according to a relative position relationship between the first position information and the second position information, The determining of the marking result of the text information according to the relative position relationship between the first position information and the second position information includes: Expanding the text area indicated by the second position information in the image to be recognized toward adjacent pixels to obtain an expanded area; and In a case where the symbol position indicated by the first position information is located in the extension area, determining that the marking result of the text information is marked; The text information includes at least two texts, the text area indicated by the second position information includes at least two text sub-areas corresponding to the at least two texts respectively; and the step of extending the text area indicated by the second position information in the image to be recognized to adjacent pixels to obtain the extended area includes: For each sub-region of the at least two text sub-regions, determining an adjacent sub-region of the at least two text sub-regions; and Each of the sub-regions is extended by a preset distance in a direction close to the adjacent sub-region to obtain an extended sub-region for each of the sub-regions; The preset distance is less than or equal to half of the distance between each sub-region and the adjacent sub-region.
2. The method according to claim 1, wherein In a case where the symbol position indicated by the first position information is located in the extended area, determining that the marking result of the text information is marked includes: In a case where the symbol position is located in any extended sub-region in the extended region, it is determined that the marking result of the text corresponding to any extended sub-region is marked.
3. The method according to claim 1, further comprising: In the case where the symbol position is outside the extended area, determining that the marking symbol is abnormal; as well as Outputting prompt information indicating that the marking symbol is abnormal.
4. The method according to claim 1, further comprising: Detecting the text information using a character recognition model to obtain characters included in the text information; as well as When the marking result of the text information is marked, the character is output.
5. The method according to claim 1, wherein Detecting marker symbols in the image to be recognized includes: A single-look detector is used to detect the marker symbol in the image to be recognized.
6. An image recognition device comprising: a first position information determining module, configured to detect a marking symbol in an image to be identified, and obtain first position information of the marking symbol if the marking symbol exists in the image to be identified; A second position information determining module, configured to detect text information in the image to be recognized and obtain second position information of the text information; as well as a marking result determining module, configured to determine a marking result of the text information based on a relative position relationship between the first position information and the second position information, The marking result determination module is further configured to: Expanding the text area indicated by the second position information in the image to be recognized toward adjacent pixels to obtain an expanded area; and In a case where the symbol position indicated by the first position information is located in the extension area, determining that the marking result of the text information is marked; The text information includes at least two texts, the text area indicated by the second position information includes at least two text sub-areas corresponding to the at least two texts respectively; and the step of extending the text area indicated by the second position information in the image to be recognized to adjacent pixels to obtain the extended area includes: For each sub-region of the at least two text sub-regions, determining an adjacent sub-region of the at least two text sub-regions; and Each of the sub-regions is extended by a preset distance in a direction close to the adjacent sub-region to obtain an extended sub-region for each of the sub-regions; The preset distance is less than or equal to half of the distance between each sub-region and the adjacent sub-region.
7. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 5.
9. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Method and system for outputting text line content after document image checkbox state recognition
CN110659574A