Image information recognition method, device and storage medium

By jointly identifying business card images with multiple algorithms and leveraging the advantages of each detection algorithm, the problems of low efficiency and unstable recognition in existing technologies are solved, and high-accuracy and high-practicality business card information entry is achieved.

CN115187992BActive Publication Date: 2025-10-03HANGZHOU HIKVISION SYST TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210879317.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-25
Publication Date
2025-10-03
Estimated Expiration
2042-07-25

AI Technical Summary

Technical Problem

Existing business card entry methods are inefficient and prone to errors, especially when affected by external factors such as light and angle deviation, the recognition effect is unstable.

Method used

It adopts a multi-algorithm joint recognition method to identify the target image separately through multiple detection algorithms. By combining the advantages of each algorithm, it can determine the target text area and content, eliminate interference items, and improve recognition accuracy.

Benefits of technology

Under the influence of various external factors, high-accuracy and high-practicality image information recognition is achieved, which improves the efficiency and reliability of business card entry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115187992B_ABST
    Figure CN115187992B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, and storage medium for image information recognition, relating to the field of image recognition technology and capable of improving the accuracy and practicality of image information recognition. The method comprises: identifying a target image based on multiple detection algorithms, obtaining detection results corresponding to each detection algorithm; wherein the detection results corresponding to each detection algorithm include: regional coordinates of an initial text region in the target image, obtained based on the detection algorithm; obtaining a target detection result based on the detection results corresponding to each detection algorithm; wherein the target detection result includes the regional coordinates of a target text region for each text region in the target image; the target text region of a text region is determined based on the overlap of multiple initial text regions in the text region, determined by the multiple detection algorithms; and recognizing the text content of each text region from the target image based on the target detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and in particular to an image information recognition method, device, and storage medium. Background Art

[0002] As an important medium in daily business activities, business cards can effectively transmit basic personal information and become an effective means of quickly establishing connections. With the rapid development of terminal technologies such as smartphones and mobile tablets, the ability to efficiently and accurately copy and store information on traditional business cards in mobile terminal address books can effectively avoid the inconvenience of carrying various business cards and the problem of business card loss.

[0003] Traditional business card entry methods include manual entry and optical character recognition (OCR). Manual entry is not only inefficient but also prone to errors, making it difficult to scale. OCR is susceptible to external factors, such as lighting, angle deviation, and camera quality, resulting in unstable recognition results.

[0004] Therefore, a business card entry method with high accuracy and strong practicality is urgently needed. Summary of the Invention

[0005] The present application provides a method, device and storage medium for image information recognition, which can improve the accuracy and practicality of image information recognition.

[0006] In a first aspect, the present application provides a method for image information recognition, comprising: identifying a target image based on multiple detection algorithms, and obtaining a detection result corresponding to each detection algorithm; wherein the detection result corresponding to a detection algorithm includes: the regional coordinates of an initial text area where each text segment in the target image is located, obtained based on a detection algorithm; a detection result includes the regional coordinates of an initial text area; a target detection result is obtained based on the detection result corresponding to each detection algorithm; wherein the target detection result includes the regional coordinates of a target text area for each text segment in the target image; the target text area of ​​a text segment is determined based on the overlap of multiple initial text areas where a text segment is located, determined by multiple detection algorithms; and according to the target detection result, the text content of each text segment is identified from the target image.

[0007] It is understood that the image information recognition method provided by the present application is based on the joint recognition of image information using multiple algorithms, and multiple detection algorithms are all used in parallel mode. First, the target image is recognized based on the multiple detection algorithms, and the detection results corresponding to each detection algorithm are obtained. Then, the target detection result is obtained based on the detection results corresponding to each detection algorithm. Then, based on the target detection results, the text content of each text segment is recognized from the target image. In this way, compared with the detection methods of the prior art that use a single algorithm, the embodiments of the present application use a multi-algorithm joint recognition method to simultaneously recognize the target image based on multiple detection algorithms, fully utilizing the advantages of various detection algorithms to achieve better detection results for the target image under the influence of various external factors. For example, if algorithm A excels at recognizing horizontally distributed text and algorithm B excels at recognizing multi-directional text, then using the multi-algorithm joint method provided by the embodiments of the present application to simultaneously recognize the target image based on algorithm A and algorithm B can obtain a more accurate result. For another example, if algorithm C excels at handling scenes with poor lighting and algorithm D excels at handling scenes with angular deviation, then using the multi-algorithm joint method provided by the embodiments of the present application to simultaneously recognize the target image based on algorithm C and algorithm D can overcome the effects of poor lighting and angular deviation on the results. In this way, based on the method provided in the embodiments of the present application, highly accurate results can be obtained in various scenarios, effectively improving the accuracy and practicality of image information recognition.

[0008] In one possible implementation, a detection result also includes: an area number of an initial text area; the area numbers of the initial text areas where the same text segment is located are the same; the above-mentioned target detection result is obtained based on the detection results corresponding to each detection algorithm, including: grouping the detection results corresponding to each detection algorithm according to the area number of the initial text area to obtain multiple detection groups; wherein each detection group includes the area coordinates of multiple initial text areas where the same text segment is located; for the first detection group among the multiple detection groups, from the multiple initial text areas included in the first detection group, determine that the overlapping area whose number of overlaps meets the first preset condition is the target text area of ​​the first detection group; the target detection result includes the target text area of ​​each detection group.

[0009] It will be appreciated that each detection group includes the coordinates of multiple initial text regions within the same text segment. Thus, the target text region can be determined based on the coordinates of the multiple initial text regions within the same text segment. For example, the target text region can be determined based on the overlapping region of the multiple initial text regions with the greatest overlap. Because the region included in each initial text region is highly likely the region within the detection group's text segment, the target text region determined based on the overlapping region of the multiple initial text regions is more accurate.

[0010] In another possible implementation, a detection result also includes: a confidence level corresponding to the regional coordinates of an initial text area; the above-mentioned determination of the overlapping area with the largest number of overlaps from the multiple initial text areas included in the first detection group as the target text area of ​​the first detection group includes: determining a valid text area from the multiple initial text areas included in the first detection group; the confidence level corresponding to the regional coordinates of the valid text area is greater than a first preset threshold; and determining the target text area of ​​the first detection group from the valid text areas included in the first detection group.

[0011] It can be understood that based on the solution provided in the present application, valid text areas with a confidence level greater than a first preset threshold are screened out from multiple initial text areas, and then the target text area is determined based on the valid text area. This can eliminate interference items (invalid text areas) and improve the accuracy of image recognition results.

[0012] In another possible implementation, the above-mentioned identifying the text content of each text segment from the target image based on the target detection result includes: identifying the regional coordinates of the target text area of ​​each text segment based on multiple recognition algorithms and the regional coordinates of the target text area of ​​each text segment, and obtaining the recognition result corresponding to each recognition algorithm; wherein the recognition result corresponding to a recognition algorithm includes: the initial text content of the target text area of ​​each text segment obtained based on a recognition algorithm; a recognition result includes the initial text content of the target text area of ​​a text segment; obtaining a target recognition result based on the recognition result corresponding to each recognition algorithm; wherein the target recognition result includes the target text content of the target text area of ​​each text segment; the target text content of the target text area of ​​a text segment is determined based on the repetition of multiple initial text contents of the target text area of ​​a text segment determined by multiple recognition algorithms.

[0013] It is understood that the method provided by this application uses multiple recognition algorithms to identify the text content of each text segment from the target image. Thus, compared to the existing recognition methods that use a single algorithm, the embodiments of this application utilize multiple algorithms for joint recognition, fully leveraging the advantages of various recognition algorithms to achieve better recognition results for the target image regardless of the influence of various external factors, effectively improving the accuracy and practicality of image information recognition.

[0014] In another possible implementation, a recognition result also includes: an area number of a target text area; the above-mentioned target recognition result is obtained according to the recognition result corresponding to each recognition algorithm, including: grouping the recognition result corresponding to each recognition algorithm according to the area number of the target text area to obtain multiple recognition groups; wherein each recognition group includes multiple initial text contents of the target text area of ​​the same text segment; for the first recognition group among the multiple recognition groups, from the multiple initial text contents of the target text area included in the first recognition group, determine the text content whose repetition number meets the second preset condition as the target text content of the target text area of ​​the first recognition group; the target recognition result includes the target text content of the target text area of ​​each recognition group.

[0015] It is understood that each recognition group includes multiple initial text contents for the target text area of ​​the same text segment. Thus, the target text content can be determined based on the multiple initial text contents. For example, the text content with the most repetitions among the multiple initial text contents can be used as the target text content. Because the text contained in each initial text content is highly likely to be the text content corresponding to the text segment of the recognition group, the text content determined based on the repetition of the multiple initial text contents is more accurate.

[0016] In another possible implementation, a recognition result also includes: a confidence level corresponding to the initial text content of a target text area of ​​a text segment; the above-mentioned determination of the text content whose repetition times meet the second preset condition from the multiple initial text contents of the target text area included in the first recognition group as the target text content of the target text area of ​​the first recognition group includes: determining valid text content from the multiple initial text contents of the target text area included in the first recognition group; the confidence level corresponding to the valid text content is greater than the second preset threshold; and determining the target text content of the target text area of ​​the first recognition group from the valid text contents of the target text area included in the first recognition group.

[0017] It can be understood that based on the solution provided in this application, valid text content with a confidence level greater than a second preset threshold is screened out from multiple initial text contents, and then the target text content is determined based on the valid text content. This can eliminate interference items (invalid text content) and improve the accuracy of image recognition results.

[0018] In another possible implementation, the number of the plurality of detection algorithms is an odd number; and / or the number of the plurality of recognition algorithms is an odd number.

[0019] It is understood that since this application determines the target text area based on the number of overlaps of the initial text areas, to avoid the same number of overlaps, the number of initial text areas is preferably an odd number, that is, the number of multiple detection algorithms is an odd number. Similarly, the target text content is determined based on the number of repetitions of the initial text content. To avoid the same number of repetitions, the number of initial text content is preferably an odd number, that is, the number of multiple recognition algorithms is an odd number.

[0020] In a second aspect, the present application provides an image information recognition device, comprising: a detection module, configured to identify target images based on multiple detection algorithms, and obtain detection results corresponding to each detection algorithm; wherein the detection result corresponding to a detection algorithm includes: the regional coordinates of the initial text area where each text segment in the target image is located, obtained based on a detection algorithm; a detection result includes the regional coordinates of an initial text area; a processing module, configured to obtain a target detection result based on the detection result corresponding to each detection algorithm; wherein the target detection result includes the regional coordinates of the target text area for each text segment in the target image; the target text area of ​​a text segment is determined based on the overlap of multiple initial text areas where a text segment is located, determined by multiple detection algorithms; and a recognition module, configured to recognize the text content of each text segment from the target image based on the target detection result.

[0021] In one possible implementation, a detection result also includes: an area number of an initial text area; the initial text areas where the same text segment is located have the same area number; a processing module, specifically used to group the detection results corresponding to each detection algorithm according to the area number of the initial text area to obtain multiple detection groups; wherein each detection group includes the area coordinates of multiple initial text areas where the same text segment is located; for the first detection group among the multiple detection groups, from the multiple initial text areas included in the first detection group, determine that the overlapping area whose number of overlaps meets the first preset condition is the target text area of ​​the first detection group; the target detection result includes the target text area of ​​each detection group.

[0022] In another possible implementation, a detection result also includes: a confidence level corresponding to the area coordinates of an initial text area; a processing module specifically used to determine a valid text area from multiple initial text areas included in the first detection group; the confidence level corresponding to the area coordinates of the valid text area is greater than a first preset threshold; and determining a target text area of ​​the first detection group from the valid text areas included in the first detection group.

[0023] In another possible implementation, the recognition module is specifically configured to identify the regional coordinates of the target text area of ​​each text segment based on multiple recognition algorithms and the regional coordinates of the target text area of ​​each text segment, and obtain a recognition result corresponding to each recognition algorithm; wherein the recognition result corresponding to a recognition algorithm includes: the initial text content of the target text area of ​​each text segment obtained based on a recognition algorithm; a recognition result includes the initial text content of the target text area of ​​a text segment; a target recognition result is obtained based on the recognition result corresponding to each recognition algorithm; wherein the target recognition result includes the target text content of the target text area of ​​each text segment; the target text content of the target text area of ​​a text segment is determined based on the repetition of multiple initial text contents of the target text area of ​​a text segment determined by multiple recognition algorithms.

[0024] In another possible implementation, a recognition result also includes: an area number of a target text area; a recognition module, specifically used to group the recognition results corresponding to each recognition algorithm according to the area number of the target text area to obtain multiple recognition groups; wherein each recognition group includes multiple initial text contents of the target text area of ​​the same text segment; for the first recognition group among the multiple recognition groups, from the multiple initial text contents of the target text area included in the first recognition group, determine the text content whose repetition number meets the second preset condition as the target text content of the target text area of ​​the first recognition group; the target recognition result includes the target text content of the target text area of ​​each recognition group.

[0025] In another possible implementation, a recognition result also includes: a confidence level corresponding to the initial text content of a target text area of ​​a text segment; a recognition module specifically used to determine valid text content from multiple initial text contents of the target text area included in the first recognition group; the confidence level corresponding to the valid text content is greater than a second preset threshold; and the target text content of the target text area of ​​the first recognition group is determined from the valid text contents of the target text area included in the first recognition group.

[0026] In another possible implementation, the number of the multiple detection algorithms is an odd number; and / or the number of the multiple recognition algorithms is an odd number.

[0027] In a third aspect, the present application provides an image information recognition device, comprising a memory and a processor. The memory and processor are coupled. The memory is configured to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the image information recognition device performs the image information recognition method according to the first aspect and any possible design thereof.

[0028] In a fourth aspect, the present application provides a chip system for use in an image information recognition device. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via circuits. The interface circuits are configured to receive signals from a memory of the image information recognition device and send signals to the processors, the signals including computer instructions stored in the memory. When the processors execute the computer instructions, the image information recognition device performs the image information recognition method according to the first aspect and any of its possible designs.

[0029] In a fifth aspect, the present application provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on a computer, the computer executes the image information recognition method as in the first aspect and any possible design thereof.

[0030] In a sixth aspect, the present application provides a computer program product, which includes computer instructions. When the computer instructions are run on a computer, the computer executes the image information recognition method as in the first aspect and any possible design thereof.

[0031] For the specific descriptions of the second to sixth aspects and their various implementations in this application, reference can be made to the detailed descriptions in the first aspect and its various implementations; and for the beneficial effects of the second to sixth aspects and their various implementations, reference can be made to the analysis of the beneficial effects in the first aspect and its various implementations, which will not be repeated here.

[0032] These and other aspects of the present application will become more readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 Schematic diagram of the implementation environment involved in a method for identifying image information provided in an embodiment of the present application Figure 1 ;

[0034] Figure 2 Schematic diagram of the implementation environment involved in a method for identifying image information provided in an embodiment of the present application Figure 2 ;

[0035] Figure 3 The process of a method for identifying image information provided in the embodiment of the present application Figure 1 ;

[0036] Figure 4 A schematic diagram of a text area provided in an embodiment of the present application Figure 1 ;

[0037] Figure 5 A schematic diagram of a text area provided in an embodiment of the present application Figure 2 ;

[0038] Figure 6 A schematic diagram of a text area provided in an embodiment of the present application Figure 3 ;

[0039] Figure 7 The process of a method for identifying image information provided in the embodiment of the present application Figure 2 ;

[0040] Figure 8 A schematic diagram of a text area provided in an embodiment of the present application Figure 4 ;

[0041] Figure 9 A schematic diagram of a text area provided in an embodiment of the present application Figure 5 ;

[0042] Figure 10 A schematic diagram of a text area provided in an embodiment of the present application Figure 6 ;

[0043] Figure 11 A schematic diagram of a text area provided in an embodiment of the present application Figure 7 ;

[0044] Figure 12 A schematic diagram of a text area provided in an embodiment of the present application Figure 8 ;

[0045] Figure 13 The process of a method for identifying image information provided in the embodiment of the present application Figure 3 ;

[0046] Figure 14 A schematic diagram of the structure of an image information recognition device provided in an embodiment of the present application Figure 1 ;

[0047] Figure 15 A schematic diagram of the structure of an image information recognition device provided in an embodiment of the present application Figure 2 . DETAILED DESCRIPTION

[0048] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0049] The terms "first" and "second" and the like in the specification and drawings of this application are used to distinguish different objects, or to distinguish different processing of the same object, rather than to describe a specific order of objects.

[0050] Furthermore, the terms "including," "having," and any variations thereof, as used in the description of this application are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.

[0051] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being more preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0052] In the description of the present application, unless otherwise specified, “plurality” means two or more.

[0053] As mentioned in the background, traditional business card entry methods are ineffective. For example, manual entry is not only inefficient but also prone to errors, making it difficult to popularize. Optical character recognition (OCR) is susceptible to external factors, such as lighting, angle deviation, and camera quality. Consequently, OCR-based entry methods have low generalization capabilities and unstable recognition results.

[0054] Therefore, a business card entry method with high accuracy and strong practicality is urgently needed.

[0055] To address the above technical issues, embodiments of the present application provide a method for image information recognition. The method employs multiple algorithms to jointly recognize image information, employing multiple detection algorithms in parallel. First, the target image is recognized based on the multiple detection algorithms, obtaining detection results corresponding to each detection algorithm. A target detection result is then obtained based on the detection results corresponding to each detection algorithm. Furthermore, based on the target detection results, the text content of each text segment is recognized from the target image. Thus, compared to existing detection methods that employ a single algorithm, embodiments of the present application employ a multi-algorithm joint recognition approach, simultaneously recognizing the target image based on multiple detection algorithms. This fully leverages the advantages of each detection algorithm, enabling better detection results for the target image regardless of the influence of various external factors. For example, if algorithm A excels at recognizing horizontally distributed text, and algorithm B excels at recognizing text in multiple directions, then using the multi-algorithm joint approach provided by embodiments of the present application, simultaneously recognizing the target image based on both algorithms A and B, can yield highly accurate results. For another example, if algorithm C excels at handling scenes with poor lighting, and algorithm D excels at handling scenes with angular deviation, then using the multi-algorithm joint approach provided by embodiments of the present application, simultaneously recognizing the target image based on both algorithms C and D, can overcome the effects of poor lighting and angular deviation on the results. In this way, based on the method provided in the embodiments of the present application, highly accurate results can be obtained in various scenarios, effectively improving the accuracy and practicality of image information recognition.

[0056] Please refer to Figure 1 , which shows a schematic diagram of the implementation environment involved in a method for identifying image information provided in an embodiment of the present application. Figure 1 As shown, the implementation environment may include: a collection device 100 and an electronic device 200.

[0057] In some embodiments, the acquisition device 100 may be integrated into the electronic device 200 as a submodule of the electronic device 200; or, the acquisition device 100 and the electronic device 200 may be two independent devices.

[0058] The acquisition device 100 is used to acquire target images.

[0059] Exemplarily, the acquisition device 100 can obtain the target image by taking pictures or scanning using a camera, a mobile phone with a built-in camera, or a scanning device.

[0060] In some embodiments, the acquisition device 100 is further configured to send the acquired target image to the electronic device 200 .

[0061] The electronic device 200 is used to perform information recognition and information processing on the image, and finally generate a complete version of the target image information.

[0062] Specifically, the electronic device 200 obtains a target image, performs text detection and text recognition on the image using OCR technology, and then obtains text information on the target image.

[0063] For example, the electronic device 200 in the embodiment of the present application may be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) or virtual reality (VR) device, etc. The embodiment of the present application does not impose any particular limitation on the specific form of the electronic device.

[0064] Specifically, such as Figure 2 As shown, the acquisition device 100 sends the acquired target image to the electronic device 200; the electronic device 200 first uses multiple detection algorithms to perform text detection on the received target image to obtain the text area of ​​each text segment in the target image; then the electronic device 200 uses multiple recognition algorithms to recognize the text area of ​​each text segment in the target image to obtain the text content of the text area of ​​each text segment in the target image.

[0065] The following is a detailed description of a method for identifying image information provided in an embodiment of the present application.

[0066] The present invention provides a method for identifying image information, which is applied to Figure 1 The electronic equipment shown in . Figure 3 As shown, the method includes the following steps S101-S103.

[0067] S101 , identifying target images based on multiple detection algorithms respectively, and obtaining detection results corresponding to each detection algorithm.

[0068] In some embodiments, the target image may be a picture of a business card or a picture of a resume, etc. For example, the embodiment of the present application is described by taking the target image as a picture of a business card as an example.

[0069] In some embodiments, the detection algorithm can identify a text segment in a target image and represent the text area where the text segment is located in the form of a rectangular frame. The coordinates of the rectangular frame are the area coordinates of the text area where the text segment is located.

[0070] Among them, a text segment refers to a collection of multiple texts. Exemplarily, "Name: Zhang San" in a business card is a text segment; "Gender: Female" in a resume is a text segment.

[0071] In some embodiments, the detection result corresponding to a detection algorithm includes: the regional coordinates of the initial text region where each text segment in the target picture obtained based on a detection algorithm is located. It can be understood that the detection result corresponding to a detection algorithm can be one or more. For example, in the case where there are multiple text segments in the target picture, the detection result corresponding to a detection algorithm can be multiple. Exemplarily, if the target picture includes 4 text segments, namely: "Zhang San", "Project Manager", "XX Digital Technology Co., Ltd.", "Phone: 021 - 3800****", and if the multiple detection algorithms include: Detection Algorithm A, then the detection result corresponding to Detection Algorithm A has 4. Among them, Detection Result 1 includes: the regional coordinates of the initial text region where the text segment "Zhang San" is located; Detection Result 2 includes: the regional coordinates of the initial text region where the text segment "Project Manager" is located; Detection Result 3 includes: the regional coordinates of the initial text region where the text segment "XX Digital Technology Co., Ltd." is located; Detection Result 4 includes: the regional coordinates of the initial text region where the text segment "Phone: 021 - 3800****" is located.

[0072] In some embodiments, a detection result includes the regional coordinates of an initial text region.

[0073] Among them, the initial text region refers to the text region detected by a detection algorithm for the same text segment. The initial text regions obtained by different detection algorithms for the same text segment can be the same or different. Exemplarily, if the multiple detection algorithms include: Detection Algorithm A, Detection Algorithm B, and Detection Algorithm C, and the target picture includes the text segment "Zhang San", where Detection Algorithm A obtains the first text region of the text segment "Zhang San" by identifying the target picture; Detection Algorithm B obtains the second text region of the text segment "Zhang San" by identifying the target picture; Detection Algorithm C obtains the third text region of the text segment "Zhang San" by identifying the target picture, then the first text region, the second text region, and the third text region are all called the initial text regions of the text segment "Zhang San".

[0074] In some embodiments, the regional coordinates of the initial text region where a text segment is located refer to the coordinates of the initial text region (such as a rectangular box) where the text segment is located. Exemplarily, the regional coordinates of the initial text region can be represented by (x i , y i ), where i represents the number of the four vertices of the rectangular box of each text region, and the value range of i is [1, 2, 3, 4].

[0075] Exemplarily, such as Figure 4 As shown, in the target image, the regional coordinates of the initial text area where the text segment is "Zhang San" can be expressed as: [(x1, y1), (x2, y2), (x3, y3), (x4, y4)], where the four coordinates respectively represent the coordinates of the four vertices of the rectangular box of the initial text area. For example, (x1, y1) represents the coordinate of the first vertex of the rectangular box of the initial text area.

[0076] In some embodiments, the multiple detection algorithms are OCR algorithms. Among them, OCR is a technology for recognizing optical characters through image processing and pattern recognition techniques, which can analyze and recognize image files of text materials to obtain text and layout information.

[0077] Exemplarily, the multiple detection algorithms may include:

[0078] The connectionist text proposal network (CTPN) algorithm. The advantages of the CTPN algorithm are: it can effectively detect horizontally distributed text in complex scenes and has a relatively fast detection speed. The disadvantages of the CTPN algorithm are: it can only detect horizontal or vertical (vertical text can be detected after changing the anchor ratio) text and cannot detect text in other directions.

[0079] The efficient and accuracy scene text (EAST) algorithm. The advantages of the EAST algorithm are: it uses the fusion detection of feature maps at multiple scales to detect text regions of different scales; the predicted text boxes are angled and can detect text in any direction; the detection process is simple and the detection speed is relatively fast. The disadvantages of the EAST algorithm are: due to the limitations of the receptive field and the size of the anchor, it is difficult to detect long text and curved text.

[0080] The SegLink algorithm. The advantages of the SegLink algorithm are: it can detect the entire text line at one time. The disadvantages of the SegLink algorithm are: it cannot detect text blocks with a large interval.

[0081] The rotation region proposal networks (RRPN) algorithm. The advantage of the RRPN algorithm is that it can achieve multi-directional text detection. The disadvantages of the RRPN algorithm are: the detection speed is slow and the computational complexity is large.

[0082] The Mask R-CNN algorithm has the advantages of enabling end-to-end detection and recognizing curved text. However, its disadvantage is that it performs poorly for Chinese text.

[0083] It can be seen that each of the multiple detection algorithms has its own advantages and disadvantages. Using only a single detection algorithm for detection will produce poor results. For example, the CTPN algorithm can only detect horizontal text. If there are text in multiple directions in the target image, using only the CTPN algorithm will not be able to detect text in other directions. Therefore, based on the method provided in the embodiments of the present application, the advantages of various detection algorithms can be fully utilized to achieve good detection results for the target image under the influence of various factors.

[0084] In some embodiments, before respectively identifying the target image based on multiple detection algorithms, the method further includes preprocessing the target image to facilitate subsequent recognition of the image by the detection algorithm.

[0085] Exemplarily, preprocessing may include cropping the original target image to obtain an optimally recognized image region of the target image, thereby reducing data dimensionality. Also, preprocessing may include binarizing the original target image to eliminate noise interference with image recognition.

[0086] In some embodiments, each of the plurality of detection algorithms corresponds to an algorithm number, and a detection result also includes the algorithm number corresponding to the detection result. Exemplarily, the region coordinates of the initial text region included in the detection result and the algorithm number corresponding to the detection result can be expressed in the following form: (x i ,y i ) n , where n represents the algorithm number.

[0087] S102: Obtain target detection results based on the detection results corresponding to each detection algorithm.

[0088] The target detection result includes the region coordinates of the target text area of ​​each text segment in the target image; the target text area of ​​a text segment is determined based on the overlap of multiple initial text areas where a text segment is located determined by multiple detection algorithms.

[0089] In some embodiments, a detection result further includes: an area number of an initial text area. For example, the area coordinates of the initial text area and the area number of the initial text area included in a detection result can be represented by (x mi ,y mi) It is represented that m represents the region number of the initial text region, and i represents the numbers of the four vertices of the rectangular frame of each text region, where the value range of i is [1, 2, 3, 4].

[0090] Exemplarily, as Figure 5 shown, in the target picture, the region coordinates of the initial text region with the text segment "Zhang San" can be represented as: [(x 11 , y 11 ), (x 12 , y 12 ), (x 13 , y 13 ), (x 14 , y 14 )], where the four coordinates respectively represent the four vertex coordinates of the rectangular frame of the initial text region with region number 1. For example, (x 11 , y 11 ) represents the coordinate of the first vertex of the rectangular frame of the initial text region with region number 1. The region coordinates of the initial text region with the text segment "Project Manager" can be represented as: [(x 21 , y 21 ), (x 22 , y 22 ), (x 23 , y 23 ), (x 24 , y 24 )], where the four coordinates respectively represent the four vertex coordinates of the rectangular frame of the initial text region with region number 2. For example, (x 21 , y 21 ) represents the coordinate of the first vertex of the rectangular frame of the initial text region with region number 2.

[0091] In some embodiments, the region numbers of the initial text regions where the same text segment is located are the same. It can be understood that the region coordinates of the initial text regions obtained by different detection algorithms for the same text segment may be different, but the region numbers are the same because the region number is used to represent the text region where the same text segment is located.

[0092] Exemplarily, if the multiple detection algorithms include detection algorithm A and detection algorithm B, then as Figure 6 shown, if a text segment in the target picture is "Zhang San", the detection result corresponding to detection algorithm A can be [(1 11 , 1 11 ), (3 12 [[ID=, 1 12 ), (1 13 , 2 13 ), (3 14 , 2 14)], the detection results corresponding to detection algorithm B can be [(2 11 , 1 11 ), (4 12 , 1 12 ), (2 13 , 2 13 ), (4 14 , 2 14 )]. It can be seen that for the same text segment "Zhang San", the regional coordinates of the initial text regions obtained by detection algorithm A and detection algorithm B can be different, but the region numbers are the same, both being 1.

[0093] Optionally, as Figure 7 shown, the above step S102 can be implemented as the following steps:

[0094] S1021. Group the detection results corresponding to each detection algorithm according to the region numbers of the initial text regions, and obtain multiple detection groups.

[0095] Among them, each detection group includes the regional coordinates of multiple initial text regions where the same text segment is located.

[0096] Exemplarily, if multiple detection algorithms include detection algorithm A, detection algorithm B, and detection algorithm C, and the target picture includes: text segment "Zhang San", text segment "Project Manager", then the initial text regions where the text segment "Zhang San" identified by detection algorithm A, the initial text regions where the text segment "Zhang San" identified by detection algorithm B, and the initial text regions where the text segment "Zhang San" identified by detection algorithm C are used as one detection group; the initial text regions where the text segment "Project Manager" identified by detection algorithm A, the initial text regions where the text segment "Project Manager" identified by detection algorithm B, and the initial text regions where the text segment "Project Manager" identified by detection algorithm C are used as another detection group.

[0097] S1022. For the first detection group among the multiple detection groups, determine the overlapping region whose overlapping times meet the first preset condition as the target text region of the first detection group.

[0098] Among them, the target detection result includes the target text regions of each detection group.

[0099] Optionally, the first detection group can be any one of the multiple detection groups.

[0100] In some embodiments, the first preset condition can be the overlapping region with the most overlapping times.

[0101] Exemplarily, if multiple detection algorithms include detection algorithm A, detection algorithm B, and detection algorithm C, and the target picture includes: a first text segment, then the first detection group includes the initial text region A where the first text segment is recognized by detection algorithm A, the initial text region B where the first text segment is recognized by detection algorithm B, and the initial text region C where the first text segment is recognized by detection algorithm C; then according to the regional coordinates of the initial text region A, the initial text region B, and the initial text region C, the overlapping region obtained is as Figure 8 shown. If the first preset condition is the overlapping region with the most overlapping times, then the target text region of the first detection group is the overlapping region with 3 overlapping times.

[0102] In some embodiments, according to the regional coordinates of the multiple initial text regions included in each detection group, the regional coordinates of the target text region of each detection group can be calculated. Exemplarily, the regional coordinates of the target text region can be expressed as (x ki , y ki ), where k represents the number of the detection group, and i represents the number of the four vertices of the rectangular frame of each text region, and the value range of i is [1, 2, 3, 4]. For example, the coordinates of the target text region can be expressed as: [(x 11 , y 11 ), (x 12 , y 12 ), (x 13 , y 13 ), (x 14 , y 14 )], where (x 11 , y 11 ) represents the coordinates of the first vertex of the rectangular frame of the target text region of the first detection group.

[0103] Exemplarily, as Figure 6 shown, if multiple detection algorithms include detection algorithm A and detection algorithm B, and a text segment included in the target picture is "Zhang San", then the first detection group can include: the regional coordinates of the initial text region A of the text segment "Zhang San" recognized by detection algorithm A: [(1 11 , 1 11 ), (३ 12 , 1 12 ), (1 13 , 2 13 ), (3 14 , 2 14 )], the regional coordinates of the initial text region B of the text segment "Zhang San" recognized by detection algorithm B: [(2 11 , 1 11 ), (4 12 , 1 12 ), (2 13 , 213 ), (4 14 , 2 14 )]. Based on the coordinates of the initial text area A and the initial text area B, the overlapping area is obtained as follows Figure 9 As shown, if the first preset condition is the overlapping area with the largest number of overlaps, the area coordinates of the target text area of ​​the first detection group can be [(1 11 , 1 11 ), (3 12 , 1 12 ), (2 13 , 2 13 ), (3 14 , 2 14 )].

[0104] In some embodiments, a detection result further includes: a confidence level corresponding to the region coordinates of an initial text region.

[0105] Confidence is also called reliability, confidence level, or confidence coefficient. It reflects the credibility of the test results.

[0106] In some embodiments, the region coordinates of the initial text region, the region number of the initial text region, and the confidence level corresponding to the region coordinates of the initial text region included in a detection result can be expressed in the following form: [(x mi ,y mi ), C m ], where C represents the coordinates of the initial text area (x mi ,y mi ) corresponding to the confidence level. For example, if a test result is {[(x 11 ,y 11 ), (x 12 ,y 12 ), (x 13 ,y 13 ), (x 14 ,y 14 )],C1}, then C1 represents the region coordinate [(x 11 ,y 11 ), (x 12 ,y 12 ), (x 13 ,y 13 ), (x 14 ,y 14 )] confidence level.

[0107] In some embodiments, when there are multiple target text areas in the first detection group, the confidence level corresponding to the area coordinates of each target text area is calculated, and the target text area with the highest confidence level is used as the final target text area.

[0108] The confidence level corresponding to the region coordinates of the target text region is the average value of the sum of the confidence levels corresponding to the region coordinates of the initial text regions (or valid text regions) constituting the target text region.

[0109] For example, Figure 10 As shown, if the target text area of ​​the first detection group includes: target text area 1 and target text area 2, wherein the target text area 1 is composed of the initial text area A (confidence is 0.8), the initial text area B (confidence is 0.6) and the initial text area C (confidence is 0.7); the target text area 2 is composed of the initial text area B (confidence is 0.6), the initial text area C (confidence is 0.7) and the initial text area D (confidence is 0.4), then the confidence of the target text area 1 is: (0.8+0.6+0.7) / 3=0.7; the confidence of the target text area 2 is: (0.6+0.7+0.45) / 3=0.58, so the target text area 1 is the final target text area.

[0110] In some embodiments, the above step S1022 can be implemented as the following steps:

[0111] Step a1: Determine a valid text area from a plurality of initial text areas included in the first detection group.

[0112] The confidence level corresponding to the region coordinates of the valid text region is greater than a first preset threshold.

[0113] It is understandable that when the confidence level corresponding to the region coordinates of the initial text region is low, it indicates that the initial text region may not cover the text segment and is an invalid text region.

[0114] For example, Figure 11 As shown, if the first detection group includes: initial text area A, initial text area B and initial text area C, where the confidence of the area coordinates of the initial text area C is less than the first preset threshold (for example, the first preset threshold can be 0.5), the initial text area C does not contain a text segment and is an invalid text area.

[0115] It can be seen that the overlapping area with the greatest number of overlaps among initial text area A, initial text area B, and initial text area C does not cover the text segment. Therefore, if initial text area C is also used to determine the target text area, the determined target text area may be invalid. Therefore, if the method provided in the embodiments of the present application is not used to exclude invalid text areas with low confidence, the determined target text area may also be invalid.

[0116] Step a2: Determine the target text area of ​​the first detection group from the valid text areas included in the first detection group.

[0117] In some embodiments, from the valid text areas included in the first detection group, the overlapping areas whose overlapping times meet the first preset condition are determined as the target text areas of the first detection group.

[0118] It can be understood that based on the solution provided in the embodiment of the present application, valid text areas with a confidence level greater than a first preset threshold are screened out from multiple initial text areas, and then the target text area is determined based on the valid text area. This can eliminate interference items (invalid text areas) and improve the accuracy of image recognition results.

[0119] As a possible implementation method, a method of intra-group voting is adopted to determine the target text area from the multiple initial text areas included in the first detection group.

[0120] Among them, the method of intra-group voting is to vote among the multiple initial text areas included in each detection group. When there are overlapping areas among the initial text areas, the overlapping area is covered by the rectangular frame of the initial text area once and gets one vote. The overlapping area with the highest number of votes is the target text area.

[0121] For example, it is assumed that the first detection group includes initial text area A, initial text area B and initial text area C, such as Figure 12 As shown, according to the area coordinates, the positions of the initial text area A, initial text area B and initial text area C on the image can be determined. Then, according to the voting rules, the number of votes for each overlapping area is as follows: Figure 12 As shown, the shaded area is the intersection of the three initial text areas, so it has the highest number of votes. Figure 12 The middle shaded area is the target text area of ​​the first detection group.

[0122] As another possible implementation, a valid text area is determined from a plurality of initial text areas included in the first detection group; and then, a target text area is determined from the valid text areas included in the first detection group by adopting an intra-group voting method.

[0123] Specifically, when there are overlapping areas among the valid text areas, each time the overlapping area is covered by the rectangular frame of the valid text area, it gets one vote, and the overlapping area with the highest number of votes becomes the target text area.

[0124] In some embodiments, the number of the multiple detection algorithms is an odd number. It is understandable that, since the embodiments of the present application determine the target text area based on the number of overlaps of the initial text areas (or based on the voting method within the group), in order to avoid the situation where the number of overlaps (or the number of votes) is the same, the number of the initial text areas is preferably an odd number, that is, the number of the multiple detection algorithms is an odd number.

[0125] In some embodiments, based on multiple influencing factors that have a greater impact on the detection effect, multiple algorithms that are good at processing the multiple influencing factors are selected as multiple detection algorithms.

[0126] For example, if poor lighting, low camera pixels, and shooting angle deviation are the three factors that have the greatest impact on the detection results when acquiring the target image, algorithm A that is good at processing scenes with poor lighting, algorithm B that is good at processing scenes with angle deviation, and algorithm C that is good at processing scenes with low pixels can be selected as multiple detection algorithms.

[0127] It is understandable that since each algorithm requires resources during execution, the more algorithms there are, the more resources are required. Therefore, in actual applications, the required detection algorithm can be selected based on the actual resource situation and factors affecting the detection effect.

[0128] S103: Identify the text content of each text segment from the target image according to the target detection result.

[0129] In some embodiments, a recognition algorithm is used to identify the text content of each text segment from the target image based on the target detection result. The recognition algorithm may be an OCR recognition algorithm.

[0130] Optional, such as Figure 13 As shown, the above step S103 can be implemented as the following steps:

[0131] S1031 : Identify the regional coordinates of the target text area of ​​each text segment based on multiple recognition algorithms and the regional coordinates of the target text area of ​​each text segment, and obtain a recognition result corresponding to each recognition algorithm.

[0132] In some embodiments, a recognition result corresponding to a recognition algorithm includes: initial text content of a target text area of ​​each text segment in the target image obtained based on the recognition algorithm. A recognition result includes the initial text content of a target text area of ​​a text segment.

[0133] The initial text content refers to the text content identified by different recognition algorithms for the same target text area. For example, if multiple recognition algorithms include: recognition algorithm A, recognition algorithm B, and recognition algorithm C, and the target detection result includes a first target text area, wherein recognition algorithm A obtains the first text content of the first target text area by recognizing the target detection result; recognition algorithm B obtains the second text content of the first target text area by recognizing the target detection result; and recognition algorithm C obtains the third text content of the first target text area by recognizing the target detection result, then the first text content, the second text content, and the third text content are all referred to as the initial text content of the first target text area.

[0134] S1032. Obtain a target recognition result according to the recognition result corresponding to each recognition algorithm.

[0135] The target recognition result includes the target text content of the target text area of ​​each text segment in the target image; the target text content of the target text area of ​​a text segment is determined based on the repetition of multiple initial text contents of the target text area of ​​a text segment determined by multiple recognition algorithms.

[0136] In some embodiments, a recognition result further includes: an area number of a target text area. For example, a recognition result includes the initial text content of a target text area and the area number of the target text area, which can be represented as V m , where V represents the initial text content of the target text area, m represents the area number of the target text area, and V m Represents the initial text content of the mth target text area.

[0137] In some embodiments, each of the plurality of recognition algorithms corresponds to an algorithm number, and a recognition result also includes the algorithm number corresponding to the recognition result. Exemplarily, the recognition result includes the initial text content of the target text area and the area number of the target text area, and the algorithm number corresponding to the recognition result can be expressed in the following form: (V m ) n , where n represents the algorithm number.

[0138] In some embodiments, the above step S1032 may be specifically implemented as the following steps:

[0139] Step b1: group the recognition results corresponding to each recognition algorithm according to the area number of the target text area to obtain multiple recognition groups.

[0140] Each recognition group includes a plurality of initial text contents in a target text area of ​​the same text segment.

[0141] Exemplarily, if multiple recognition algorithms include recognition algorithm A, recognition algorithm B, and recognition algorithm C, and the target detection result includes: a first target text region and a second target text region, then the initial text content of the first target text region recognized by recognition algorithm A, the initial text content of the first target text region recognized by recognition algorithm B, and the initial text content of the first target text region recognized by recognition algorithm C are taken as a recognition group; the initial text content of the second target text region recognized by recognition algorithm A, the initial text content of the second target text region recognized by recognition algorithm B, and the initial text content of the second target text region recognized by recognition algorithm C are taken as another recognition group.

[0142] Step b2: For the first recognition group among multiple recognition groups, determine the text content with the repetition times meeting the second preset condition from the multiple initial text contents of the target text regions included in the first recognition group as the target text content of the target text region of the first recognition group.

[0143] Among them, the target recognition result includes the target text content of the target text region of each recognition group.

[0144] Optionally, the first recognition group can be any one of the multiple recognition groups.

[0145] In some embodiments, the second preset condition can be the text content with the most repetition times.

[0146] Exemplarily, if multiple recognition algorithms include recognition algorithm A, recognition algorithm B, and recognition algorithm C, and the target detection result includes a first target text region, then the first recognition group includes the initial text content A of the first target text region recognized by recognition algorithm A, the initial text content B of the first target text region recognized by recognition algorithm B, and the initial text content C of the first target text region recognized by recognition algorithm C. If the initial text content A is "Zhang San", the initial text content B is "Zhang San", and the initial text content C is "Zhang Er", where "Zhang San" is repeated twice and "Zhang Er" is repeated once, if the second preset condition is the text content with the most repetition times, then the target text content of the first recognition group is "Zhang San".

[0147] In some embodiments, a recognition result further includes: the confidence corresponding to the initial text content of the target text region of a text segment.

[0148] Exemplarily, the initial text content of the target text region of a text segment included in a recognition result, the region number of the target text region, and the confidence corresponding to the initial text content of the target text region can be expressed in the following form: (V m , C m), where C represents the confidence level of the initial text content of the target text area, m represents the area number of the target text area, and C m Indicates the confidence level corresponding to the initial text content of the mth target text area.

[0149] In some embodiments, when there are multiple target text contents in the first recognition group, the confidence level corresponding to each target text content is calculated, and the target text content with the highest confidence level is used as the final target text content.

[0150] The confidence level corresponding to the target text content is the average value of the sum of the confidence levels corresponding to all initial text contents identical to the target text content in the first recognition group.

[0151] For example, if the initial text content in the first recognition group includes: [("Zhang San", 0.97), ("Zhang San", 0.95), ("Zhang San", 0.99), ("Zhang Er", 0.91), ("Zhang Er", 0.92), ("Zhang Er", 0.90), ("Zhang Yi", 0.87), ("Zhang Yi", 0.90)], then there are two target text contents in the first recognition group, namely: target text content 1: "Zhang San", "target text content 2: "Zhang Er", among which the confidence of target text content 1 is: (0.97+0.95+0.99) / 3=0.97; the confidence of target text content 2 is: (0.91+0.92+0.90) / 3=0.91, therefore, target text content 1 is the final target text content.

[0152] In some embodiments, the above step b2 can be implemented as the following steps:

[0153] Step b2-1: Determine valid text content from a plurality of initial text contents in the target text area included in the first recognition group.

[0154] The confidence level corresponding to the valid text content is greater than a second preset threshold.

[0155] It is understandable that when the confidence level corresponding to the initial text content is low, it indicates that the initial text content may not completely include all characters of the text segment and is invalid text content.

[0156] For example, if the first recognition group includes: initial text content A, initial text B and initial text content C, where the initial text content A is ("XX Digital Technology Co., Ltd.", 0.98), the initial text content B is ("XX Digital Technology Co., Ltd.", 0.99), and the initial text content C is ("XX", 0.51), it can be seen that the confidence of the initial text content C is low, and the initial text content C does not completely include all the characters of the text segment, and is invalid text content.

[0157] Step b2-2: Determine the target text content of the target text area of ​​the first recognition group from the valid text content of the target text area included in the first recognition group.

[0158] In some embodiments, from the valid text contents of the target text area included in the first recognition group, text contents whose repetition times meet the second preset condition are determined as the target text contents of the target text area of ​​the first recognition group.

[0159] It can be understood that based on the solution provided in the embodiment of the present application, valid text content with a confidence level greater than a second preset threshold is screened out from multiple initial text contents, and then the target text content is determined based on the valid text content. This can eliminate interference items (invalid text content) and improve the accuracy of image recognition results.

[0160] As a possible implementation method, a group voting method is adopted to determine the target text content from the multiple initial text contents included in the first recognition group.

[0161] The voting method within the group is to vote among the multiple initial text contents included in each recognition group. When there are repeated contents in each initial text content, each repetition gets one vote, and the repeated content with the highest number of votes is the target text content.

[0162] For example, if the initial text content in the first recognition group includes: ("Zhang San", "Zhang San", "Zhang San", "Zhang Er", "Zhang Er"), then according to the voting rules, the initial text content "Zhang San" gets 3 votes, and the initial text content "Zhang Er" gets 2 votes, so the initial text content "Zhang San" is the target text content of the first recognition group.

[0163] As another possible implementation, valid text content is determined from a plurality of initial text contents included in the first recognition group; and then, target text content is determined from the valid text contents included in the first recognition group by voting within the group.

[0164] Specifically, when there are duplicate contents among the valid text contents, each duplicate content receives one vote, and the duplicate content with the highest number of votes becomes the target text content.

[0165] In some embodiments, the number of the multiple recognition algorithms is an odd number. It is understandable that, since the embodiments of the present application determine the target text content based on the number of repetitions of the initial text content (or based on a voting method within a group), to avoid the situation where the number of repetitions (or the number of votes) is the same, the number of the initial text content is preferably an odd number, that is, the number of the multiple recognition algorithms is an odd number.

[0166] In some embodiments, based on multiple influencing factors that have a greater impact on the recognition effect, multiple algorithms that are good at processing the multiple influencing factors are selected as multiple recognition algorithms.

[0167] It is understandable that the method provided in the embodiments of the present application uses multiple algorithms to jointly identify image information, with multiple detection algorithms and multiple recognition algorithms all being used in parallel. First, the target image is identified based on the multiple detection algorithms, obtaining detection results corresponding to each detection algorithm. Then, a target detection result is obtained based on the detection results corresponding to each detection algorithm. Furthermore, based on the multiple recognition algorithms and the target detection results, the text content of each text segment is identified from the target image. Thus, compared to the prior art detection methods that use a single algorithm, the embodiments of the present application use a multi-algorithm joint recognition approach, simultaneously identifying the target image based on multiple detection algorithms and multiple recognition algorithms, fully leveraging the advantages of various detection algorithms and various recognition algorithms to achieve better recognition results for the target image regardless of the influence of various external factors. For example, if Algorithm A is good at recognizing horizontally distributed text, and Algorithm B is good at recognizing multi-directional text, then the multi-algorithm combination method provided in the embodiment of the present application is adopted, and the target image is recognized based on Algorithm A and Algorithm B at the same time, which can obtain a result with higher accuracy. For another example, if Algorithm C is good at processing scenes with poor lighting, and Algorithm D is good at processing scenes with angle deviation, then the multi-algorithm combination method provided in the embodiment of the present application is adopted, and the target image is recognized based on Algorithm C and Algorithm D at the same time, which can overcome the influence of light difference and angle deviation on the result. In this way, based on the method provided in the embodiment of the present application, it is possible to obtain a result with higher accuracy in various scenarios, effectively improving the accuracy and practicality of image information recognition.

[0168] In some embodiments, after identifying the text content of each text segment in the target image, corresponding entity tags can be matched for the text information in the target image based on a combination of matching rules and a deep learning model.

[0169] Specifically, first, the content of the text segment in the content set of the text segment is identified based on the entity matching rule to obtain a first recognition result; then, based on the first recognition result, it is determined whether it is necessary to perform pattern recognition on the content set of the text segment; when it is necessary to perform pattern recognition on the content set of the text segment, a language recognition model is used to identify the content of some text segments in the text segment set to obtain a second recognition result.

[0170] In this way, a final recognition result is obtained based on the first recognition result and the second recognition result, and complete image information is generated based on the final recognition result.

[0171] like Figure 14 As shown, the embodiment of the present application provides a picture information recognition device for performing the following Figure 3 The picture information recognition method shown in FIG. The picture information recognition device 300 includes: a detection module 301 , a processing module 302 and a recognition module 303 .

[0172] Detection module 301 is used to identify the target image based on multiple detection algorithms and obtain detection results corresponding to each detection algorithm; wherein the detection result corresponding to a detection algorithm includes: the regional coordinates of the initial text area where each text segment in the target image is located based on a detection algorithm; a detection result includes the regional coordinates of an initial text area.

[0173] Processing module 302 is used to obtain a target detection result based on the detection results corresponding to each detection algorithm; wherein the target detection result includes the regional coordinates of the target text area of ​​each text segment in the target image; the target text area of ​​a text segment is determined based on the overlap of multiple initial text areas determined by multiple detection algorithms for a text segment.

[0174] The recognition module 303 is configured to recognize the text content of each text segment from the target image according to the target detection result.

[0175] In one possible implementation, a detection result also includes: an area number of an initial text area; the area numbers of the initial text areas where the same text segment is located are the same; a processing module 302 is specifically used to group the detection results corresponding to each detection algorithm according to the area number of the initial text area to obtain multiple detection groups; wherein each detection group includes the area coordinates of multiple initial text areas where the same text segment is located; for a first detection group among the multiple detection groups, from the multiple initial text areas included in the first detection group, an overlapping area whose number of overlaps meets a first preset condition is determined as the target text area of ​​the first detection group; the target detection result includes the target text area of ​​each detection group.

[0176] In another possible implementation, a detection result also includes: a confidence level corresponding to the regional coordinates of an initial text area; a processing module 302, specifically used to determine a valid text area from multiple initial text areas included in the first detection group; the confidence level corresponding to the regional coordinates of the valid text area is greater than a first preset threshold; and determining a target text area of ​​the first detection group from the valid text areas included in the first detection group.

[0177] In another possible implementation, the recognition module 303 is specifically configured to identify the regional coordinates of the target text area of ​​each text segment based on multiple recognition algorithms and the regional coordinates of the target text area of ​​each text segment, and obtain a recognition result corresponding to each recognition algorithm; wherein the recognition result corresponding to one recognition algorithm includes: the initial text content of the target text area of ​​each text segment obtained based on one recognition algorithm; one recognition result includes the initial text content of the target text area of ​​one text segment; a target recognition result is obtained based on the recognition result corresponding to each recognition algorithm; wherein the target recognition result includes the target text content of the target text area of ​​each text segment; the target text content of the target text area of ​​one text segment is determined based on the repetition of multiple initial text contents of the target text area of ​​one text segment determined by multiple recognition algorithms.

[0178] In another possible implementation, a recognition result also includes: an area number of a target text area; a recognition module 303 is specifically used to group the recognition results corresponding to each recognition algorithm according to the area number of the target text area to obtain multiple recognition groups; wherein each recognition group includes multiple initial text contents of the target text area of ​​the same text segment; for the first recognition group among the multiple recognition groups, from the multiple initial text contents of the target text area included in the first recognition group, the text content with the number of repetitions that meets the second preset condition is determined as the target text content of the target text area of ​​the first recognition group; the target recognition result includes the target text content of the target text area of ​​each recognition group.

[0179] In another possible implementation, a recognition result also includes: a confidence level corresponding to the initial text content of a target text area of ​​a text segment; a recognition module 303 is specifically used to determine valid text content from multiple initial text contents of the target text area included in the first recognition group; the confidence level corresponding to the valid text content is greater than a second preset threshold; and the target text content of the target text area of ​​the first recognition group is determined from the valid text contents of the target text area included in the first recognition group.

[0180] In another possible implementation, the number of the multiple detection algorithms is an odd number; and / or the number of the multiple recognition algorithms is an odd number.

[0181] Of course, the image information recognition device 300 provided in the embodiment of the present application includes but is not limited to the above modules.

[0182] In the case of implementing the functions of the above-mentioned integrated modules in the form of hardware, the embodiment of the present application provides another structural diagram of the image information recognition device. Figure 15 As shown, the image information recognition device 400 includes: a processor 402 , a communication interface 403 , and a bus 404 . Optionally, the image information recognition device 400 may further include a memory 401 .

[0183] Processor 402 may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 402 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 402 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, or a combination of a DSP and a microprocessor.

[0184] The communication interface 403 is used to connect to other devices via a communication network, such as Ethernet, wireless access network, wireless local area network (WLAN), etc.

[0185] The memory 401 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0186] As a possible implementation, memory 401 can exist independently of processor 402. Memory 401 can be connected to processor 402 via bus 404 to store instructions or program codes. When processor 402 calls and executes the instructions or program codes stored in memory 401, the image information recognition method provided in the embodiment of the present application can be implemented.

[0187] In another possible implementation, the memory 401 may also be integrated with the processor 402 .

[0188] The bus 404 may be an extended industry standard architecture (EISA) bus, etc. The bus 404 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 15 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0189] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the electronic device can be divided into different functional modules to complete all or part of the functions described above.

[0190] Another embodiment of the present application provides a chip system for use in an image information recognition device. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via circuitry. The interface circuits are configured to receive signals from the memory of the image information recognition device and send signals to the processors, the signals including computer instructions stored in the memory. When the processors execute the computer instructions, the image information recognition device executes each step of the image information recognition method described in the above method embodiment.

[0191] Another embodiment of the present application further provides a computer-readable storage medium having computer instructions stored therein. When the computer instructions are executed on a computer, the computer executes each step of the image information recognition method shown in the above method embodiment.

[0192] Another embodiment of the present application further provides a computer program product, which includes computer instructions. When the computer instructions are executed on a computer, the computer executes each step of the image information recognition method shown in the above method embodiment.

[0193] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer execution instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more media that can be integrated. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).

[0194] The above is only a specific embodiment of the present application. Those skilled in the art may conceive of changes or substitutions based on the specific embodiment provided in this application, and all such changes or substitutions shall fall within the scope of protection of this application.

Claims

1. A method for recognizing image information, characterized in that: include: Identifying the target image based on multiple detection algorithms and obtaining detection results corresponding to each detection algorithm; wherein the detection result corresponding to each detection algorithm includes: regional coordinates of an initial text area where each text segment in the target image is located, obtained based on the detection algorithm; and each detection result includes the regional coordinates of an initial text area; Obtaining a target detection result based on the detection results corresponding to each detection algorithm; wherein the target detection result includes the region coordinates of a target text area for each text segment in the target image; the target text area for a text segment is determined based on the overlap of multiple initial text areas determined by the multiple detection algorithms where the text segment is located; and the target text area for a text segment is the intersection of the multiple initial text areas where the text segment is located; According to the target detection result, the text content of each text segment is identified from the target image.

2. The method according to claim 1, characterized in that A detection result also includes: an area number of an initial text area; the area numbers of the initial text areas where the same text segment is located are the same; Obtaining a target detection result according to the detection result corresponding to each detection algorithm includes: Grouping the detection results corresponding to each detection algorithm according to the region number of the initial text region to obtain multiple detection groups; wherein each detection group includes the region coordinates of multiple initial text regions where the same text segment is located; For the first detection group among the multiple detection groups, from the multiple initial text areas included in the first detection group, the overlapping area whose number of overlaps meets the first preset condition is determined as the target text area of ​​the first detection group; the target detection result includes the target text area of ​​each detection group.

3. The method according to claim 2, characterized in that A detection result also includes: a confidence level corresponding to the coordinates of an initial text area; The step of determining, from among the multiple initial text regions included in the first detection group, an overlapping region with the largest number of overlaps as the target text region of the first detection group includes: Determining a valid text area from the multiple initial text areas included in the first detection group; the confidence level corresponding to the area coordinates of the valid text area is greater than a first preset threshold; A target text area of ​​the first detection group is determined from valid text areas included in the first detection group.

4. The method according to claim 1, wherein The step of identifying the text content of each text segment from the target image according to the target detection result includes: identifying the regional coordinates of the target text area of ​​each text segment based on multiple recognition algorithms and the regional coordinates of the target text area of ​​each text segment, and obtaining a recognition result corresponding to each recognition algorithm; wherein the recognition result corresponding to each recognition algorithm includes: initial text content of the target text area of ​​each text segment obtained based on the recognition algorithm; and a recognition result includes the initial text content of the target text area of ​​a text segment; A target recognition result is obtained based on the recognition results corresponding to each recognition algorithm; wherein the target recognition result includes the target text content of the target text area of ​​each text segment; the target text content of the target text area of ​​a text segment is determined based on the repetition of multiple initial text contents of the target text area of ​​the text segment determined by the multiple recognition algorithms.

5. The method according to claim 4, characterized in that A recognition result also includes: a region number of the target text region; Obtaining a target recognition result according to the recognition result corresponding to each recognition algorithm includes: Grouping the recognition results corresponding to each recognition algorithm according to the region number of the target text region to obtain a plurality of recognition groups; wherein each recognition group includes a plurality of initial text contents of the target text region of the same text segment; For the first recognition group among the multiple recognition groups, from the multiple initial text contents of the target text area included in the first recognition group, the text content whose repetition times meet the second preset condition is determined as the target text content of the target text area of ​​the first recognition group; the target recognition result includes the target text content of the target text area of ​​each recognition group.

6. The method according to claim 5, characterized in that A recognition result also includes: a confidence level corresponding to the initial text content of a target text area of ​​a text segment; The step of determining, from among the multiple initial text contents of the target text area included in the first recognition group, text contents having a repetition number that meets a second preset condition as target text contents of the target text area of ​​the first recognition group includes: Determining valid text content from a plurality of initial text contents of the target text area included in the first recognition group; wherein the confidence level corresponding to the valid text content is greater than a second preset threshold; The target text content of the target text area of ​​the first recognition group is determined from the valid text content of the target text area included in the first recognition group.

7. A device for recognizing image information, characterized in that: include: a detection module, configured to identify a target image based on multiple detection algorithms and obtain detection results corresponding to each detection algorithm; wherein the detection result corresponding to each detection algorithm includes: region coordinates of an initial text region where each text segment in the target image is located, obtained based on the detection algorithm; and a detection result includes the region coordinates of an initial text region; a processing module configured to obtain a target detection result based on the detection results corresponding to each detection algorithm; wherein the target detection result includes the region coordinates of a target text area of ​​each text segment in the target image; the target text area of ​​a text segment is determined based on the overlap of multiple initial text areas determined by the multiple detection algorithms in which the text segment is located; and the target text area of ​​a text segment is the intersection of the multiple initial text areas in which the text segment is located; The recognition module is configured to recognize the text content of each text segment from the target image according to the target detection result.

8. The device according to claim 7, characterized in that A detection result further includes: an area number of an initial text area; initial text areas in the same text segment have the same area number; the processing module is specifically configured to group the detection results corresponding to each detection algorithm according to the area number of the initial text area to obtain multiple detection groups; wherein each detection group includes the area coordinates of multiple initial text areas in which the same text segment is located; for a first detection group among the multiple detection groups, an overlapping area whose number of overlaps meets a first preset condition is determined from among the multiple initial text areas included in the first detection group as a target text area for the first detection group; the target detection result includes the target text area of ​​each detection group; A detection result further includes: a confidence level corresponding to the region coordinates of an initial text region; the processing module is specifically configured to determine a valid text region from the multiple initial text regions included in the first detection group; the confidence level corresponding to the region coordinates of the valid text region is greater than a first preset threshold; and determining a target text region of the first detection group from the valid text regions included in the first detection group; The recognition module is specifically configured to respectively identify the regional coordinates of the target text area of ​​each text segment based on multiple recognition algorithms and the regional coordinates of the target text area of ​​each text segment, and obtain a recognition result corresponding to each recognition algorithm; wherein the recognition result corresponding to each recognition algorithm includes: the initial text content of the target text area of ​​each text segment obtained based on the recognition algorithm; a recognition result includes the initial text content of the target text area of ​​a text segment; a target recognition result is obtained based on the recognition results corresponding to each recognition algorithm; wherein the target recognition result includes the target text content of the target text area of ​​each text segment; the target text content of the target text area of ​​a text segment is determined based on the repetition of multiple initial text contents of the target text area of ​​the text segment determined by the multiple recognition algorithms; A recognition result further includes: an area number of a target text area; the recognition module is specifically configured to group the recognition results corresponding to each recognition algorithm according to the area number of the target text area to obtain a plurality of recognition groups; wherein each recognition group includes a plurality of initial text contents of the target text area of ​​the same text segment; for a first recognition group among the plurality of recognition groups, from the plurality of initial text contents of the target text area included in the first recognition group, text contents having a repetition number that meets a second preset condition are determined as target text contents of the target text area of ​​the first recognition group; the target recognition result includes the target text contents of the target text area of ​​each recognition group; A recognition result also includes: a confidence level corresponding to the initial text content of a target text area of ​​a text segment; the recognition module is specifically used to determine valid text content from multiple initial text contents of the target text area included in the first recognition group; the confidence level corresponding to the valid text content is greater than a second preset threshold; and the target text content of the target text area of ​​the first recognition group is determined from the valid text contents of the target text area included in the first recognition group.

9. A device for recognizing image information, characterized in that: include: one or more processors; one or more memories; Wherein, the one or more memories are used to store computer program codes, and the computer program codes include computer instructions. When the one or more processors execute the computer instructions, the image information recognition device executes the image information recognition method according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed on a computer, the computer is caused to execute the image information recognition method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • End-to-end identification method for scene text with random shape

    CN108549893A

  • Method and device for recognizing text in image, computer equipment and computer storage medium

    CN111291629A