A method, device, electronic device and readable storage medium for number identification
By calculating and correcting the inclination angle of the image in the numbering recognition method and segmenting it into a single character image for recognition, the problem of low recognition accuracy caused by lighting changes in the prior art is solved, and higher recognition accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202010767362.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-03
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2040-08-03
AI Technical Summary
The existing number identification method can easily lead to failure of tilt correction when the light changes, resulting in a low recognition accuracy.
By obtaining the numbered text area of the image, calculating its inclination angle and performing inclination correction, dividing it into a single character image for recognition, and finally outputting the positive character recognition result.
Improve the accuracy of number identification and enhance the reliability of identification results, especially in the case of light changes.
Smart Images

Figure CN114092941B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine vision, and more specifically, to a number recognition method, device, electronic device and readable storage medium. Background Art
[0002] When performing image recognition on serial numbers, the existing character area positioning methods mostly use template matching and general target detection to locate the character area. The detection only obtains the rectangular positive non-rotated rectangular frame of the target object where the characters are located in the entire image, and the target obtained by segmentation has a large degree of deflection. For example: With the development of artificial intelligence technology, image processing methods have begun to be used to identify lead seal serial numbers, but the current seal serial number is easily affected by light during the recognition process, which can lead to the failure of tilt correction. It can be seen that the current accuracy of identifying serial numbers is relatively low. Summary of the invention
[0003] The embodiments of the present application provide a number recognition method, device, electronic device and readable storage medium to solve the problem of low number recognition accuracy.
[0004] In a first aspect, an embodiment of the present application provides a number identification method, including:
[0005] Obtaining numbered text regions of an image to obtain a first image;
[0006] Calculating a tilt angle of the first image, and performing tilt correction on the first image according to the tilt angle to obtain a second image;
[0007] Segmenting the second image into characters and recognizing the segmented characters;
[0008] Output the positive result of the character.
[0009] In a second aspect, an embodiment of the present application further provides a device, including:
[0010] An acquisition unit, used for acquiring a numbered text area of an image to obtain a first image;
[0011] a first processing unit, configured to calculate a tilt angle of the first image, and perform tilt correction on the first image according to the tilt angle to obtain a second image;
[0012] A second processing unit, configured to segment characters of the second image and recognize the segmented characters;
[0013] An output unit is used to output the positive result of the character.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a program or instruction stored in the memory and running on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the number identification method disclosed in the first aspect of the embodiment of the present application.
[0015] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the number identification method disclosed in the first aspect of the embodiment of the present application are implemented.
[0016] Thus, in the embodiment of the present application, after the tilt correction is performed on the acquired image number text area, it is segmented into single character images for recognition, and the recognition result is judged and a positive character recognition result is output, thereby achieving the technical effect of improving the accuracy of number recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] Figure 1 It is a flowchart of a number identification method provided in an embodiment of the present application;
[0019] FIG. 2( a ) is a schematic diagram of an inversion result of number recognition provided in an embodiment of the present application;
[0020] FIG2( b ) is a schematic diagram of a positive result of number recognition provided in an embodiment of the present application;
[0021] Figure 3 It is a flowchart of another number identification method provided in an embodiment of the present application;
[0022] Figure 4 is an image containing numbered text provided in an embodiment of the present application;
[0023] Figure 5 It is a structural schematic diagram of a number identification device provided in an embodiment of the present application;
[0024] Figure 6 is a structural schematic diagram of another number identification device provided in an embodiment of the present application;
[0025] Figure 7 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0027] See also Figure 1 , Figure 1 is a flowchart of a number identification method provided in an embodiment of the present application, such as Figure 1 As shown, the following steps are included:
[0028] Step 101: Acquire a numbered text area of an image to obtain a first image.
[0029] Among them, the above-mentioned acquisition can be a pre-extraction of the numbered text area of the image, for example: extracting the numbered text area of the image through natural scene text detection (Detecting Text in Natural Image with Connectionist Text Proposal Network, CTPN), non-maximum suppression algorithm (Non-Maximum Suppression, NMS), maximally stable extremal region (Robust wide-baseline stereo from maximally stable extremal regions, MSER), stroke width change (Stroke Width Transform, SWT) and other methods.
[0030] The above-mentioned image may be any image containing a number text area to be identified, for example, a photographed seal number image, a photographed license plate number image, and the like.
[0031] The above numbering may include pure digital numbers and numbers in a specific style, for example, a number whose first two characters are specific letters followed by numbers.
[0032] Step 102: Calculate the tilt angle of the first image, and perform tilt correction on the first image according to the tilt angle to obtain a second image.
[0033] Among them, the above-mentioned tilt angle can be generated by the tilt of the numbered text area in the image, or it can be the tilt generated when the image is acquired, for example: the tilt angle generated when the image is acquired; the numbered text area itself has a tilt angle relative to the entire image, then a tilt angle will also be generated.
[0034] Among them, the above-mentioned tilt angle can be obtained in many ways, for example: determining the image tilt angle based on Radon transform, obtaining the tilt angle by obtaining the angle of the minimum circumscribed rotation rectangle of the character image, obtaining the tilt angle by calculating the maximum transformation angle of the vertical projection variance, etc.
[0035] The above-mentioned tilt correction can be achieved by converting the calculated tilt angle into a corresponding affine matrix, and then using the affine matrix to map the image, thereby achieving the tilt correction of the image.
[0036] Step 103: segment the second image into characters and recognize the segmented characters.
[0037] Among them, the above segmentation can be a precise segmentation of characters, for example: jump processing, vertical projection, initial contour segmentation, merging overlapping contour boxes, splitting oversized contour boxes, merging contour boxes, projection checking, character screening and repair, unifying character width, non-maximum suppression algorithm and other operation processes.
[0038] Of course, in this embodiment, the above segmentation includes but is not limited to the segmentation methods listed above. For example, the above segmentation can first use a simple contour segmentation method to calculate the connected domain of the binary sealed text and obtain the circumscribed rectangular frames of each character. When the contour segmentation does not obtain the corresponding number of characters, the method of contour segmentation plus binary image projection correction is used to accurately segment the characters.
[0039] Step 104: output the positive result of the character.
[0040] Among them, the above results may include both positive and negative situations. For example, after the above step 103, the character recognition results obtained include both positive and negative situations, as shown in Figure 2(a). Figure 2(a) is an inverted recognition result. The orientation of the character can be determined by combining the uniqueness of the character itself and the recognition confidence. Figure 2(a) is rotated 180° to obtain the positive result of number recognition as shown in Figure 2(b), thereby improving the accuracy and reliability of the recognition results.
[0041] In this embodiment, after the tilt correction is performed on the acquired image number text area, it is segmented into single character images for recognition, and the recognition result is judged and a positive character recognition result is output, thereby achieving the technical effect of improving the accuracy of number recognition.
[0042] See also Figure 3 , Figure 3 is a flow chart of another number identification method provided in an embodiment of the present application, such as Figure 3 As shown, the following steps are included:
[0043] Step 301: Acquire a numbered text area of an image to obtain a first image.
[0044] Optionally, step 301 may include:
[0045] Acquire an image including a number;
[0046] Locating the position of the numbered text in the image to obtain a plurality of images including the numbered text;
[0047] Repeated images in the multiple images are filtered out, and the filtered images are merged to obtain a first image including all numbered texts.
[0048] Among them, the above-mentioned acquired image can be an original image, and the position of the numbered text is located according to the above-mentioned image, so as to obtain a first image including all the numbered texts, for example: an area image including text is obtained from the photographed seal number image (original image), and the above-mentioned image including the text area eliminates the area image not including text compared with the original image.
[0049] Among them, the above-mentioned positioning can adopt a text detection model, for example: adopt the yolo_v3 text detection model, and perform character area positioning based on a model combined with natural scene text detection. It combines the fast and accurate recognition performance of yolo_v3 for small character targets with the unique performance of natural scene text detection for text sequence detection, and uses the text combination strategy to combine the obtained text sequence small boxes, which can accurately locate the area where only text exists.
[0050] Among them, the above filtering can filter out repeated images in the numbered text images obtained above, for example: first eliminate images with confidence levels less than a certain threshold, and then filter out repeated images through confidence score sorting and a standard non-maximum suppression algorithm.
[0051] This implementation can obtain an image including all numbered texts, so that the area where only texts exist can be accurately located, and the obtained text area image contains less background interference.
[0052] Step 302: Calculate the tilt angle of the first image, and perform tilt correction on the first image according to the tilt angle to obtain a second image.
[0053] The tilt correction may include corrections in the horizontal and vertical directions. For example, when calculating the first image, the horizontal tilt angle may be calculated first, and then the vertical tilt correction may be performed after the horizontal tilt correction.
[0054] Optionally, step 302 may include:
[0055] Calculating a horizontal tilt angle of the first image, and performing tilt correction on the first image according to the horizontal tilt angle to obtain a third image;
[0056] The vertical tilt angle of the third image is calculated, and the tilt correction is performed on the third image according to the vertical tilt angle to obtain a second image.
[0057] The horizontal tilt angle can be obtained by the Radon algorithm, for example, by finding the angle of the maximum projection value of the image through the method of fixed-direction projection superposition, so as to determine the image tilt angle. According to the characteristics of tilted text, when calculating the current angle Radon response value, it is necessary to determine whether it is greater than a certain threshold value t R , if greater than t R Only when the characters are superimposed on the image will the response value be 0 and the tilt correction will fail. In order to avoid the failure of Radon correction leading to the failure of horizontal tilt correction, the algorithm performs morphological processing on the binary image of the seal number, extracts the connected domain formed by the characters, and obtains the angle of its minimum circumscribed rotation rectangle. When the tilt angle obtained by Radon correction is too small, but the minimum circumscribed rotation rectangle of the character has a certain angle, the Radon correction is considered to be invalid, and the angle of the minimum circumscribed rotation rectangle is taken as its horizontal tilt angle. Finally, according to the calculated tilt angle, the affine matrix is calculated, and the affine transformation is used to correct the tilt of the text image.
[0058] Among them, the vertical tilt angle of the above-mentioned third image can be obtained by calculating the variance of the vertical projection. Since the vertically tilted text binary image is in the shape of a parallelogram after morphological processing, the three vertices of the parallelogram are selected as source points according to the vertical tilt angle, and the three vertices of the corrected rectangle are set as target points. The affine matrix is calculated, and the tilt is corrected using affine transformation.
[0059] In this implementation, the horizontal tilt angle obtained by Radon transformation and the angle of the minimum circumscribed tilted rectangle and the vertical tilt angle obtained by the variance of the vertical projection make the obtained horizontal tilt angle and vertical tilt angle more reliable, which can effectively solve the tilt problem, and considering that the text binary image is in the form of a parallelogram after morphological processing when vertically tilted, therefore, when correcting the vertical tilt, the three vertices of the parallelogram are used as source points, and the three vertices of the corrected rectangle are set as target points, and the affine matrix of the image is calculated, thereby reducing the problem of diamond distortion in the binary image of the acquired numbered character image due to the presence of tilt.
[0060] Optionally, calculating the horizontal tilt angle of the first image includes:
[0061] If the tilt angle of the Radon transform of the first image is greater than or equal to a preset threshold, obtaining a horizontal tilt angle of the first image according to the tilt angle of the Radon transform;
[0062] If the tilt angle of the Radon transform of the first image is less than a preset threshold, the horizontal tilt angle of the first image is obtained according to the angle of the minimum circumscribed rotated rectangle of the first image, wherein the binary image of the first image is morphologically processed to extract the connected domain formed by the characters therein to obtain the minimum circumscribed rotated rectangle of the first image.
[0063] Among them, the above-mentioned preset threshold can be obtained through empirical values. For example: when correcting the tilt angle of a tilted text image, it is determined whether the tilt angle of the Radon transform is greater than or equal to the preset threshold. If the image character pixels are too few, then the Radon transform will cause the horizontal tilt angle correction to fail. At this time, the binary image of the image is morphologically processed, the connected domain formed by its characters is extracted, and the angle of its minimum circumscribed rotation rectangle is obtained. The angle of the minimum circumscribed rotation rectangle is the horizontal tilt angle of the image.
[0064] The morphological processing may be to dilate the image characters. For example, after the characters in the binary image are dilated, adjacent characters will be connected into a whole after appropriate dilation processing, and the connected domain formed by the characters can be extracted.
[0065] In this implementation, the horizontal tilt angle is calculated by combining Radon transform with the minimum circumscribed rectangular frame, and the obtained tilt angle has higher accuracy. When calculating the tilt angle, different tilt angle calculation methods are used to select the most reasonable tilt angle, rather than simply relying on one tilt correction method. The obtained tilt angle has higher accuracy, which improves the adaptability of the tilt correction operation in different environments.
[0066] Step 303: segment the second image into characters and recognize the segmented characters.
[0067] Optionally, step 303 may include:
[0068] Calculate the text connected domain of the second binarized image;
[0069] According to the calculation, segment the second image characters to obtain a plurality of circumscribed rectangular frames;
[0070] Removing non-character rectangular frames from the plurality of circumscribed rectangular frames to obtain a plurality of character circumscribed rectangular frames;
[0071] Identify the characters in the rectangular box surrounding multiple characters after segmentation.
[0072] The above segmentation may adopt a contour segmentation method. For example, the connected domain of the binary image text may be segmented by the contour segmentation method to obtain the circumscribed rectangular frame of a single character.
[0073] Among them, the above-mentioned character segmentation of the second image may result in non-character outline frames, for example: non-character outline frames segmented out when interference or adhesion occurs in the binarization of the text image due to highlighting, image blur, etc., and for numbers whose first few digits are fixed English characters, their strong prior information is used to identify the characters and then remove the outline frames in front of the specific characters; for other types of seals, the search is conducted from the beginning to the end at the same time to remove the outline frames with lower confidence.
[0074] In addition, the character text in the character circumscribed rectangular frame may be in the forward direction or inverted direction. When the non-character rectangular frames in the plurality of circumscribed rectangular frames are removed, it can be determined whether the text is inverted.
[0075] The above characters may be unique in themselves. For example, when identifying a seal number, if the first few characters are letters, then the seal number with the highest confidence is first determined as the recognition result. If the following characters are numbers, the characters with the highest confidence among 0-9 are selected as the recognition result. If all the above characters are numbers, the recognition of all characters in the text box only takes the characters with the highest confidence among 0-9 as the recognition result.
[0076] In this implementation, the accuracy of character segmentation can be improved, the interference of non-characters in the image recognition process can be reduced, and the technical effect of improving the accuracy of number recognition can be achieved.
[0077] Step 304: output the positive result of the character.
[0078] Among them, step 304 can be understood as further processing of the recognition result of step 303. For example, the recognized character string may appear forward or inverted. If there is an inverted recognition result, it is necessary to adjust to obtain a forward result and then output the forward result.
[0079] Optionally, step 304 may include:
[0080] According to the obtained character recognition results, the character string obtained in the direction with the maximum sum of character confidences is selected as the output result; or,
[0081] According to the obtained character recognition result, the character direction is adjusted to the forward direction, and the obtained character string is used as the output result.
[0082] Among them, the above-mentioned character recognition results can be inverted. For example: after obtaining a single character recognition result, when the character string is all numbers, the sum of the confidences of the original direction and the direction rotated 180° of the character string recognition result are calculated respectively, and the recognition result obtained in the direction with the larger sum of character confidences is selected as the positive result of character recognition; in the case of special characters, the character string can be adjusted by judging the direction of the character to obtain a positive result of character recognition.
[0083] In this implementation, during the recognition process, by performing judgment processing on two situations of character string recognition and outputting a positive result, the adaptability and recognition accuracy of number recognition can be improved.
[0084] Optionally, after step 301, the method may further include the following steps:
[0085] Step 305: scaling and expanding the first image size.
[0086] Among them, the above-mentioned scaling can be to reduce the size of the image. For example, when the size of the image taken by a mobile phone is relatively large, the overly large image size is reduced, which is beneficial to reduce the time spent on subsequent image processing operations and improve the real-time performance of number recognition.
[0087] Among them, the above-mentioned expansion can be performed on the reduced image. For example, for an image of suitable size, in order to avoid interference from small high-grayscale areas at the corners during image processing operations, and to facilitate subsequent character segmentation, rectangular expansion can be performed as needed, and the image after tilt correction can be expanded to a certain extent.
[0088] In this implementation, scaling the image size can reduce the time consumption of subsequent image processing, and expanding the image can reduce the interference of the corner area on subsequent image operations, thereby improving the real-time performance and accuracy of number recognition.
[0089] Step 306: Obtain the maximum stable extreme value region of the first image after the scaling and expansion, and extract the first area image of the target position in the maximum stable extreme value region, where the target position is the character region of the first image before the scaling and expansion.
[0090] The maximum stable extreme value region can be a region containing numbered text. The image is binarized by a method similar to the watershed algorithm. In all binary images obtained, some connected regions in the image change very little or even do not change. These regions are the maximum stable extreme value regions. For example, in an image containing numbered text characters, the grayscale values of the numbered characters are the same, and the threshold value will not be covered for a period of time when it continues to increase. It will not be submerged until the threshold value rises to the grayscale value of the character itself, and the numbered text region in the image can be found.
[0091] In this implementation, the numbered text region in the original image is extracted through the maximum stable extremum region, so that the numbered text region can be quickly located.
[0092] Step 307: extracting a second region image having a preset color from the first region image.
[0093] The above-mentioned preset color can be the background color of the numbered characters, for example: Figure 4 As shown, image 401 includes the numbered text "123456". According to prior information, the background color of the target numbered text area is blue. Therefore, when blue is set as the preset color for extraction, the "456" area image 402 with a blue background color can be extracted for the next step of tilt correction.
[0094] In this implementation, the target character area is accurately extracted by combining the maximum stable extreme value area with the background color prior information, thereby achieving the technical effect of improving the accuracy of number recognition.
[0095] Step 308: Calculate the tilt angle of the second region image, and perform tilt correction on the second region image according to the tilt angle to obtain a second image.
[0096] Among them, the above-mentioned tilt angle may include a horizontal tilt angle and a vertical tilt angle, for example: by calculating the tilt angles of the image in the horizontal and vertical directions, converting the angles into corresponding affine matrices, and then using the affine matrix to map the image, thereby achieving image tilt correction.
[0097] Among them, the second image can be obtained after the second area image is tilted and corrected. For example, when calculating the horizontal tilt angle, the second area image binary image can be morphologically processed to extract the connected domain formed by its characters and obtain the angle of its minimum circumscribed rotation rectangle; when calculating the vertical tilt angle, the second area image is horizontally tilted and corrected. Since the vertically tilted text binary image is in the form of a parallelogram after morphological processing, the three vertices of the parallelogram are selected as source points according to the vertical tilt angle, and the three vertices of the corrected rectangle are set as target points, and the affine matrix is calculated, and the tilt is corrected using affine transformation. In addition, the tilt angle of the second area image can be consistent with the tilt angle of the first image, and the number recognition result of the second area image obtained by processing the first image has a higher accuracy rate.
[0098] In this implementation, different tilt angle calculation methods are used when calculating the tilt angle, and the most reasonable tilt angle is selected instead of relying solely on one tilt correction method, thereby improving the adaptability of the tilt correction operation in different environments.
[0099] See also Figure 5 , Figure 5 is a schematic diagram of a structure of a number identification device provided in an embodiment of the present application, such as Figure 5 As shown, the device 500 includes:
[0100] An acquisition unit 501 is used to acquire a numbered text area of an image to obtain a first image;
[0101] A first processing unit 502 is used to calculate a tilt angle of the first image, and perform tilt correction on the first image according to the tilt angle to obtain a second image;
[0102] The second processing unit 503 is used to segment the second image into characters and recognize the segmented characters;
[0103] The output unit 504 is used to output the positive result of the character.
[0104] Optionally, the acquisition unit 501 may be used to:
[0105] Acquire an image including a number;
[0106] Locating the position of the numbered text in the image to obtain a plurality of images including the numbered text;
[0107] Repeated images in the multiple images are filtered out, and the filtered images are merged to obtain a first image including all numbered texts.
[0108] Optionally, the first processing unit 502 may be configured to:
[0109] Calculating a horizontal tilt angle of the first image, and performing tilt correction on the first image according to the horizontal tilt angle to obtain a third image;
[0110] The vertical tilt angle of the third image is calculated, and the tilt correction is performed on the third image according to the vertical tilt angle to obtain a second image.
[0111] Optionally, calculating the horizontal tilt angle of the first image includes:
[0112] If the tilt angle of the Radon transform of the first image is greater than or equal to a preset threshold, obtaining a horizontal tilt angle of the first image according to the tilt angle of the Radon transform;
[0113] If the tilt angle of the Radon transform of the first image is less than a preset threshold, the horizontal tilt angle of the first image is obtained according to the angle of the minimum circumscribed rotated rectangle of the first image, wherein the binary image of the first image is morphologically processed to extract the connected domain formed by the characters therein to obtain the minimum circumscribed rotated rectangle of the first image.
[0114] Optionally, the second processing unit 503 may be configured to:
[0115] Calculate the text connected domain of the second binarized image;
[0116] According to the calculation, segment the second image characters to obtain a plurality of circumscribed rectangular frames;
[0117] Removing non-character rectangular frames from the plurality of circumscribed rectangular frames to obtain a plurality of character circumscribed rectangular frames;
[0118] Identify the characters in the rectangular box surrounding multiple characters after segmentation.
[0119] Optionally, the output unit 504 may be used to:
[0120] According to the obtained character recognition results, the character string obtained in the direction with the maximum sum of character confidences is selected as the output result; or,
[0121] According to the obtained character recognition result, the character direction is adjusted to the forward direction, and the obtained character string is used as the output result.
[0122] Optional, such as Figure 6 As shown, the above device 500 may further include:
[0123] A third processing unit 505 is used to scale and expand the size of the first image;
[0124] A fourth processing unit 506 is configured to obtain a maximum stable extreme value region of the first image after scaling and expansion, and extract a first region image of a target position in the maximum stable extreme value region, wherein the target position is a character region of the first image before scaling and expansion;
[0125] A fifth processing unit 507, configured to extract a second region image having a preset color from the first region image;
[0126] The first processing unit 502 may be configured to:
[0127] The tilt angle of the second region image is calculated, and the tilt correction is performed on the second region image according to the tilt angle to obtain a second image.
[0128] The device 500 can realize Figures 1 to 4 The various processes implemented by the device in the method embodiment are not described here in detail to avoid repetition. The device 500 can achieve the technical effect of improving the accuracy of number recognition.
[0129] See also Figure 7 , Figure 7 is a schematic diagram of the structure of an electronic device provided by the present application, such as Figure 7 As shown, the electronic device 700 includes: a processor 701 , a memory 702 , and a program or instruction stored in the memory 702 and executable on the processor 701 .
[0130] It can be understood that the memory 702 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 702 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0131] In the embodiment of the present application, by calling the program or instruction stored in the memory 702, the processor 701 is used to:
[0132] Obtaining numbered text regions of an image to obtain a first image;
[0133] Calculating a tilt angle of the first image, and performing tilt correction on the first image according to the tilt angle to obtain a second image;
[0134] Segmenting the second image into characters and recognizing the segmented characters;
[0135] Output the positive result of the character.
[0136] The method disclosed in the above embodiment of the present application can be applied to the processor 701, or implemented by the processor 701. The processor 701 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 701. The above processor 701 can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor are combined and executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 702, and the processor 701 reads the information in the memory 702 and completes the steps of the above method in combination with its hardware.
[0137] It is understood that the embodiments described herein may be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), general purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in the present application, or a combination thereof.
[0138] For software implementation, the techniques described herein can be implemented by modules (e.g., procedures, functions, etc.) that perform the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0139] Optionally, the step of acquiring the numbered text area of the image to obtain the first image may include:
[0140] Acquire an image including a number;
[0141] Locating the position of the numbered text in the image to obtain a plurality of images including the numbered text;
[0142] Repeated images in the multiple images are filtered out, and the filtered images are merged to obtain a first image including all numbered texts.
[0143] Optionally, after acquiring the numbered text area of the image to obtain the first image, the processor 701 may also be configured to:
[0144] Scaling and expanding the first image size;
[0145] Acquire a maximum stable extreme value region of the first image after the scaling and expansion, and extract a first region image of a target position in the maximum stable extreme value region, wherein the target position is a character region of the first image before the scaling and expansion;
[0146] Extracting a second area image having a preset color from the first area image;
[0147] The calculating the tilt angle of the first image and performing tilt correction on the first image according to the tilt angle to obtain the second image includes:
[0148] The tilt angle of the second region image is calculated, and the first image is tilt-corrected according to the tilt angle to obtain a second image.
[0149] Optionally, calculating the tilt angle of the first image and performing tilt correction on the first image according to the tilt angle to obtain the second image may include:
[0150] Calculating a horizontal tilt angle of the first image, and performing tilt correction on the first image according to the horizontal tilt angle to obtain a third image;
[0151] The vertical tilt angle of the third image is calculated, and the tilt correction is performed on the third image according to the vertical tilt angle to obtain a second image.
[0152] Optionally, calculating the horizontal tilt angle of the first image may include:
[0153] If the tilt angle of the Radon transform of the first image is greater than or equal to a preset threshold, obtaining a horizontal tilt angle of the first image according to the tilt angle of the Radon transform;
[0154] If the tilt angle of the Radon transform of the first image is less than a preset threshold, the horizontal tilt angle of the first image is obtained according to the angle of the minimum circumscribed rotated rectangle of the first image, wherein the binary image of the first image is morphologically processed to extract the connected domain formed by the characters to obtain the minimum circumscribed rotated rectangle of the first image.
[0155] Optionally, the step of segmenting the second image into characters and identifying the segmented characters may include:
[0156] Calculate the text connected domain of the second binarized image;
[0157] According to the calculation, segment the second image characters to obtain a plurality of circumscribed rectangular frames;
[0158] Removing non-character rectangular frames from the plurality of circumscribed rectangular frames to obtain a plurality of character circumscribed rectangular frames;
[0159] Identify the characters in the rectangular box surrounding multiple characters after segmentation.
[0160] Optionally, the positive result of outputting the character may include:
[0161] According to the obtained character recognition results, the character string obtained in the direction with the maximum sum of character confidences is selected as the output result; or,
[0162] According to the obtained character recognition result, the character direction is adjusted to the forward direction, and the obtained character string is used as the output result.
[0163] The electronic device 700 can implement each process of the above-mentioned number recognition method embodiment and can achieve the same technical effect, so it will not be described here to avoid repetition. The electronic device 700 can achieve the technical effect of improving the accuracy of number recognition.
[0164] It should be noted that the electronic devices in the embodiments of the present application include mobile electronic devices and non-mobile electronic devices.
[0165] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned number identification method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0166] The processor is a processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0167] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0168] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0169] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A number identification method, characterized in that: include: Obtaining numbered text regions of an image to obtain a first image; Calculating a tilt angle of the first image, and performing tilt correction on the first image according to the tilt angle to obtain a second image; Segmenting the second image into characters and recognizing the segmented characters; Output the positive result of the character; The calculating the tilt angle of the first image and performing tilt correction on the first image according to the tilt angle to obtain the second image comprises: If the tilt angle of the Radon transform of the first image is greater than or equal to a preset threshold, the horizontal tilt angle of the first image is obtained according to the tilt angle of the Radon transform; if the tilt angle of the Radon transform of the first image is less than the preset threshold, the horizontal tilt angle of the first image is obtained according to the angle of the minimum circumscribed rotation rectangle of the first image, wherein the binary image of the first image is morphologically processed, and the connected domain formed by the characters is extracted to obtain the minimum circumscribed rotation rectangle of the first image; the first image is tilt-corrected according to the horizontal tilt angle to obtain a third image; The vertical tilt angle of the third image is calculated, and the tilt correction is performed on the third image according to the vertical tilt angle to obtain a second image.
2. The method according to claim 1, characterized in that The acquiring of the numbered text area of the image to obtain the first image comprises: Acquire an image including a number; Locating the position of the numbered text in the image to obtain a plurality of images including the numbered text; Repeated images in the multiple images are filtered out, and the filtered images are merged to obtain a first image including all the numbered texts.
3. The method according to claim 1, characterized in that After acquiring the numbered text area of the image to obtain the first image, the method further includes: Scaling and expanding the first image size; Acquire a maximum stable extreme value region of the first image after the scaling and expansion, and extract a first region image of a target position in the maximum stable extreme value region, wherein the target position is a character region of the first image before the scaling and expansion; Extracting a second area image having a preset color from the first area image; The calculating the tilt angle of the first image and performing tilt correction on the first image according to the tilt angle to obtain the second image includes: The tilt angle of the second region image is calculated, and the tilt correction is performed on the second region image according to the tilt angle to obtain a second image.
4. The method according to claim 1, characterized in that The segmenting of the second image into characters and identifying the segmented characters comprises: Calculate the text connected domain of the second binarized image; According to the calculation, segment the second image characters to obtain a plurality of circumscribed rectangular frames; Removing non-character rectangular frames from the plurality of circumscribed rectangular frames to obtain a plurality of character circumscribed rectangular frames; Identify the characters in the rectangular box surrounding multiple characters after segmentation.
5. The method according to claim 1, characterized in that The positive result of outputting the character includes: According to the obtained character recognition results, the character string obtained in the direction with the maximum sum of character confidences is selected as the output result; or, According to the obtained character recognition result, the character direction is adjusted to the forward direction, and the obtained character string is used as the output result.
6. A number recognition device, characterized in that: The device comprises: An acquisition unit, used for acquiring a numbered text area of an image to obtain a first image; a first processing unit, configured to calculate a tilt angle of the first image, and perform tilt correction on the first image according to the tilt angle to obtain a second image; A second processing unit, configured to segment the second image into characters and recognize the segmented characters; An output unit, used for outputting the positive result of the character; The first processing unit is used for: If the tilt angle of the Radon transform of the first image is greater than or equal to a preset threshold, the horizontal tilt angle of the first image is obtained according to the tilt angle of the Radon transform; if the tilt angle of the Radon transform of the first image is less than the preset threshold, the horizontal tilt angle of the first image is obtained according to the angle of the minimum circumscribed rotation rectangle of the first image, wherein the binary image of the first image is morphologically processed, and the connected domain formed by the characters therein is extracted to obtain the minimum circumscribed rotation rectangle of the first image; the first image is tilt-corrected according to the horizontal tilt angle to obtain a third image; The vertical tilt angle of the third image is calculated, and the tilt correction is performed on the third image according to the vertical tilt angle to obtain a second image.
7. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and running on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the number recognition method according to any one of claims 1 to 5.
8. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the number recognition method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Method for detecting defects of cosmetic paper label
CN108548820A
License plate recognition method, device and system
CN108985137A
A method and a system for detecting and recognizing traffic signs
CN109255279A