Character recognition device and character recognition method
The method addresses distorted character recognition on steel plates by converting images to a frontal view using rectangular information, ensuring accurate character identification through a two-step recognition process.
Patent Information
- Application Number
- JP2022050722
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-10-23
- Estimated Expiration
- 2042-03-25
AI Technical Summary
Existing character recognition technologies struggle with distorted images captured from varying angles, leading to inaccurate identification of characters on objects like steel plates, especially when the camera and object are not aligned, resulting in unrecognizable or incorrect character recognition.
A character recognition method that utilizes a deep learning model to first identify rectangular information of characters, then uses this information to convert the distorted image into a frontal view, allowing a second character recognition step to accurately identify characters by aligning the image based on horizontal and vertical directions.
Ensures high accuracy in character recognition even with distorted images by converting them to a frontal view, enabling reliable identification of characters on steel plates using a normal character recognition engine.
Smart Images

Figure 0007758951000003 
Figure 0007758951000004 
Figure 0007758951000005
Abstract
Description
[Technical Field]
[0001] The present invention relates to a character recognition device and a character recognition method for recognizing characters displayed on the surface of an object such as a steel plate. [Background technology]
[0002] On steel products such as thick plates, various types of information such as management information are displayed on the surface using a stencil, by attaching a label with the information printed on it, or by carving an inscription with the information, and this information is sometimes used for inventory management (to identify each product and confirm the location of the product).
[0003] A technology for identifying information displayed on such objects such as thick plates is described, for example, in Patent Document 1. Patent Document 1 discloses that an image of a product is displayed on a screen on a crane machine, and after masking part or all of the product number related to the displayed product, the crane operator visually confirms the serial number and then inputs the confirmed product number by voice or text.
[0004] In recent years, many industrial fields have been actively using classifiers trained by machine learning methods to extract experience and knowledge from huge amounts of data and use this knowledge to automate processes. In particular, in the field of image recognition, the use of classifiers based on neural networks, including deep learning, has dramatically improved classification accuracy for image recognition problems such as image classification and object detection, which have long been considered important issues.
[0005] For example, Non-Patent Document 1 discloses a technology that applies a deep learning model to character recognition of serial numbers and the like. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Patent No. 6003918 [Patent Document 2] Japanese Patent Application Laid-Open No. 2012-063869 [Non-patent literature]
[0007] [Non-Patent Document 1] Shinichiro Omachi, "Relay Commentary: The Potential of Machine Learning <<Part 5>> Machine Learning and Character Detection and Recognition: Detecting and Recognizing Text in the Environment," Measurement and Control, Vol. 58, No. 8, August 2019 Summary of the Invention [Problem to be solved by the invention]
[0008] However, the technology described in Patent Document 1 is based on the premise that the crane operator can visually recognize the serial number. Therefore, if there is variation in the angle at which the crane holds the object, or if the product surface is scratched or dirty, there is a risk that all or part of the information displayed on the product, such as the serial number, may not be recognizable. In addition, there is a problem in that the crane operator cannot identify the information displayed on the product unless he or she is near the product.
[0009] Therefore, it has been desired to capture an image of information displayed on a product (object) and improve the accuracy of identifying information contained in the captured image obtained by capturing the image.
[0010] Furthermore, even if a deep learning model is used to recognize characters printed with a stencil, characters on a label attached to the surface of an object such as a thick plate, or characters engraved on the surface of a product from a captured image, the accuracy of the character recognition may be poor.
[0011] This is due to the following reason. For example, when capturing an image of characters displayed on the surface of an object for character recognition, if the positional relationship between the camera and the object is not constant, the optical axis of the camera will not be directly aligned with the surface of the object, resulting in a distorted image of the character area displayed on the surface of the object. Performing character recognition in a distorted captured image increases the difficulty of character recognition, resulting in not only recognizable characters but also unrecognizable characters. This poses a problem in that the results of this character recognition cannot be used as the final character recognition results as is.
[0012] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a character recognition device and a character recognition method that can obtain sufficient recognition accuracy when reading characters displayed on the surface of an object such as a steel product using a deep learning model, even under conditions where the image is distorted. [Means for solving the problem]
[0013] As mentioned above, when performing character recognition under conditions where the image is distorted, even if it is not possible to determine the type of character the character is, it has been found that there is a high possibility of recognition if the information is limited to the "rectangular information" corresponding to the position and size information of the character to be recognized (i.e., information corresponding to a rectangular area that indicates the position and size of the character; a bounding box).
[0014] The inventors then came up with the idea that if they could use this rectangular information to convert the distorted captured image into an undistorted image facing the camera directly, then even a normal character recognition engine would be able to correctly recognize characters.
[0015] Specifically, the present invention is as follows.
[0016] In order to solve the above problem, according to one aspect of the present invention, Steel A character recognition device for recognizing characters displayed on the surface of the Steel When the positional relationship with Steel an imaging unit that captures an image of the surface of the Steel a first character recognition unit that performs character recognition based on the captured image; and a processing unit that recognizes characters displayed on the surface of the captured image. Among the characters displayed on the surface of the steel material The character recognized by the first character recognition unit Character rectangle information when the first character recognition unit acquires rectangular information of a predetermined number of characters in one line, each consisting of a font of the same size, and a predetermined number of characters in another line, each consisting of a font of the same size, the first character recognition unit grasps a parallelogram consisting of four points, the position of the rectangular information of the first character in one line, the position of the rectangular information of the predetermined number of characters in one line, the position of the rectangular information of the first character in the other line, and the position of the rectangular information of the predetermined number of characters in the other line, and of the two sets of two sides that are parallel and opposite to each other and included in the parallelogram, the straight lines of one set are regarded as straight lines corresponding to the horizontal direction of the characters displayed on the surface of the steel material, and the straight lines of the other set are regarded as straight lines corresponding to the vertical direction of the characters displayed on the surface of the steel material, The aforementioned Steel a straight line corresponding to the horizontal direction of the characters displayed on the surface of said Steel a line acquisition unit that acquires a line corresponding to the horizontal direction and a line corresponding to the vertical direction of the character displayed on the surface of the character; and Steel The character recognition device has a frontal facing conversion unit that converts the captured image to generate a converted image so that the characters displayed on the surface of the image are viewed from the front, and a second character recognition unit that performs character recognition based on the converted image, and outputs the characters recognized by the second character recognition unit.
[0017] In order to solve the above problems, according to another aspect of the present invention, Using a character recognition device, A character recognition method for recognizing characters displayed on a surface of a The imaging unit The aforementioned Steel When the positional relationship with Steel an imaging step of imaging the surface of the object to generate an image; The first character recognition unit a first character recognition step of performing character recognition based on the captured image; Among the characters displayed on the surface of the steel material The character recognized by the first character recognition unit Character rectangle information when the first character recognition unit acquires rectangular information of a predetermined number of characters in one line, each consisting of a font of the same size, and a predetermined number of characters in another line, each consisting of a font of the same size, the first character recognition unit grasps a parallelogram consisting of four points, the position of the rectangular information of the first character in one line, the position of the rectangular information of the predetermined number of characters in one line, the position of the rectangular information of the first character in the other line, and the position of the rectangular information of the predetermined number of characters in the other line, and of the two sets of two sides that are parallel and opposite to each other and included in the parallelogram, the straight lines of one set are regarded as straight lines corresponding to the horizontal direction of the characters displayed on the surface of the steel material, and the straight lines of the other set are regarded as straight lines corresponding to the vertical direction of the characters displayed on the surface of the steel material, The aforementioned Steel a straight line corresponding to the horizontal direction of the characters displayed on the surface of said Steel a line acquisition step of acquiring a line corresponding to the vertical direction of the character displayed on the surface of the The facing transformation unit is Based on the straight line corresponding to the horizontal direction and the straight line corresponding to the vertical direction acquired in the straight line acquisition step, Steel a frontal facing transformation step for generating a transformed image by transforming the captured image so that the characters displayed on the surface of the image are viewed from the front; The second character recognition unitand a second character recognition step of performing character recognition based on the converted image. [Effects of the Invention]
[0018] According to the present invention, when character recognition is performed using a deep learning model based on a captured image in a situation where the captured image is distorted, even a normal character recognition engine can recognize characters with sufficient recognition accuracy. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a diagram illustrating a relationship between a character recognition device according to an embodiment of the present invention and an object. [Figure 2] 10A and 10B are diagrams showing a display form of management information on the surface of an object according to an embodiment of the present invention. [Figure 3] FIG. 10 is a diagram showing an example of a captured image of characters displayed on the surface of an object, captured from a position facing the object. [Figure 4] FIG. 10 is a diagram showing an example of a captured image of characters displayed on the surface of an object captured from a distorted position. [Figure 5] FIG. 2 is a diagram illustrating a configuration of a processing unit according to an embodiment of the present invention. [Figure 6] 5A and 5B are diagrams for explaining acquisition of a straight line by a straight line acquisition unit according to an embodiment of the present invention; [Figure 7] 3 is a flowchart illustrating processing in the character recognition device according to the embodiment of the present invention. [Figure 8] FIG. 2 is a diagram illustrating an example of a hardware configuration of a processing unit according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0020] A character recognition device and a character recognition method according to an embodiment of the present invention will be described below with reference to the drawings.
[0021] A character recognition device according to an embodiment of the present invention will be described with reference to Fig. 1. Fig. 1 shows the relationship between a character recognition device 100 according to an embodiment of the present invention and an object 1.
[0022] The character recognition device 100 is a device that recognizes characters displayed on the surface of an object 1, and includes at least an imaging unit 3 and a calculation processing unit 4.
[0023] As shown in FIG. 1(A), the character recognition device 100 may be a mobile terminal such as a smartphone or tablet that incorporates an imaging unit 3 and an arithmetic processing unit 4. Alternatively, as shown in FIG. 1(B), the character recognition device 100 may be configured with an imaging unit 3 and an arithmetic processing unit 4 that are separate and connected to each other. In this case, it is assumed that the object 1 is held by a holding unit 2 such as a crane. Alternatively, the character recognition device 100 may be an unmanned aerial vehicle (drone) that includes the imaging unit 3 and the arithmetic processing unit 4 and flies within a factory. In either case, the positional relationship between the object 1 and the imaging unit 3 is not constant for each object 1, and the positional relationship between the object 1 and the imaging unit 3 varies.
[0024] In the following description, unless otherwise specified, a case will be described in which a mobile terminal is used as the character recognition device 100.
[0025] The object 1 is an object on which characters to be recognized by the character recognition device 100 are displayed. As the object 1, various objects (products) can be used as appropriate as long as they have character information such as management information displayed on their surface. The object 1 may be, for example, various thick plates (e.g., thick steel plates) manufactured in steelworks, other steel materials such as thin plates, steel pipes, and wire rods, resin materials, wood, etc.
[0026] In the following explanation, unless otherwise specified, a thick plate will be used as the object 1, and an example will be given in which character recognition is performed on management information for managing thick plates in a factory where various processes such as heat treatment and cutting are performed on the thick plates until the thick plates are shipped to thick plate consumers.
[0027] The flow of shipping heavy plates includes, for example, that heavy plates are manufactured at a steelworks and loaded onto a transport ship. The transport ship docks at a quay near a post-processing plant, and the heavy plates are unloaded from the transport ship and loaded onto a transport vehicle (truck, etc.). After being transported by the transport vehicle, the heavy plates arrive at the processing plant and are placed in a storage area within the post-processing plant. After being placed in the storage area, the heavy plates are subjected to processes such as heat treatment and cutting (gas cutting) when the time is right, and then placed in a warehouse. The heavy plates are then shipped from the post-processing plant to consumers.
[0028] In such cases, for management purposes, it is necessary to identify each plate at any time between when the plate is produced at the steelworks and when it is shipped to a customer. For this reason, management information is displayed on the surface of the plate by printing it with a stencil, by attaching a label with the management information printed on it, or by engraving the management information on the surface, so that it can be used for product management, etc.
[0029] FIG. 2 shows a display form of management information on the surface of the object 1 according to an embodiment of the present invention.
[0030] FIG. 2(A) shows a first example of a display form for a display area in which management information on the surface of a thick plate, which is the object 1, is displayed.
[0031] 2A, display area 101 is a display area where management information is printed using a stencil, display area 102 is a display area where management information is engraved, and display area 103 is a display area where a label (side label) with printed management information is attached.
[0032] Display areas 101 and 102 are located on the top surface of the thick board, and display area 103 is located on the side surface of the thick board. Note that the dashed lines shown in Figures 2(A) to 2(C) are imaginary lines and are not actually displayed on the surface of the thick board.
[0033] FIG. 2(B) shows a second example of the display form for the display area in which management information for the surface of the thick plate, which is the object 1, is displayed.
[0034] 2(B), display area 111 is a display area where management information is printed using a stencil, display area 112 is a display area where management information is engraved, and display area 113 is a display area where a label (side label) with printed management information is attached.
[0035] Display areas 111 and 112 are on the top surface of the plank, and display area 113 is on the side surface of the plank.
[0036] As shown in Figures 2(A) and 2(B), the arrangement of stencil display areas 101, 111, engraved display areas 102, 112, and label display areas 103, 113 can be set appropriately on the surface of the object 1, and may be different for each steel mill that produces thick plates.
[0037] In this embodiment, actual product management (identification and position management of each plank) is performed based on management information of at least one of stencil marking, marking by engraving, and marking by labeling.
[0038] The display area and display form for the thick plate are not limited to those shown in Figures 2(A) and 2(B). For example, the distance between the stencil display area and the engraved display area may be large, or there may be no engraved display area.
[0039] Figure 2(C) shows an example of a stack of multiple thick plates. As shown in Figure 2(C), the stencil markings and engraved markings of the lower thick plates of the stack cannot be seen. In such a case, actual product management can be performed based on the markings by labels 103, 123, and 133 on the sides of the thick plates.
[0040] 2(C) shows a case where the longitudinal direction of one thick plate is parallel to the longitudinal direction of the other thick plate, but the method of stacking the thick plates is not limited to this. For example, the longitudinal direction of one thick plate may be stacked perpendicular or at an angle close to perpendicular to the longitudinal direction of the other thick plate.
[0041] Fig. 3 shows an example of an image captured from a position facing directly at characters displayed on the surface of the object 1. Fig. 3(A) is a diagram showing a first example of a stencil display and an engraved display, and Fig. 3(B) is a diagram showing a second example of a stencil display and an engraved display.
[0042] 3(A) and 3(B), the stencil display includes marks 201, 211, customer names 202, 212, specifications 203, 213, sizes 204, 214, IDs 205, 215, customer codes 206, 216, order numbers 207, 217, and identification information 208, 218. The engraved display includes IDs 209, 219.
[0043] Marks 201 and 211 are marks representing the manufacturer of the thick plate. Customer names 202 and 212 are information indicating the customer (purchaser) of the thick plate. Standards 203 and 213 are information indicating the standard of the thick plate. Sizes 204 and 214 are information indicating the size (thickness x width x length) of the thick plate. IDs 205, 215, 209 and 219 are information for uniquely identifying the thick plate and are plate numbers (identification numbers for the thick plate). Therefore, the same ID is not assigned to different thick plates. Customer codes 206 and 216 are information that customers specify to the thick plate manufacturer to assign to the thick plate. Order numbers 207 and 217 are part of the number for identifying the order of the thick plate from the customer.
[0044] The marks 201, 211, customer names 202, 212, specifications 203, 213, sizes 204, 214, IDs 205, 215, 209, 219, customer codes 206, 216, and order numbers 207, 217 are information that may generally be displayed on a thick plate. Note that information other than the IDs 205, 215, 209, and 219 may not be included in the stencil display and the engraved display. Furthermore, information other than the information described above may be included in the stencil display and the engraved display. For example, the engraved display may display display items equivalent to the stencil display items (e.g., at least one of the mark, customer name, specifications, size, customer code, order information, and identification information).
[0045] In the following description, characters or character strings displayed on the surface of the object 1 to be identified may be referred to as a character group, even if they span multiple lines.
[0046] The imaging unit 3 is a camera that captures an image of the surface of the object 1 while its positional relationship with the object 1 is not constant, and generates a captured image. As shown by the dashed line in FIG. 1, the imaging unit 3 captures an image of a certain range including at least one display area on the surface of the object 1 as an imaging field of view, and generates a captured image consisting of a two-dimensional image. The captured image may be a still image or a video. The imaging unit 3 is not limited to capturing an image of the top surface of the object 1, and may also capture an image of the side of the object 1, or the top and side surfaces simultaneously.
[0047] If the character recognition device 100 is a mobile terminal such as a tablet or smartphone, the imaging unit 3 is a camera attached to the mobile terminal; if the character recognition device 100 has an imaging unit 3 separate from the processing unit 4, the imaging unit 3 is an area camera that can capture images on its own; and if the character recognition device 100 is an unmanned aerial vehicle such as a drone, the imaging unit 3 is a camera installed on board the unmanned aerial vehicle.
[0048] FIG. 4 shows an example of an image of characters displayed on the surface of the object 1 captured from a distorted position.
[0049] As described above, in this embodiment, the object 1 is imaged via a handheld mobile terminal, a crane, or an unmanned aerial vehicle, so the positional relationship between the object 1 and the imaging unit 3 is not constant for each object 1 but varies.
[0050] Therefore, under ideal conditions where the display on the surface of the object 1 and the imaging unit 3 are directly facing each other, an image without distortion as shown in Figure 3(A) or Figure 3(B) should be obtained. However, in general, the display on the surface of the object 1 and the imaging unit 3 are not directly facing each other, so an image is generated in which the characters on the surface of the object 1 are distorted, as shown in Figure 4.
[0051] The arithmetic processing unit 4 is a functional unit that recognizes characters displayed on the surface of the target object 1 based on the captured image generated by the imaging unit 3. The arithmetic processing unit 4 is realized by, for example, a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), a communication device, etc.
[0052] 5 is a diagram illustrating the configuration of the arithmetic processing unit 4 according to the embodiment of the present invention. As shown in FIG. 5, the arithmetic processing unit 4 includes a captured image acquisition unit 41, a first character recognition unit 43, a straight line acquisition unit 45, a facing conversion unit 47, a second character recognition unit 49, and a recognition result output unit 51.
[0053] The captured image acquisition unit 41 is a functional unit that acquires the captured image captured by the imaging unit 3 from the imaging unit 3. The captured image acquisition unit 41 sends the acquired captured image to the first character recognition unit 43, which will be described later.
[0054] The first character recognition unit 43 is a functional unit that performs character recognition based on the captured image. The first character recognition unit 43 performs character recognition using the captured image acquired from the captured image acquisition unit 41, i.e., a distorted image in a state where the positional relationship with the target object 1 is not constant, as input. Then, as a result of the character recognition, the first character recognition unit 43 outputs information on the position (coordinates) in the image, the size, and type of character. In this case, the information indicating the position and size of the character is represented by a rectangle (area for each character) that is placed at the target position in the captured image and is distorted according to the captured image. This information indicating the position and size of the character is referred to as "rectangle information" (bounding box).
[0055] For example, a deep learning model for object detection that can simultaneously output coordinates, size, and character type can be applied to this first character recognition unit 43. For example, the one described in Non-Patent Document 1 can be used as the first character recognition unit 43. Note that the character recognition model may be a known OCR (Optical character recognition) or a technology that combines AI and OCR (so-called AI OCR).
[0056] 4 is input to the first character recognition unit 43, it was found that the reliability of the recognition of character types among the outputs from the first character recognition unit 43 is low. On the other hand, it was found that the rectangular information of each character can often be recognized relatively correctly.
[0057] For this reason, in the subsequent processing, the rectangle information recognized by the first character recognition unit 43 will be utilized regardless of the type of character that is difficult for the first character recognition unit 43 to recognize.
[0058] The first character recognition unit 43 sends the rectangular area for each recognized character to the line acquisition unit 45.
[0059] The straight line acquisition unit 45 is a functional unit that acquires a straight line corresponding to the horizontal direction of the characters displayed on the surface of the object 1 and a straight line corresponding to the vertical direction of the characters displayed on the surface of the object 1 based on the rectangular information recognized by the first character recognition unit 43.
[0060] The straight line acquisition unit 45 acquires straight lines corresponding to the horizontal direction (row direction) and the vertical direction (column direction) of the character group based on the rectangular information (character position and character size) of each character included in the character group recognized by the first character recognition unit 43.
[0061] These lines are later used in calculations by the facing conversion unit 47, and a total of four lines are obtained, for example, two lines (V1, V2) indicating the vertical direction of the distorted character area on the image and two lines (H1, H2) indicating the horizontal direction.
[0062] Fig. 6 is a diagram for explaining acquisition of straight lines by straight line acquisition unit 45 according to an embodiment of the present invention. The many rectangles shown in Fig. 6 represent rectangle information for each character obtained by performing character recognition on the captured image exemplified in Fig. 4 using first character recognition unit 43.
[0063] As a method of obtaining the straight line, as shown in H1, a straight line passing through the center of gravity of each rectangular information relating to the characters arranged in a row may be extracted, or as shown in H1', a straight line circumscribing each rectangular information relating to the characters arranged in a row may be extracted.
[0064] Also, for example, suppose rectangular information for a predetermined number of characters (e.g., five characters) in one line of a group of characters, each consisting of the same font size, and a predetermined number of characters (e.g., five characters) in another line of the same font size can be obtained. This allows a parallelogram consisting of four points: (i) the center position of the rectangular information for the first character in one line, (ii) the center position of the rectangular information for the predetermined number of characters in one line, (iii) the center position of the rectangular information for the first character in another line, and (iv) the center position of the rectangular information for the predetermined number of characters in another line. This parallelogram has two sets of parallel opposing sides, so one of these sets can be considered a straight line corresponding to the horizontal direction of the characters displayed on the surface of the object 1, and the other set can be considered a straight line corresponding to the vertical direction of the characters displayed on the surface of the object 1.
[0065] Also, for example, the horizontal side of the rectangular information for one character can be regarded as the direction of a straight line corresponding to the horizontal direction of the character displayed on the surface of the object 1, and the vertical side of the rectangular information for that one character can be regarded as the direction of a straight line corresponding to the vertical direction of the character displayed on the surface of the object 1.
[0066] In such a case, it is desirable to extract the distance Lv between V1 and V2 and the distance Lh between H1 and H2 so as to be as large as possible, from the viewpoint of improving the conversion accuracy in the subsequent facing conversion unit 47. Strictly speaking, since the rectangular information of the characters is distorted, it is not possible to accurately measure the distances Lv and Lh in real space at this point, but as an index, the difference in the number of lines between V1 and V2 can be used in place of Lv, or the difference in the number of characters from the beginning of each line from the left can be used in place of Lh.
[0067] Each line can be determined by regressing multiple rectangles. In this case, it is desirable to use a large number of characters to determine one line, and it is also desirable to use a row or column with a long distance between the characters at both ends of the multiple characters along the line.
[0068] Another way to improve conversion accuracy is to provide the line acquisition unit 45 with information about the positions and sizes of the characters to be displayed that is known in advance, such as, for example, that M characters are lined up in the Nth row from the top in Figure 6, that the first characters in each column are aligned in a single line, and that each character is the same size.By extracting four lines (H1, H2, V1, V2) based on this information, the accuracy of extracting the four lines can be further improved.
[0069] Information about the four straight lines extracted by the straight line acquisition unit 45 is sent to the facing transformation unit 47.
[0070] The front-facing conversion unit 47 is a functional unit that generates a converted image by converting the captured image so that the characters displayed on the surface of the target object 1 are viewed from the front, based on the straight lines corresponding to the horizontal direction and the vertical direction acquired by the straight line acquisition unit 45.
[0071] That is, the front-facing conversion unit 47 converts the distorted captured image into an undistorted image when the character group is viewed from the front, based on the straight lines (H1, H2) corresponding to the horizontal direction and the straight lines (V1, V2) corresponding to the vertical direction acquired by the straight line acquisition unit 45.
[0072] In the captured image generated by the imaging unit 3, the lines corresponding to the horizontal direction (H1, H2) are not parallel to each other, the lines corresponding to the vertical direction (V1, V2) are not parallel to each other, and the lines corresponding to the horizontal direction and the lines corresponding to the vertical direction are not perpendicular to each other, due to the surface of the object 1 not being directly facing the optical axis of the imaging unit 3. Therefore, if a transformation can be performed so that the lines corresponding to the horizontal direction are parallel to each other, the lines corresponding to the vertical direction are parallel to each other, and the lines corresponding to the horizontal direction and the lines corresponding to the vertical direction are perpendicular to each other, then the transformation can be used to convert the captured image into an image as if the object 1 and the imaging unit 3 were directly facing each other (i.e., an image as if the surface of the object 1 were viewed from the front).
[0073] As such a transformation, various types of transformation are known, and for example, a planar projective transformation can be used. Here, an example will be described in which the planar projective transformation method described in Patent Document 2 is used in the direct facing transformation unit 47.
[0074] In the planar projective transformation, a distorted captured image is used as a transformation target image, and an image taken in a frontal position is used as a transformed image.
[0075] The parameters of the planar projective transformation are expressed as a 3x3 matrix and can be calculated from four or more pairs of corresponding points between the transformed image and the image to be transformed. If the coordinates before transformation are (x, y) and the coordinates after transformation are (x', y'), then the parameters can be expressed by equation (1).
[0076]
number
[0077] Here, for the coordinates (x, y) before conversion, four coordinates are prepared by calculating the vertices of a rectangle formed by a line (H1, H2) corresponding to the horizontal direction and a line (V1, V2) corresponding to the vertical direction, both of which are acquired by the line acquisition unit 45. The coordinates (x', y') after conversion corresponding to the four vertices can be expressed as the difference in the number of characters and the difference in the number of lines, with the number of characters as the unit.
[0078] The difference in the number of characters is a value expressed as the number of characters, including spaces, that Lv and Lh have, with the vertex where lines H1 and V1 intersect as the origin. Methods for calculating the difference in the number of characters include counting the number of characters recognized in the section from line V1 to line V2 on line H1, or dividing the number of pixels in the section from line V1 to line V2 on line H1 by the number of pixels per character calculated using the median value of the spacing between recognized adjacent characters, and rounding the result to an integer.
[0079] Similarly, the line number difference is calculated by dividing the number of pixels in the section from line H1 to line H2 on line V1 by the number of pixels per character, and rounding the result to an integer.
[0080] In the example shown in FIG. 6, Lv is 7 characters, Lh is 5 characters, and the converted coordinates (x', y') are (0, 0), (0, 5), (7, 0), and (7, 5).
[0081] Furthermore, by introducing a scaling (magnification) parameter into the planar projection transformation, the number of pixels can be made suitable for character recognition.
[0082] The geometric distortion correction (planar projective transformation) that introduces the scaling parameters can be expressed by equation (2) when the scaling parameter in the width direction is Sx and the scaling parameter in the height direction is Sy.
[0083]
number
[0084] Here, when the size of the captured image is 60 pixels wide and 80 pixels high, the width scaling parameter Sx can be set to 60 pixels and the height scaling parameter Sy can be set to 80 pixels, which are the number of pixels per character suitable for character recognition.
[0085] When a captured image is subjected to planar projection transformation based on four straight lines to generate a new image that resembles characters displayed on the surface of an object viewed from the front, this image is referred to as a transformed image.
[0086] The converted image generated by the opposite conversion unit 47 is sent to the second character recognition unit.
[0087] The second character recognition unit 49 is a functional unit that performs character recognition based on the converted image.
[0088] That is, the second character recognition unit 49 receives the converted image obtained by the opposite conversion unit 47 as input, and outputs, as the character recognition result, information on the size and type of character present at which position (coordinate) in the converted image.
[0089] Second character recognition unit 49 may perform character recognition using the same character recognition model as that used in first character recognition unit 43, or may perform character recognition using a different character recognition model. An important difference from first character recognition unit 43 is that the input image is converted by frontal facing conversion unit 47 so that the character group is viewed from the front, eliminating image distortion, and therefore, highly reliable results are output not only for easily recognizable rectangular information (coordinates, size) but also for the type of character.
[0090] In this way, the character recognition device 100 can recognize characters displayed on the surface of the object 1 with sufficient recognition accuracy by using the character type obtained by the second character recognition unit 49 as the final result of the character recognition (regarding the character type).
[0091] Regarding the coordinates and size of the characters, the results of the first character recognition unit 43 or the results of the second character recognition unit 49 may be the final results.
[0092] The second character recognition unit 49 sends the result of the character recognition to the recognition result output unit 51 .
[0093] The recognition result output unit 51 is a functional unit that outputs the results of character recognition to a display, printer, etc. (not shown) that is included in the calculation processing unit 4 or that is connected to the calculation processing unit 4, and outputs the results of character recognition to an operator of the character recognition device 100, etc.
[0094] As an output method in the recognition result output unit 51, any known method can be used as appropriate, for example, the recognized characters for each line can be output in text format, or the recognized character types can be overlaid on the converted image at each position and displayed. The output results can also be visually confirmed by an operator, or can be sent to another computer and then collated with a database of steel products.
[0095] Next, the processing in the character recognition device 100 according to the embodiment of the present invention will be described with reference to Fig. 7. Fig. 7 shows a flowchart for explaining the processing in the character recognition device 100 according to the embodiment of the present invention.
[0096] (Step S001)
[0097] When the character recognition device 100 according to this embodiment starts processing related to the character recognition method, processing in step S001 is performed. In step S001, the imaging unit 3 captures an image of the surface of the object 1 to generate a captured image (imaging step).
[0098] At this time, the imaging unit 3 is used to capture an image of an area on the surface of the object 1 that includes a group of characters to be recognized, thereby generating a captured image.
[0099] As described above, in this embodiment, the object 1 is imaged using a handheld mobile terminal, a large-scale crane, or an unmanned aerial vehicle that can move in a three-dimensional manner, so the positional relationship between the object 1 and the imaging unit 3 is not constant for each object 1 but varies.
[0100] Therefore, under ideal conditions where the display on the surface of the object 1 and the imaging unit 3 are directly facing each other, an undistorted image such as that shown in Figure 3(A) or Figure 3(B) should be obtained. However, in general, the characters on the surface of the object 1 and the imaging unit 3 are not directly facing each other, and an image in which the characters on the surface of the object 1 are distorted will be generated, as shown in Figure 4.
[0101] When the generation of the captured image is completed, the captured image is sent to the captured image acquisition unit 41, and the process proceeds to step S003.
[0102] (Step S003)
[0103] In step S003, the captured image generated in step S001 is acquired wirelessly or via a wire from the imaging unit 3 using the captured image acquisition unit 41. The acquired captured image is stored in the arithmetic processing unit 4, and can be used in various processes of the arithmetic processing unit 4.
[0104] Once the captured image is acquired, the process proceeds to step S005.
[0105] (Step S005)
[0106] In step S005, the first character recognition unit 43 performs character recognition based on the captured image (first character recognition step).
[0107] In step S005, for example, a deep learning model for object detection that can simultaneously output the coordinates, size, and type of characters can be used.
[0108] As described above, the first character recognition unit 43 generally receives a distorted captured image such as that shown in Figure 4, and therefore the type of character output from the first character recognition unit 43 is not very reliable even if it can be recognized, but the rectangular information (position, size) for each character can be recognized relatively accurately.
[0109] Therefore, in step S005, if character rectangle information can be recognized for a certain number of characters (this can be set appropriately within that range if the processing in the straight line acquisition unit 45 and the direct opposite conversion unit 47 can be executed), the character rectangle information is sent to the straight line acquisition unit 45 regardless of whether the type of character can be recognized or not, and the process proceeds to step S007.
[0110] (Step S007)
[0111] In step S007, based on the rectangular information for each character recognized in step S005, a straight line corresponding to the horizontal direction of the character printed on the surface of the object 1 and a straight line corresponding to the vertical direction of the character printed on the surface of the object 1 are obtained (straight line acquisition step).
[0112] In step S007, the straight line acquisition unit 45 acquires two straight lines corresponding to the horizontal direction and two straight lines corresponding to the vertical direction of the group of characters displayed on the surface of the target object 1 based on the shape and arrangement of the rectangular information (position, size) for each character recognized in step S005.
[0113] If the straight line acquisition unit 45 can acquire two straight lines corresponding to the horizontal direction and two straight lines corresponding to the vertical direction, the information is sent to the facing conversion unit 47, and the process proceeds to step S009.
[0114] (Step S009)
[0115] In step S009, based on the acquired straight lines corresponding to the horizontal direction and the vertical direction, a converted image is generated by converting the captured image so that the characters printed on the surface of the object 1 are viewed from the front (frontal conversion step).
[0116] That is, in step S009, the front facing transformation unit 47 uses a known planar projection transformation based on the total of four straight lines acquired in step S007 to transform the four straight lines so that they form a rectangle, thereby generating a transformed image from the captured image that is an image in which the group of characters is viewed from the front.
[0117] If the converted image can be generated, the converted image is sent to the second character recognition unit 49, and the process proceeds to step S011.
[0118] (Step S011)
[0119] In step S011, character recognition is performed based on the converted image (second character recognition step).
[0120] That is, in step S011, second character recognition unit 49 performs character recognition again based on the converted image, generated in step S009, in which the characters displayed on the surface of object 1 are converted so as to be viewed from the front. In this case, the same character recognition model as that used in the first character recognition unit may be used, or a different character recognition model may be used.
[0121] In the character recognition in step S011, the input image is a converted image that has been converted in step S009 so that the group of characters is viewed from the front, and therefore image distortion has been removed. As a result, highly reliable recognition results are output not only for the rectangular information (position, size) of characters that are easy to recognize, but also for the type of character.
[0122] In this way, the character recognition device 100 can recognize characters displayed on the surface of the object 1 with high accuracy by using the character type obtained by the second character recognition unit 49 as the final result of the character recognition (regarding the character type).
[0123] When the character recognition is completed in step S011, the character recognition results are stored in the arithmetic processing unit 4 so that they can be handled appropriately, and then the process proceeds to step S013.
[0124] (Step S013)
[0125] In step S013, the recognition result output unit 51 outputs the character recognition result obtained in step S011 to the operator of the character recognition device 100, etc., so that the operator, etc. can utilize the output character recognition result with sufficient accuracy.
[0126] By performing the above processing, the processing in the character recognition device 100 according to the embodiment of the present invention is completed.
[0127] As explained above, according to the present invention, when attempting to recognize characters displayed on the surface of an object 1, even if the positional relationship between the object 1 and the imaging unit 3 is not constant and varies for each object 1, the captured image in which the characters to be recognized are captured is first subjected to character recognition, and based on the resulting positional area of the characters that are easy to recognize, two lines corresponding to the horizontal direction and two lines corresponding to the vertical direction are obtained, and based on these four lines, the captured image is converted so that it is viewed from the front, and then character recognition is performed again, making it possible to recognize the type of character with sufficient recognition accuracy.
[0128] It should be noted that the above-described embodiments of the present invention are merely examples of specific embodiments for carrying out the present invention, and the technical scope of the present invention should not be construed as being limited by these. In other words, the present invention can be embodied in various forms without departing from its technical concept or main features.
[0129] <Hardware>
[0130] An example of the hardware configuration of the arithmetic processing unit according to an embodiment of the present invention is shown in Fig. 8. An example of hardware for realizing the arithmetic processing unit 4 will be described with reference to Fig. 8.
[0131] 8, the arithmetic processing unit 4 has a CPU 1201, a main memory device 1202, an auxiliary memory device 1203, a communication circuit 1204, a signal processing circuit 1205, an image processing circuit 1206, an I / F circuit 1207, a user interface 1208, a display 1209, and a bus 1210.
[0132] The CPU 1201 controls the entire arithmetic processing unit 4. The CPU 1201 uses the main memory device 1202 as a work area to execute programs stored in the auxiliary memory device 1203. The main memory device 1202 temporarily stores data. The auxiliary memory device 1203 stores various types of data in addition to the programs executed by the CPU 1201.
[0133] The communication circuit 1204 is a circuit for communicating with the outside of the arithmetic processing unit 4. The communication circuit 1204 may perform wireless communication or wired communication with the outside of the arithmetic processing unit 4.
[0134] The signal processing circuit 1205 performs various signal processing on signals received by the communication circuit 1204 and signals input under the control of the CPU 1201 .
[0135] The image processing circuit 1206 performs various types of image processing on the input signal under the control of the CPU 1201. The signal that has undergone this image processing is output to a display 1209, for example.
[0136] The user interface 1208 is a part through which an operator gives instructions to the arithmetic processing unit 4. The user interface 1208 has, for example, buttons, switches, dials, etc. The user interface 1208 may also have a graphical user interface using a display 1209.
[0137] The display 1209 displays an image based on a signal output from the image processing circuit 1206. The I / F circuit 1207 exchanges data with devices connected to the I / F circuit 1207. In FIG. 12, a user interface 1208 and a display 1209 are shown as devices connected to the I / F circuit 1207. However, the devices connected to the I / F circuit 1207 are not limited to these. For example, a portable storage medium may be connected to the I / F circuit 1207. Furthermore, at least a part of the user interface 1208 and the display 1209 may be located outside the arithmetic processing unit 4.
[0138] The output unit 415 is realized by using, for example, at least one of the communication circuit 1204 and signal processing circuit 1205, and the image processing circuit 1206, I / F circuit 1207, and display 1209.
[0139] The CPU 1201, main memory device 1202, auxiliary memory device 1203, signal processing circuit 1205, image processing circuit 1206, and I / F circuit 1207 are connected to a bus 1210. Communication between these components is performed via the bus 1210. The hardware of the arithmetic processing unit 4 is not limited to that shown in FIG. 8 as long as it can realize the functions of the arithmetic processing unit 4 described above. [Industrial Applicability]
[0140] The present invention can be used, for example, to identify products when managing steel products in-situ. [Explanation of symbols]
[0141] 100 character recognition device 1. Object 2 Holding part 3. Imaging unit 4. Processing unit 41 Image acquisition unit 43 1st character recognition section 45 Straight line acquisition part 47 Opposite conversion section 49 Second character recognition section
Claims
1. A character recognition device that recognizes characters displayed on the surface of a steel material, an imaging unit that captures an image of the surface of the ferrous material and generates a captured image while the positional relationship with the ferrous material is not constant; a processing unit that recognizes characters displayed on the surface of the steel material based on the captured image; and The arithmetic processing unit a first character recognition unit that performs character recognition based on the captured image; a straight line acquisition unit which, when the first character recognition unit acquires rectangular information of a predetermined number of characters in a certain line and a predetermined number of characters in another line, each made of the same font size, grasps a parallelogram consisting of four points: the position of the rectangular information of the first character in the certain line, the position of the rectangular information of the predetermined number of characters in the certain line, the position of the rectangular information of the first character in the other line, and the position of the rectangular information of the predetermined number of characters in the other line; and of two sets of parallel opposing sides included in the parallelogram, regards one set of straight lines as a straight line corresponding to the horizontal direction of the characters displayed on the surface of the steel material, and regards the other set of straight lines as a straight line corresponding to the vertical direction of the characters displayed on the surface of the steel material, thereby acquiring the straight line corresponding to the horizontal direction of the characters displayed on the surface of the steel material and the straight line corresponding to the vertical direction of the characters displayed on the surface of the steel material; a front-facing conversion unit that generates a converted image by converting the captured image so that the characters displayed on the surface of the steel material are viewed from the front, based on the straight line corresponding to the horizontal direction and the straight line corresponding to the vertical direction acquired by the straight line acquisition unit; a second character recognition unit that performs character recognition based on the converted image; and a character recognition device that outputs the characters recognized by the second character recognition unit;
2. A character recognition method for recognizing characters displayed on the surface of the steel material using the character recognition device according to claim 1, comprising: an imaging step in which the imaging unit images a surface of the ferrous material and generates an image in a state where the positional relationship with the ferrous material is not constant; a first character recognition step in which the first character recognition unit performs character recognition based on the captured image; a straight line acquisition step in which, when the first character recognition unit acquires rectangular information of a predetermined number of characters in a certain line and a predetermined number of characters in another line, each consisting of a font of the same size, the straight line acquisition unit grasps a parallelogram consisting of four points: the position of the rectangular information of the first character in the certain line, the position of the rectangular information of the predetermined number of characters in the certain line, the position of the rectangular information of the first character in the other line, and the position of the rectangular information of the predetermined number of characters in the other line, and of two sets of parallel opposing sides included in the parallelogram, the straight lines of one set are regarded as straight lines corresponding to the horizontal direction of the characters displayed on the surface of the steel material, and the straight lines of the other set are regarded as straight lines corresponding to the vertical direction of the characters displayed on the surface of the steel material, thereby acquiring the straight lines corresponding to the horizontal direction of the characters displayed on the surface of the steel material and the straight lines corresponding to the vertical direction of the characters displayed on the surface of the steel material; a facing transformation step in which the facing transformation unit generates a transformed image by transforming the captured image so that the characters displayed on the surface of the steel material are viewed from the front, based on the straight line corresponding to the horizontal direction and the straight line corresponding to the vertical direction acquired in the straight line acquisition step; a second character recognition step in which the second character recognition unit performs character recognition based on the converted image; A character recognition method comprising:
Citation Information
Patent Citations
Method for centering tail part of strip
JP1985003918A
Apparatus and method for image processing
JP2003288588A
Image processing apparatus, image processing method and program
JP2012022413A
License plate reader
JP2012063869A
Information processor and control method thereof and program
JP2020149184A