Image processing apparatus, image processing method, and image processing program

The image processing apparatus enhances recognition accuracy by aligning multiple objects in an image to face the front direction, addressing the challenge of varying orientations and improving character recognition efficiency.

JP7701569B2Active Publication Date: 2025-07-01RAKUTEN GROUP INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024534761
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-07-27
Publication Date
2025-07-01
Estimated Expiration
2043-07-27

AI Technical Summary

Technical Problem

Conventional image recognition techniques face challenges in improving the accuracy of recognizing multiple objects in an image when their writing surfaces face different directions or projections relative to the camera, leading to suboptimal recognition results.

Method used

An image processing apparatus that estimates normal vectors for each object, determines a representative normal vector direction, and geometrically deforms the image to align all objects facing the front direction, enhancing the suitability for character recognition.

Benefits of technology

This approach improves the recognition accuracy of characters in multiple objects within an image by aligning them optimally for recognition in a single transformation and recognition process, maintaining positional relationships and reducing the number of required processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701569000001
    Figure 0007701569000001
  • Figure 0007701569000002
    Figure 0007701569000002
  • Figure 0007701569000003
    Figure 0007701569000003
Patent Text Reader

Abstract

The present invention addresses the problem of performing image conversion suitable for character recognition for a query image in which a plurality of regions are captured. An image processing device 1 comprises: a query image acquisition unit 21 that acquires a query image in which a plurality of objects are captured; a normal line inference unit 25 that infers, for each of the plurality of objects captured in the query image, a normal direction indicating the front direction of the object; a representative normal line determination unit 26 that determines a representative normal direction representing the entire query image, on the basis of the plurality of normal directions inferred for the plurality of objects; and a conversion unit 27 that generates a corrected query image by geometrically deforming the query image on the basis of the representative normal direction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to image processing technology.

Background Art

[0002] Conventionally, there has been proposed a recognition device including a vehicle detection unit that cuts out a peripheral image of a license plate from an image acquired by an imaging device, a license plate detection unit that extracts a front image of the license plate from the peripheral image of the license plate, a character recognition unit that recognizes a vehicle registration number from the front image of the license plate, a reliability determination unit that determines the ability to extract the front image of the license plate detection unit based on the recognition difficulty of the character recognition unit, and a model change unit that changes the ability to extract the front image of the license plate detection unit according to the determination result of the reliability determination unit (see Patent Document 1).

[0003] Also conventionally, a technique for estimating the displacement of a camera between two images of a planar object by decomposing a homography matrix is known (see Non-Patent Document 1).

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Non-Patent Documents

[0005]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] Conventionally, various techniques for recognizing a predetermined object (such as an imaged object or character) included in an image have been proposed, but there is room for improvement in the recognition accuracy of the object in the image. In view of the above problems, an object of the present disclosure is to improve the recognition accuracy of a predetermined object included in an image.

Means for Solving the Problems

[0007] An example of the present disclosure includes query image acquisition means for acquiring a query image in which a plurality of objects are imaged, normal vector estimation means for estimating, for each of the plurality of objects imaged in the query image, a normal vector direction indicating the front direction of the object, representative normal vector determination means for determining a representative normal vector direction representing the entire query image based on the plurality of normal vector directions estimated for the plurality of objects, and conversion means for generating a corrected query image by geometrically deforming the query image based on the representative normal vector direction, and is an image processing apparatus.

[0008] The present disclosure can be understood as an image processing apparatus, a system, a method executed by a computer, or a program to be executed by a computer. Further, the present disclosure can also be understood as a recording medium in which such a program is recorded and can be read by a computer or other devices, machines, etc. Here, the recording medium readable by a computer or the like refers to a recording medium that accumulates information such as data and programs by an electrical, magnetic, optical, mechanical, or chemical action and can be read by a computer or the like.

Advantages of the Invention

[0009] According to the present disclosure, it is possible to improve the recognition accuracy of a predetermined object included in an image.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0011] Hereinafter, embodiments of an image processing apparatus, method, and program according to the present disclosure will be described with reference to the drawings. However, the embodiments described below are illustrative of the embodiments, and do not limit the image processing apparatus, method, and program according to the present disclosure to the specific configurations described below. In practice, specific configurations according to the implementation mode may be appropriately adopted, and various improvements and modifications may be made.

[0012] Conventionally, after performing a geometric transformation that faces the region to be subjected to character recognition, which is included in an image acquired by imaging or the like, toward the front, character recognition has been performed. However, when a plurality of regions to be subjected to character recognition are included in an image, and the writing surfaces (fronts) of the characters to be recognized in these plurality of regions face different directions from each other, or the projections are different from each other depending on the positional relationship with the camera used for imaging, even if a geometric transformation that faces any one of the regions toward the front is performed, the region that is oriented toward the front will be in a state suitable for character recognition, but the other regions will not be in a state suitable for character recognition. That is, there is room for improvement in the conventional technology in terms of improving the recognition accuracy in a case where a plurality of regions to be subjected to character recognition are located separately within the same image. In view of the above-described problems, the present disclosure aims to perform an image transformation suitable for character recognition on a query image in which a plurality of regions have been imaged.

[0013] In the present embodiment, an embodiment will be described in which the technique according to the present disclosure is applied to a query image in which a plurality of objects have been imaged, an image transformation suitable for character recognition is performed, and then a character recognition process is performed on the transformed image. However, the technique according to the present disclosure can be widely used for performing an image transformation suitable for character recognition on a query image in which a plurality of regions have been imaged, and the application target of the present disclosure is not limited to the examples shown in the embodiment. For example, the object may not be a physical object, but an object drawn in a virtual space or an object drawn using drawing software. Further, for example, the image may not be an image of the real world that has been imaged, but an image of a virtual world that has been drawn or an image drawn using drawing software.

[0014] <Configuration of the System> FIG. 1 is a schematic diagram showing the configuration of the system according to the present embodiment. The system according to the present embodiment includes an image processing apparatus 1, an imaging apparatus 81, and a user terminal 9 that can communicate with each other by being connected to a network.

[0015] The image processing apparatus 1 is a computer including a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a storage device 14 such as an EEPROM (Electrically Erasable and Programmable Read Only Memory) or an HDD (Hard Disk Drive), a communication unit 15 such as a NIC (Network Interface Card), and the like. However, regarding the specific hardware configuration of the image processing apparatus 1, appropriate omissions, replacements, and additions can be made according to the embodiment. Also, the image processing apparatus 1 is not limited to a device consisting of a single housing. The image processing apparatus 1 may be realized by a plurality of devices using so-called cloud or distributed computing technologies and the like.

[0016] The imaging device 81 obtains a query image described later by imaging an object. As the imaging device 81, a general digital camera or other device capable of recording light incident from the object may be used, and its specific configuration is not limited.

[0017] The user terminal 9 is a terminal device used by a user. The user terminal 9 is a computer including a CPU, a ROM, a RAM, a storage device, a communication unit, an input device, an output device, etc. (illustrations are omitted). However, regarding the specific hardware configuration of the user terminal 9, appropriate omissions, replacements, and additions can be made according to the embodiment. Also, the user terminal 9 is not limited to a device consisting of a single housing. The user terminal 9 may be realized by a plurality of devices using so-called cloud or distributed computing technologies and the like. The user uses various services provided by the system according to this embodiment via these user terminals 9.

[0018] FIG. 2 is a diagram showing an outline of the functional configuration of the image processing apparatus 1 according to the present embodiment. In the image processing apparatus 1, the program recorded in the storage device 14 is read into the RAM 13 and executed by the CPU 11, and each hardware provided in the image processing apparatus 1 is controlled, so that the query image acquisition unit 21, the object detection unit 22, the shape data acquisition unit 23, the homography matrix calculation unit 24, the normal vector estimation unit 25, the representative normal vector determination unit 26, the conversion unit 27, and the character recognition unit 28 are provided. In the present embodiment and other embodiments described later, each function provided in the image processing apparatus 1 is executed by the CPU 11 which is a general-purpose processor, but a part or all of these functions may be executed by one or a plurality of dedicated processors.

[0019] The query image acquisition unit 21 acquires a query image in which a plurality of objects to be subjected to image processing and / or character recognition are simultaneously imaged. The method of acquiring the query image is not limited, but in the present embodiment, an example of acquiring a query image imaged using the imaging device 81 via the user terminal 9 will be described.

[0020] FIG. 3 is a diagram showing an example of a query image in the present embodiment. In this figure, a plurality of smartphones in a state where product guides are displayed on the touch panel display of the smartphone installed on a table in the store, and an explanatory medium (hereinafter simply referred to as "product information card") on which information such as explanations and prices for each of the plurality of smartphones is described and which is installed near the smartphone, are imaged in the query image. The person in charge takes a photo so that these objects, that is, the smartphones and the product information cards, which are the objects of image processing and character recognition in the present embodiment, are included in the same image, and uses this as the query image.

[0021] FIG. 4 is a diagram showing an example in which a geometric transformation according to a conventional method is applied to a query image in the present embodiment. According to this figure, for example, when image conversion is performed so that the product information card in the center of the query image faces the camera, the central product information card facing the camera after the conversion is in a state suitable for character recognition. However, for other objects, especially the smartphone at the screen edge where the front direction (in this embodiment, since the object is the target of character recognition, hereinafter, the normal direction of the plane on which the recognition target characters are displayed on the object is defined as the front direction) is significantly different from that of the central product information card, it can be seen that the distortion increases and the accuracy of character recognition will decrease. Also, before the image conversion, the positions of the smartphone and the product information card corresponding to the smartphone have a vertical alignment relationship (see FIG. 3). However, after the image conversion, the positions of the smartphone and the product information card corresponding to the smartphone are shifted in the vertical alignment, and it can be seen that it is also difficult to associate the smartphone with the product information card.

[0022] When trying to solve the above problems, it is conceivable to perform image conversion that turns each of the multiple objects included in the query image to face the front and perform character recognition on the object. However, when such a method is adopted, the same number of image conversion and character recognition processes as the number of objects in the query image are required. Therefore, in the system according to the present embodiment, while suppressing the number of image conversion and character recognition processes for one query image, the accuracy of character recognition is increased.

[0023] The object detection unit 22 detects a plurality of object images from the query image. In the present embodiment, the object detection unit 22 uses a machine learning model that outputs the position or range in the query image in which an object related to the type or attribute of the object to be the target of character recognition is imaged, in response to an input of the type or attribute of the object. The plurality of objects imaged in the query image are detected. However, the type of object detection technology used for object detection from an image is not limited, and any object detection technology known currently or developed in the future may be used.

[0024] The shape data acquisition unit 23 acquires shape data related to the front shape of the object. For example, the shape data may be the aspect ratio, size, design data, front image, etc. of the object that are previously held in the RAM 13 or the storage device 14, etc. Further, the shape data may be data that uniquely determines at least the front shape of an object related to the same type or attribute.

[0025] However, even for objects related to the same type or attribute, depending on the specified type or attribute, the front shape of the object related to the type or attribute may not be uniquely determined. For example, when the object is specified as "smartphone manufactured by XX company, model number XX-XX" or "B8-size card", its shape can be uniquely determined, but when the object is specified only as "smartphone", "smartphone manufactured by XX company", or "product information card", there is a range in its shape. For this reason, in the present embodiment, the shape data may be shape data having a range for an object related to the same type or attribute. The shape data having a range may be, for example, the range of the aspect ratio of the object, the range of the size, a plurality of front images, etc. In the present embodiment, the case where the range of the aspect ratio of the object is used as the shape data will be described as an example.

[0026] The homography matrix calculation unit 24 calculates a homography matrix between the front shape of the object specified according to the shape data and the object image related to the object in the query image for each of the plurality of objects. Here, when the shape data has a range as described above, the homography matrix calculation unit 24 calculates a homography matrix related to the minimum value and a homography matrix related to the maximum value in the range of the shape data for each of the plurality of objects.

[0027] The normal vector estimation unit 25 estimates the normal vector direction indicating the front direction of each of the plurality of objects imaged in the query image. In the present embodiment, the normal vector estimation unit 25 estimates the normal vector direction (normal vector) of the object based on the homography matrix calculated by the homography matrix calculation unit 24 for each of the plurality of objects, thereby estimating the normal vector direction of the object according to the comparison result between the front shape of the object specified according to the shape data and the object image in the query image.

[0028] However, the specific method for estimating the normal vector direction of the object is not limited to calculating the normal vector from the homography matrix. Also, when the shape data has a range as described above, the normal vector estimation unit 25 calculates the normal vector direction related to the minimum value and the normal vector direction related to the maximum value in the range of the shape data for each of the plurality of objects.

[0029] The representative normal vector determination unit 26 determines a representative normal direction that represents the entire query image based on a plurality of estimated normal directions for a plurality of objects. In the present embodiment, the representative normal vector determination unit 26 sets the average value of a plurality of estimated normal vectors for a plurality of objects as the representative normal vector, and determines the direction of the representative normal vector as the representative normal direction. However, the representative normal direction may be obtained by a method of obtaining a representative direction using a statistical method, and may be calculated by a method other than calculating the average value of the normal vectors. Further, when the shape data has a range as described above, the representative normal vector determination unit 26 determines, as the representative normal direction, a normal direction that represents a plurality of normal directions corresponding to the minimum value in the range of the shape data estimated for a plurality of objects and a plurality of normal directions corresponding to the maximum value in the range of the shape data estimated for each of the plurality of objects.

[0030] The conversion unit 27 generates a corrected query image by geometrically deforming the query image based on the representative normal direction. In the present embodiment, the conversion unit 27 generates a corrected query image by performing a projective transformation on the query image so that the representative normal direction becomes the front direction of the image after conversion.

[0031] The character recognition unit 28 recognizes characters described in an object to be the target of character recognition by performing character recognition processing on the corrected query image. The technology according to the present disclosure aims to provide image conversion suitable for character recognition, and the specific algorithm adopted for character recognition is not limited. For character recognition, conventionally used or future-developed character recognition technologies may be adopted.

[0032] <Flow of processing> Next, the flow of processing executed by the image processing apparatus according to the present embodiment will be described. Note that the specific content and processing order of the processing described below are an example for implementing the present disclosure. The specific processing content and processing order may be appropriately selected according to the embodiment of the present disclosure.

[0033] 5 is a flowchart showing the flow of image conversion and character recognition processing according to this embodiment. The processing shown in this flowchart is executed when an instruction to start processing by a user is accepted.

[0034] In step S101, a query image is acquired. The worker captures an image of an object using the imaging device 81, and inputs image data of the obtained query image to the image processing device 1. In this embodiment, the image capture object is a table in a store on which a smartphone and a product information card are installed as predetermined objects. Although the imaging method and the method of inputting image data to the image processing device 1 are not limited, in this embodiment, the object is captured using the imaging device 81, and the image data transferred from the imaging device 81 to the user terminal 9 via communication or a recording medium is further transferred to the image processing device 1 via a network, whereby the image data of the query image is input to the image processing device 1. When the query image is acquired by the query image acquisition unit 21, the process proceeds to step S102.

[0035] In step S102, an object to be the target of character recognition is specified. The object detection unit 22 acquires object specification information indicating the type or attribute of the object to be the target of character recognition among the objects imaged in the query image, which is input by the operator. In the present embodiment, an example will be described in which, among the objects imaged in a query image in which a table in a store is imaged, a smartphone displayed as a product and a product information card of the smartphone are the targets of character recognition. For this reason, in the present embodiment, for example, "smartphone" and "tag" are input as object identification information. Note that, in the present embodiment, an example will be described in which object specification information is acquired as text data input by the user from a prompt of the UI, but the method of acquiring object identification information is not limited to the example in the present embodiment. For example, the object identification information may be acquired in a format other than text data (for example, a code indicating the type or attribute of the object), or the object identification information may be acquired by being set in the system in advance before the acquisition of the query image. Thereafter, the process proceeds to step S103.

[0036] In steps S103 and S104, an object is detected from the query image, and the range in which the object in the query image is imaged is cut out as an object image. The object detection unit 22 detects a plurality of objects of the type or attribute specified by the object specification information acquired in step S102 from the query image acquired in step S101 (step S103). Here, it is preferable that all of the objects of the type or attribute specified by the object specification information among the objects imaged in the query image are detected. In the present embodiment, for example, a zero-shot object detector using a convolutional neural network (CNN) may be used for object detection from the query image. However, the type of algorithm used for object detection from an image is not limited, and any algorithm currently known or developed in the future may be used.

[0037] When a plurality of objects are detected from the query image, the object detection unit 22 cuts out, from the query image, the range in which the object is imaged as an object image for the plurality of detected objects (step S104). However, the image data of the object image does not necessarily have to be newly generated, and it is sufficient that the range in the query image is specified for the object image. Thereafter, the process proceeds to step S105.

[0038] In step S105, the aspect ratio (width-to-height ratio) of the object is set. The shape data acquisition unit 23 sets, for each of the plurality of objects detected in step S103, an aspect ratio corresponding to the type or attribute of the object. Specifically, the shape data acquisition unit 23 selects, based on the type or attribute of the detected object, an aspect ratio related to a part or the whole of the configuration of the object when viewed from the front (in plan view), which is held in advance for each type or attribute of the object, and sets it as the aspect ratio of the object. Here, the type or attribute of the object can be identified based on the object designation information input in step S102 or the output obtained from the object detector in step S103. In the present embodiment, for the object "smartphone" detected from the query image, the aspect ratio held in advance as the aspect ratio of the smartphone is set, and for the object "product information card", the aspect ratio held in advance as the aspect ratio of the product information card is set.

[0039] Note that the aspect ratio set here may have a range. For example, the aspect ratio of a smartphone varies depending on the model. For this reason, in the present embodiment, the aspect ratio of the smartphone is set in a range from a minimum value to a maximum value, such as "between 0.479 and 0.486". When the aspect ratio (width-to-height ratio) of the object is set, the process proceeds to step S106.

[0040] In steps S106 and S107, the homography matrix and the normal vector of the object are calculated. The homography matrix calculation unit 24 calculates the homography matrix for each of the plurality of objects detected in step S103 using the aspect ratio range set in step S105. First, the homography matrix calculation unit 24 detects a predetermined component of the object from the object image for each object and matches it with the predetermined component related to the set aspect ratio. In the present embodiment, the objects are a smartphone and a product information card, both of which generally have a substantially rectangular shape. Therefore, the homography matrix calculation unit 24 detects the vertices (corners) of the rectangle related to the object and matches the detected vertices with the vertices of the rectangle related to the set aspect ratio.

[0041] Here, as a means for detecting a predetermined component (in this embodiment, the vertices of the rectangle) of the object from the object image, for example, an image analysis technique for detecting feature points in the image may be used. However, the type of the image analysis technique used for detecting the predetermined component from the object image is not limited, and any image analysis technique known currently or developed in the future may be used.

[0042] When the matching between the predetermined component detected from the object and the predetermined component related to the set aspect ratio is completed, the homography matrix calculation unit 24 calculates the homography matrix of the object by comparing the rectangle of the object in the object image that has been matched with the set aspect ratio (step S106).

[0043] When the homography matrix is calculated, the normal vector estimation unit 25 obtains the normal vector of the object by decomposing the homography matrix obtained in step S106 for each of the plurality of object images detected in step S103 (step S107. For details of the method of estimating the normal vectors of two planar objects by decomposing the homography matrix, refer to Non-Patent Document 1.).

[0044] Figure 6 shows the normal vectors N A and N B estimated for two objects A and B in different orientations in this embodiment. Since the aspect ratio used when calculating the homography matrix is the aspect ratio of the object when viewed from the front, the normal vector indicates the front direction of the object. Also, since the normal vector calculated in this embodiment is calculated based on an aspect ratio with a range, it is a normal vector having a range from the normal vector calculated based on the minimum aspect ratio (hereinafter referred to as the "minimum ratio normal vector") to the normal vector calculated based on the maximum aspect ratio (hereinafter referred to as the "maximum ratio normal vector").

[0045] Figure 7 is a diagram showing examples of the minimum ratio normal vectors (indicated by broken lines) and the maximum ratio normal vectors (indicated by solid lines) estimated for each of a plurality of objects in the query image in this embodiment. When normal vectors are calculated for the detected plurality of objects, the process proceeds to step S108.

[0046] In step S108, a representative normal vector is determined. The representative normal vector determination unit 26 determines a representative normal vector representing the query image including these objects based on the plurality of normal vectors calculated in step S107 for the plurality of objects. In this embodiment, the representative normal vector determination unit 26 determines the average of the normal vectors calculated for each of the plurality of objects as the representative normal vector. However, the method of determining the representative normal vector is not limited to calculating the average value, and a method of obtaining a representative vector using a statistical method may be adopted.

[0047] Figure 8 shows a plurality of minimum ratio normal vectors N A1 and N B1and a plurality of maximum ratio normal vectors N A2 and N B2 FIG. is a diagram showing a concept in which a representative normal vector N is calculated based on and. As described above, the plurality of normal vectors calculated in step S107 are each in the range from the minimum ratio normal vector to the maximum ratio normal vector (for object A in the figure, the minimum ratio normal vector N A1 to the maximum ratio normal vector N A2 range. For object B in the figure, the minimum ratio normal vector N B1 to the maximum ratio normal vector N B2 range.). Therefore, the representative normal vector determination unit 26 calculates a representative value of the plurality of minimum ratio normal vectors (N1 in the figure. Hereinafter, referred to as the "minimum ratio representative normal vector") and a representative value of the plurality of maximum ratio normal vectors (N2 in the figure. Hereinafter, referred to as the "maximum ratio representative normal vector"), and calculates a representative value (N in the figure. For example, the average value) of the minimum ratio representative normal vector and the maximum ratio representative normal vector to determine the representative normal vector. When the representative normal vector is determined, the process proceeds to step S109.

[0048] In step S109, the query image is projective-transformed based on the representative normal vector. The conversion unit 27 performs projective transformation (perspective correction by homography transformation in this embodiment) on the query image so that the representative normal vector determined in step S108 faces the front, thereby obtaining a corrected query image. In other words, the conversion unit 27 calculates a homography matrix such that the XY-axis components of the representative normal vector in the corrected query image when the XY plane is used for the corrected query image after projective transformation become zero (only the Z-axis component), and performs homography transformation on the query image using the homography matrix.

[0049] FIG. 9 is a diagram showing an example in which a geometric transformation based on a representative normal vector is applied to a query image in the present embodiment. According to this figure, by performing a geometric transformation based on the representative normal vectors representing a plurality of objects in the query image, the entire query image including the plurality of objects is in a state suitable for character recognition, and it can be seen that the accuracy of character recognition for the entire query image will be improved. Also, it can be seen that even after the image transformation, the positional relationship in which the positions of the smartphone and the product information card corresponding to the smartphone are vertically aligned is generally maintained. When the projective transformation of the query image based on the representative normal vector is completed, the process proceeds to step S110.

[0050] In step S110, the characters captured in the corrected query image are recognized. The character recognition unit 28 recognizes the characters captured in the corrected query image by performing optical character recognition (OCR) processing on the corrected query image obtained in step S109. Then, the character recognition unit 28 outputs the character recognition result, and the processing shown in this flowchart ends. According to the processing shown in this flowchart, by having the above-described processing flow, it is possible to improve the recognition accuracy of the characters described in the plurality of objects captured in the query image in one image transformation and one character recognition process for the query image.

[0051] <Effect> According to the image processing apparatus, method, and program according to the present embodiment, it is possible to perform an image transformation suitable for character recognition on a query image in which a plurality of regions are captured, and thus it is possible to improve the accuracy of character recognition for the query image in which a plurality of regions are captured.

[0052] <Variation> In the above-described embodiment, the processing when the shape data has a range has been described as an example. However, as described above, the shape data may be data that uniquely determines at least the front shape of an object belonging to the same type or attribute. When such shape data is used, the calculation of the minimum ratio normal vector and the maximum ratio normal vector for each object is omitted, and the representative normal vector is determined based on the normal vectors obtained one by one for each of the plurality of objects.

Description of Reference Numerals

[0053] 1 Image processing apparatus

Claims

1. Query image acquisition means for acquiring a query image in which a plurality of objects are imaged; Shape data acquisition means for acquiring shape data related to the front shape of the object; For each of the plurality of objects imaged in the query image, according to the comparison result between the front shape of the object specified according to the shape data and the object image in the query image, normal direction estimation means for estimating the normal direction indicating the front direction of the object; Representative normal direction determination means for determining a representative normal direction representing the entire query image based on the plurality of normal directions estimated for the plurality of objects; Conversion means for generating a corrected query image by geometrically deforming the query image based on the representative normal direction; An image processing apparatus comprising:

2. The shape data is shape data having a range for objects related to the same type or attribute, The normal direction estimation means calculates, for each of the plurality of objects, the normal direction related to the minimum value and the normal direction related to the maximum value in the range of the shape data, The representative normal direction determination means determines, as the representative normal direction, a normal direction representing the plurality of normal directions related to the minimum value in the range of the shape data estimated for the plurality of objects and the plurality of normal directions related to the maximum value in the range of the shape data estimated for each of the plurality of objects, The image processing apparatus according to claim 1.

3. The apparatus further comprises homography matrix calculation means for calculating a homography matrix between the front shape of the object specified according to the shape data and the object image in the query image for each of the plurality of objects, The normal direction estimation means estimates the normal direction of the object for each of the plurality of objects based on the homography matrix calculated by the homography matrix calculation means, The image processing apparatus according to claim 1.

4. The shape data is shape data having a range for objects related to the same type or attribute, The homography matrix calculation means calculates, for each of the plurality of objects, the homography matrix related to the minimum value and the homography matrix related to the maximum value in the range of the shape data, For each of the plurality of objects, the normal direction estimating means calculates the normal direction corresponding to the minimum value and the normal direction corresponding to the maximum value within the range of the shape data. For the plurality of objects, the representative normal direction determining means determines, as the representative normal direction, a normal direction that represents the plurality of normal directions corresponding to the minimum value within the range of the shape data estimated for the plurality of objects and the plurality of normal directions corresponding to the maximum value within the range of the shape data estimated for each of the plurality of objects. The image processing apparatus according to claim 3.

5. The shape data includes the aspect ratio of the object held in advance. The range is a range of aspect ratios. The image processing apparatus according to claim 2 or 4.

6. Query image acquisition means for acquiring a query image in which a plurality of objects are imaged; Object detection means for detecting a plurality of objects imaged in the query image by using a machine learning model that outputs the position or range in the query image where an object related to the type or attribute of the object for which character recognition is desired is imaged in response to an input of the type or attribute of the object for which character recognition is desired; For each of the plurality of objects imaged in the query image, normal direction estimating means for estimating a normal direction indicating the front direction of the object; Based on the plurality of normal directions estimated for the plurality of objects, representative normal direction determining means for determining a representative normal direction representing the entire query image; Conversion means for generating a corrected query image by geometrically deforming the query image based on the representative normal direction; Character recognition means for performing character recognition processing on the corrected query image to recognize characters described in the object for which character recognition is desired. An image processing apparatus comprising:

7. For each of the plurality of objects imaged in the query image, the normal direction estimating means estimates a normal vector indicating the front direction of the object. For the plurality of objects, the representative normal direction determining means determines the average of the plurality of normal vectors estimated for the plurality of objects as the representative normal direction. The image processing apparatus according to claim 1 or 6.

8. The conversion means performs projective transformation on the query image so that the representative normal direction becomes the front direction of the image after conversion. The image processing apparatus according to claim 1 or 6.

9. A computer performs a query image acquisition step of acquiring a query image in which a plurality of objects are imaged, a shape data acquisition step of acquiring shape data related to the front shape of the object, for each of the plurality of objects imaged in the query image, according to the comparison result between the front shape of the object specified according to the shape data and the object image in the query image, a normal direction estimation step of estimating the normal direction indicating the front direction of the object, a representative normal direction determination step of determining a representative normal direction representing the entire query image based on the plurality of normal directions estimated for the plurality of objects, a conversion step of generating a corrected query image by geometrically deforming the query image based on the representative normal direction, an image processing method for executing. **Claim 10**: A computer performs a query image acquisition step of acquiring a query image in which a plurality of objects are imaged, an object detection step of detecting a plurality of objects imaged in the query image by using a machine learning model that outputs the position or range in the query image where an object related to the type or attribute of the object to be the target of character recognition is imaged, in response to the input of the type or attribute of the object to be the target of character recognition, a normal direction estimation step of estimating, for each of the plurality of objects imaged in the query image, the normal direction indicating the front direction of the object, a representative normal direction determination step of determining a representative normal direction representing the entire query image based on the plurality of normal directions estimated for the plurality of objects, a conversion step of generating a corrected query image by geometrically deforming the query image based on the representative normal direction, a character recognition step of recognizing the characters described on the object to be the target of character recognition by performing character recognition processing on the corrected query image, an image processing method for executing. **Claim 11** A computer is provided with a query image acquisition means for acquiring a query image in which a plurality of objects are imaged, a shape data acquisition means for acquiring shape data related to the front shape of the object, For each of the plurality of objects imaged in the query image, a normal direction estimation means for estimating a normal direction indicating the front direction of the object according to a comparison result between the front shape of the object specified according to the shape data and the object image in the query image; A representative normal direction determination means for determining a representative normal direction representing the entire query image based on the plurality of normal directions estimated for the plurality of objects; A conversion means for generating a corrected query image by geometrically deforming the query image based on the representative normal direction; An image processing program for functioning as.

12. A computer, A query image acquisition means for acquiring a query image in which a plurality of objects are imaged; Using a machine learning model that outputs the position or range in the query image in which an object related to the type or attribute is imaged for an input of the type or attribute of the object for which character recognition is desired, an object detection means for detecting a plurality of objects imaged in the query image; For each of the plurality of objects imaged in the query image, a normal direction estimation means for estimating a normal direction indicating the front direction of the object; A representative normal direction determination means for determining a representative normal direction representing the entire query image based on the plurality of normal directions estimated for the plurality of objects; A conversion means for generating a corrected query image by geometrically deforming the query image based on the representative normal direction; A character recognition means for recognizing the characters described in the object for which character recognition is desired by performing character recognition processing on the corrected query image; An image processing program for functioning as.

Citation Information

Patent Citations

  • Image processor

    JP2002207963A

  • Image processor, method, program, and storage medium

    JP2008077489A

  • Image processing method, display device, and inspection system

    JP2017168077A

  • Information processor and program

    JP2018198030A

  • Recognition device, recognition method and program

    JP2020160814A