Output device, output method, and program

The output device and method use a learning model to enhance the detection of image deformation by determining projective transformation, addressing the limitations of low-resolution and noisy images in existing techniques.

JP7708966B2Active Publication Date: 2025-07-15RAKUTEN GROUP INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024510299
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-02-14
Publication Date
2025-07-15
Estimated Expiration
2043-02-14

AI Technical Summary

Technical Problem

Existing image processing techniques struggle to accurately determine deformation of detection targets in images with low resolution or noise, leading to failed feature point matching and improper recognition of shape deformation.

Method used

An output device and method utilizing a learning model to determine projective transformation of detection targets in images, including a reception unit, determination unit, and output unit, which uses a neural network to extract region candidates and output information on projective transformation.

Benefits of technology

Enhances the ability to accurately determine whether detection targets in images are deformed, improving the recognition of shape transformations such as rotation, enlargement, reduction, and shear.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708966000009
    Figure 0007708966000009
  • Figure 0007708966000010
    Figure 0007708966000010
  • Figure 0007708966000011
    Figure 0007708966000011
Patent Text Reader

Abstract

Provided is an output device comprising: a reception unit that receives an input of an image obtained by imaging a detection object; a determination unit that determines, by inputting the image to a learning model, information which relates to projective transformation of the detection object; and an output unit that outputs the information which relates to the projective transformation and which has been determined by the determination unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an output device, an output method, and a program.

Background Art

[0002] It is known that a homography matrix can be calculated by extracting feature points from two images and matching the extracted feature points. For example, Patent Document 1 describes an image processing apparatus that extracts a feature point pair from two images and calculates a homography matrix using the extracted feature points.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] For example, in order to examine whether a detection target shown in a captured image is deformed, feature points of the detection target are extracted from the image in which the detection target is shown, and the extracted feature points are compared with the feature points of the detection target having a correct shape, so that it is conceivable to detect the presence or absence of deformation of the detection target. However, when extracting feature points from an image as in the technique described in Patent Document 1, if the resolution of the image is low or the image has a lot of noise, the matching of the feature points may fail, and it may not be possible to appropriately recognize the presence or absence of deformation.

[0005] Therefore, an object of the present disclosure is to provide an output device, an output method, and a program that can more appropriately determine whether a detection target shown in a captured image is deformed.

Means for Solving the Problems

[0006] An output device according to one aspect of the present invention includes a reception unit that receives an input of an image in which a detection target is photographed, a determination unit that determines information related to the projective transformation of the detection target by inputting the image into a learning model, and an output unit that outputs the information related to the projective transformation determined by the determination unit.

Effect of the Invention

[0007] According to the present disclosure, it is possible to provide an output device, an output method, and a program that can more appropriately determine whether a detection target shown in a photographed image is deformed.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Mode for Carrying Out the Invention

[0009] Embodiments of the present invention will be described with reference to the accompanying drawings. In each figure, those denoted by the same reference numerals have the same or similar configurations.

[0010] <System Configuration> FIG. 1 is a diagram showing an example of an image determination system according to the present embodiment. The image determination system 1 includes an information processing apparatus 10 and a terminal 20. The information processing apparatus 10 and the terminal 20 are connected via a wireless or wired communication network N and can communicate with each other.

[0011] The information processing apparatus 10 is an apparatus that outputs information related to projective transformation (homography transformation) indicating how the detection target shown in the image is projected and transformed as compared with the original shape of the detection target. The detection target has a predetermined shape, and includes, for example, a logo, a mark, a symbol, an icon, a sign, text, and the like. The original shape of the detection target may be referred to as the correct shape of the detection target. In the following description, the case where the detection target is a logo will be described as an example, but the present embodiment is not limited thereto.

[0012] The information related to projective transformation may be, for example, information indicating a method of projective transformation or information indicating whether the shape of the detection target is projected and transformed. The information indicating the method of projective transformation may be, for example, the values of the components in the homography matrix (projective transformation matrix), information indicating the method of projective transformation (for example, rotating an image 30 degrees clockwise), or information indicating the coordinates of a plurality of feature points in the detection target.

[0013] The information processing apparatus 10 may be configured from one or more physical servers or the like, may be configured using virtual servers operating on a hypervisor, or may be configured using cloud servers.

[0014] The terminal 20 is a terminal operated by a user who uses the image determination system, and is, for example, a personal computer (PC), a notebook PC, a smartphone, a tablet terminal, a mobile phone, or the like. Various data output from the information processing apparatus 10 is displayed on the screen of the terminal 20. Further, the user can operate the information processing apparatus 10 via the terminal 20.

[0015] When the information processing apparatus 10 inputs an image in which a detection target is shown, it determines information regarding projective transformation using a learning model that has been learned to output information regarding projective transformation.

[0016] FIG. 2 is a diagram showing an example of projective transformation of a logo. The logo L1 shows the correct shape of the logo. The logo L2 shows a state in which the logo L1 is reduced in the x-axis direction. The logo L3 shows a state in which the logo L1 is reduced in the y-axis direction. The logo L4 shows a state in which the logo L1 is sheared (skewed) in the y-axis direction. The logo L5 shows a state in which the logo L1 is rotated.

[0017] The use of the image determination system 1 is arbitrary, but for example, it may be used by a company to check whether another company is appropriately using its own logo. For example, assume a case where Company B, which is a business partner of Company A, posts logo A indicating service A of Company A at the storefront or publishes it in a printed material. Also, assume that logo A is the same as logo L1 in FIG. 2. Company A hopes that Company B uses logo A in the correct shape when using logo A, but depending on Company B, due to reasons such as printing errors, logo A may be used in a slightly distorted state (for example, the state of logo L2). In such a case, the user of Company A can easily discover a case where logo A is being used in a deformed state by using the image determination system 1.

[0018] <Hardware Configuration> FIG. 3 is a diagram showing a hardware configuration example of the information processing apparatus 10. The information processing apparatus 10 includes a processor 11 such as a CPU (Central Processing Unit) and a GPU (Graphical Processing Unit), a memory (for example, RAM or ROM), a storage device 12 such as an HDD (Hard Disk Drive) and / or an SSD (Solid State Drive), a network IF (Network Interface) 13 that performs wired or wireless communication, an input device 14 that receives an input operation, and an output device 15 that outputs information. The input device 14 is, for example, a keyboard, a touch panel, a mouse, and / or a microphone, etc. The output device 15 is, for example, a display, a touch panel, and / or a speaker, etc.

[0019] <Functional block configuration> FIG. 4 is a diagram showing a functional block configuration example of the information processing apparatus 10. The information processing apparatus 10 includes a storage unit 100, a reception unit 101, a determination unit 102, an output unit 103, and a learning unit 104. The storage unit 100 can be realized by using the storage device 12 included in the information processing apparatus 10. Also, the reception unit 101, the learning unit 104, and the output unit 103 can be realized by the processor 11 of the information processing apparatus 10 executing a program stored in the storage device 12. Further, the program can be stored in a storage medium. The storage medium storing the program may be a non-transitory computer readable medium (Non-transitory computer readable medium). The non-transitory storage medium is not particularly limited, but may be, for example, a storage medium such as a USB (Universal Serial Bus) memory or a CD-ROM (Compact Disc Read-Only Memory).

[0020] The storage unit 100 stores a learning model. The learning model includes information for determining a model structure and various parameter values.

[0021] The reception unit 101 receives the input of an image in which the object to be detected is photographed. For example, the reception unit 101 may receive the input of image data via the terminal 20. The reception unit 101 may be referred to as an input unit.

[0022] The determination unit 102 determines information related to the projective transformation of the detection target by inputting the image received by the reception unit 101 into the learning model. The learning model may be a model using a neural network. Also, the determination unit 102 may determine the display position of a bounding box (hereinafter referred to as BBOX (Bounding Box)) indicating the position where the detection target exists on the image by inputting the image into the learning model.

[0023] Also, the determination unit 102 may determine the type of the detection target shown in the image by inputting the image into the learning model. The type of the detection target may be referred to as the class of the detection target. When the learning model has the ability to detect one type of detection target, the determination unit 102 may determine, as the type of the detection target, information indicating whether the one type of detection target is shown in the image. Also, when the learning model has the ability to detect two or more types of detection targets, the determination unit 102 may determine, as the type of the detection target, information indicating which detection target is shown in the image.

[0024] The output unit 103 outputs the information related to the projective transformation determined by the determination unit 102. The output unit 103 may display the information related to the projective transformation on the screen of the terminal 20. Also, the output unit 103 may output the display position of the BBOX determined by the determination unit 102. Also, the output unit 103 may superimpose and display the BBOX on the image. Also, the output unit 103 may output the type of the detection target determined by the determination unit 102.

[0025] Further, the output unit 103 may output information indicating whether the shape of the detection target has been projective-transformed or information indicating whether the shape of the detection target has been deformed from the original shape based on the information related to the projective transformation determined by the determination unit 102.

[0026] The learning unit 104 trains a learning model using teacher data in which an image of the detection target is associated with information related to the projective transformation of the detection target.

[0027] <Processing procedure> Subsequently, the processing procedure performed by the information processing apparatus 10 will be specifically described.

[0028] FIG. 5 is a diagram showing an overview of the learning model. In FIG. 5, it is assumed that the detection target is the logo L100. The learning model M100 is a model using a neural network, and the structure of the model may be a structure in which two neural networks, the network N100 and the network N200, are connected.

[0029] Here, the network N100 may be a network having the ability to extract a region candidate of an object appearing in the input image. Further, the network N200 may be a network having the ability to output information related to the projective transformation of the detection target (hereinafter referred to as "projective transformation information") from the region candidate extracted by the network N100.

[0030] More specifically, the network N100 may have the ability to extract a region (region candidate) in the entire image where it is estimated that some object is depicted. For example, when an image P100 in which a logo L100 is depicted is input, the network N100 recognizes, in the entire image P100, the background region and the region where some object is depicted, and may extract the region where some object is depicted (here, the region where the logo L100 is depicted) as a region candidate. Further, the network N200 may output projective transformation information indicating how the logo L100 depicted in the region candidate extracted by the network N100 is projectively transformed as compared to the original shape.

[0031] Note that the network N200 may further output "class information" indicating the type of the detection target depicted in the region candidate extracted by the network N100. For example, when the image P100 is input, the network N200 may output information indicating that the detection target depicted in the image P100 is the logo L100. Further, the network N200 may further output "BBOX information" indicating the region in the image P100 where the detection target is depicted, from the region candidate extracted by the network N100.

[0032] FIG. 6 is a flowchart showing an outline of a processing procedure when the information processing apparatus 10 trains a learning model. In the description of FIGS. 6 and 7, the learning model is assumed to output three types of information: class information, BBOX information, and projective transformation information, but it is not limited thereto. For example, the learning model may output only the projective transformation information.

[0033] The reception unit 101 receives the input of learning data via the terminal 20 (S10). The learning data (also referred to as teacher data) is data in which the image data of an image in which a detection target is depicted is associated with the class of the detection target, the display position of the BBOX, and the projective transformation information.

[0034] Subsequently, the learning unit 104 generates a learning model by training the model using the learning data (S11). When the training of the model is completed, the learning unit 104 stores various parameters, which are the training results, in the storage unit 100.

[0035] FIG. 7 is a flowchart showing an outline of a processing procedure when the information processing apparatus 10 determines projective transformation information from an image.

[0036] The reception unit 101 receives an input of image data from the user via the terminal 20 (S20).

[0037] Subsequently, the determination unit 102 inputs the image data into the learning model, and acquires information indicating a class, BBOX information, and projective transformation information from the learning model, thereby determining class information, BBOX information, and projective transformation information.

[0038] The output unit 103 outputs the class information, BBOX information, and projective transformation information determined by the determination unit 102 to the screen of the terminal 20. Note that the output unit 103 may transmit the class information, BBOX information, and projective transformation information to another information processing apparatus instead of outputting the acquired information to the terminal 20.

[0039] <Specific Example> Subsequently, a plurality of specific examples of the configuration of the learning model will be described. In the following specific examples, the learning model is assumed to be a neural network that gives a neural network called Faster R-CNN (Regions with Convolutional Neural Networks) the ability to output projective transformation information. Also, the detection target is assumed to be the logo shown in FIG. 2.

[0040] <Specific Example 1> FIG. 8 is a diagram showing a learning model (specific example 1). The FC layer means a fully connected layer. When an image is input to the network N100, the learning model M100 in the specific example 1 may output, as projective transformation information, each component of a homography matrix (projective transformation matrix) estimated to be applied to the logo (detection target) before projective transformation from the network N230 connected to the network N100.

[0041] Further, the learning model M100 may include a network N210 connected to the network N100 for determining a BBOX surrounding the logo (detection target) from the region candidates of the object, and a network N220 connected to the network N100 for determining the type of the logo (detection target) from the region candidates of the object. At this time, the determination unit 102 may determine the BBOX and the type of the detection target by inputting the image to the learning model, and the output unit 103 may output the BBOX and the type of the detection target determined by the determination unit 102 (the same applies to the specific example 2 described later).

[0042] In the specific example 1, the network N100 and the network N230 may be referred to as a first network and a second network, respectively. Also, the network N210 and the network N220 may be referred to as a third network and a fourth network, respectively.

[0043] Here, assuming that the coordinates on the image before projective transformation are (x, y), the coordinates on the image after projective transformation are (x′, y′), and the homography matrix is H, the coordinates (x′, y′) can be expressed by Equation (1). Also, the homography matrix can be expressed by Equation (2). Note that from Equation (1), s = h 31 ×x + h 32 ×y + h 33 becomes. Also, it is known that the value of h 33 in Equation (2) may be 1.

[0044]

Number

[0045]

Number

[0046] The learning of the learning model M100 in Specific Example 1 may be performed according to the following procedure. First, the learning unit 104 generates a homography matrix by randomly generating nine components. At this time, h 33 may always be set to "1". Subsequently, the learning unit 104 generates an image obtained by synthesizing the logo image subjected to projective transformation using the generated homography matrix with a background image where the logo image does not exist. Subsequently, the learning unit 104 uses the generated image as input data, and uses the class information corresponding to the logo image, the position of the BBOX indicating the region where the logo image exists in the image, and the nine components of the homography matrix used when performing projective transformation on the logo image as output data to generate learning data. Note that the class information and the position of the BBOX may be specified by the user who generates the learning model. The learning unit 104 generates a large number of learning data by repeating the process of generating learning data.

[0047] Subsequently, the learning unit 104 uses the large number of generated learning data to train the learning model M100. The loss function used for training may use, for example, RMSLE (Root Mean Squared Logarithmic Error) that uses the mean squared error, but is not limited thereto.

[0048] Regarding the learning of the learning model M100 described above, depending on the components of the generated homography matrix, there is a possibility that the logo image after projective transformation may represent an inappropriate shape, such as becoming a point. Also, since it is necessary to vary the nine components of the homography matrix in various ways, the learning data may become enormous. Therefore, the learning data may not include the values of the components that cause the logo image after projective transformation to represent an inappropriate shape.

[0049] As described above, when a company uses the information processing apparatus 10 to check whether another company is appropriately using its logo, it is assumed that the patterns in which the logo is deformed are limited to deformations that can be represented by linear transformations, such as rotation, enlargement, reduction, and shear.

[0050] Here, if the coordinates on the image before linear transformation are (x, y), the coordinates on the image after linear transformation are (x′, y′), and the matrix representing the linear transformation is L, then the coordinates (x′, y′) can be expressed by Equation (3). Also, the matrix representing the linear transformation can be expressed by Equation (4).

[0051]

Equation

[0052]

Equation

[0053] As shown in the mathematical formula (4), since the components of the matrix representing the linear transformation are four, the amount of learning data required for learning the learning model M100 can be significantly reduced as compared with the case of estimating nine components.

[0054] Therefore, the determination unit 102 may determine information related to the linear transformation of the detection target (hereinafter referred to as "linear transformation information") by inputting the image into the learning model. Further, the learning model M100 may be a neural network including a network N100 that extracts a region candidate of an object shown in the image and a network N230 that outputs linear transformation information of a logo (detection target) from the region candidate of the object. Further, the linear transformation information output from the learning model M100 is the four components of the matrix representing the linear transformation applied to the logo (l in Equation 4 11 ~l 22 , or h in Equation 2 11 ~h 22 ).

[0055] The learning of the learning model M100 in this case may be performed according to the following procedure. First, the learning unit 104 is the four components (h in Equation 2 11 ~h 22 , or l in Equation 4 11 ~l 22A homography matrix (or a matrix representing a linear transformation) is generated by randomly generating ). Subsequently, the learning unit 104 generates an image obtained by synthesizing the linearly transformed logo image using the generated homography matrix (or the matrix representing the linear transformation) onto a background image where the logo image does not exist. Subsequently, the learning unit 104 uses the generated image as input data, and uses the class information corresponding to the logo image, the position of the BBOX indicating the region where the logo image exists in the image, and the four components of the homography matrix (or the matrix representing the linear transformation) used when linearly transforming the logo image as output data to generate training data. Note that the class information and the position of the BBOX may be specified by the user who generates the learning model. The learning unit 104 generates a plurality of pieces of training data by repeating the process of generating the training data. Subsequently, the learning unit 104 uses the generated plurality of pieces of training data to train the learning model M100.

[0056] By limiting it to linear transformation, the elements of the matrix output by the learning model M100 are limited to four components, so that the amount of training data can be significantly reduced, and the time required for training the learning model can be significantly shortened.

[0057] <Specific Example 2> FIG. 9 is a diagram showing a learning model (Specific Example 2). When an image is input to the network N100 in Specific Example 2, the network N231 connected to the network N100 outputs, as projective transformation information, a component representing rotation in the homography matrix, a component representing scale transformation (enlargement or reduction), and a component representing shear. The networks N210 and N220 are the same as in Specific Example 1. In Specific Example 2, the network N100 and the network N231 may be referred to as a first network and a second network, respectively. Also, the networks N210 and N220 may be referred to as a third network and a fourth network, respectively.

[0058] That is, the network N231 (second network) in the specific example 2 may include at least one of a network that outputs components of a homography matrix related to rotation, a network that outputs components of a homography matrix related to scale conversion, and a network that outputs components of a homography matrix related to shear. Further, when an image is input to the network N100, the learning model M100 may output, as information related to projective transformation, components of a homography matrix related to at least one of rotation, scale conversion, and shear, which are estimated to be applied to the logo (detection target) before projective transformation, from the network N231 connected to the network N100.

[0059] Figure 10 is a diagram showing four patterns of projective transformation. A in Figure 10 shows an example of rotating the logo by θ rot degrees clockwise. The homography matrix in this case is represented by Equation 5.

[0060]

Equation

[0061]

Equation

[0062]

Equation

[0063] [Number] The learning model M100 outputs the value of θ as a component of the homography matrix related to rotation, outputs the values of W and H as components of the homography matrix related to scale conversion, and outputs the value of θ as a component of the homography matrix related to shear in the y direction. rot The value of θ, outputs the values of W and H as components of the homography matrix related to scale conversion, outputs the value of θ as a component of the homography matrix related to shear in the y direction, and outputs the value of θ as a component of the homography matrix related to shear in the x direction. shear_y The value of θ, outputs the value of θ as a component of the homography matrix related to shear in the x direction. shear_x It may be configured to output. Also, for the values corresponding to the deformations that are irrelevant among these output values, the learning model M100 outputs a value indicating no deformation (specifically, θ rot = 0 degrees, W = 1, H = 1, θ shear_y = 0 degrees, θ shear_x = 0 degrees). For example, when the deformation of the logo is only rotation, the learning model M100 outputs the value of θ rot corresponding to the rotation angle (e.g., 10 degrees or 45 degrees, etc.), outputs 1 for each of the values of W and H, outputs 0 for the value of θ shear_y and outputs 0 for the value of θ shear_x It may be configured to output. Similarly, when the deformation of the logo is only expansion in the y direction, the learning model M100 outputs 0 for the value of θ rot outputs the expanded value (e.g., 1.5 or 2, etc.) for the value of W, outputs 1 for the value of H, outputs 0 for the value of θ shear_y and outputs 0 for the value of θ shear_x It may be configured to output.

[0064] The learning of the learning model M100 in Specific Example 2 may be performed according to the following procedure. First, the learning unit 104 determines the value of θ rot , the value of W, the value of H, the value of θ shear_y and the value of θ shear_xRandomly generate the value of. Subsequently, the learning unit 104 generates a homography matrix by multiplying the matrix represented by Equation (5), the matrix represented by Equation (6), the matrix represented by Equation (7), and the matrix represented by Equation (8). Subsequently, the learning unit 104 generates an image obtained by synthesizing the logo image subjected to projective transformation using the generated homography matrix with a background image where the logo image does not exist. Subsequently, the learning unit 104 uses the generated image as input data, the class information corresponding to the logo image, the position of the BBOX indicating the region where the logo image exists in the image, and θ rot values of, values of W, values of H, θ shear_y values of and θ shear_x values of and θ

[0065] to generate training data with the output data. Note that the class information and the position of the BBOX may be specified by the user who generates the learning model. The learning unit 104 generates a large number of training data by repeating the process of generating the training data.

[0066] When the deformation pattern of the logo is limited to any one of rotation, scale transformation in the y-axis direction, scale transformation in the x-axis direction, shear in the y-axis direction, and shear in the x-axis direction, the learning unit 104 rot values of, values of W, values of H, θ shear_y values of and θ shear_x When randomly generating the values of, the training data may be generated such that only any one of these values is changed and the other values become values indicating no deformation.

[0067] According to Specific Example 2, the amount of training data can be significantly reduced compared to Specific Example 1, and the time required for learning the learning model can be significantly shortened.

[0068] Note that in the above-described Specific Example 2, since the information processing apparatus 10 determines rotation, enlargement, reduction, and shear as the four patterns of projective transformation, it is synonymous with performing the determination of linear transformation. Therefore, in the description of Specific Example 2, the terms "projective transformation" and "projective transformation information" may be replaced with the terms "linear transformation" and "projective transformation information", respectively.

[0069] <Specific Example 3> FIG. 11 is a diagram showing a learning model (Specific Example 3). When an image is input to the network N100 in the learning model M100 in Specific Example 3, the network N232 outputs, as projective transformation information, the coordinates of a plurality of feature points existing in the logo (detection target) after projective transformation, which are relative coordinates from a predetermined reference point. The networks N210 and N220 are the same as those in Specific Example 1.

[0070] Further, the learning model M100 may include a network N210 connected to the network N100 for determining a BBOX that encloses a logo (detection target) from the region candidates of an object, and a network N220 connected to the network N100 for determining the type of the logo (detection target) from the region candidates of the object. Also, the network N232 may be connected to the network N220. At this time, the determination unit 102 may determine the BBOX and the type of the detection target by inputting an image to the learning model, and the output unit 103 may output the BBOX and the type of the detection target determined by the determination unit 102.

[0071] In Specific Example 3, the network N100 and the network N232 may be referred to as a first network and a second network, respectively. Also, the networks N210 and N220 may be referred to as a third network and a fourth network, respectively.

[0072] FIG. 12 is a diagram for explaining the feature points to be detected. As shown in FIG. 12, for logo L1, the positions of the relative coordinates (x, y) of four feature points P1 to P4 are predetermined. Note that the point (0, 0) where the x-axis and the y-axis intersect may be used as the reference point, but it is not limited thereto. Any point can be used as the reference point. Also, the number of feature points is not limited to four. For example, the number of feature points may be three, or may be five or more. The positions of the feature points are arbitrary, but it is preferable to set them at positions as far outside as possible, such as the upper left end, upper right end, lower left end, and lower right end of the logo.

[0073] FIG. 13 is a diagram showing an example of the feature points after projective transformation. A in FIG. 13 shows the relative coordinates of the feature points (P1' to P4') when the logo is rotated. B in FIG. 13 shows the relative coordinates of the feature points (P1' to P4') when the logo is enlarged or reduced. C in FIG. 13 shows the relative coordinates of the feature points (P1' to P4') when the logo is sheared in the y direction. D in FIG. 13 shows the relative coordinates of the feature points (P1' to P4') when the logo is sheared in the x direction.

[0074] For example, when an image in which the logo shown in A of FIG. 13 is shown is input to the learning model M100, the relative coordinates of the feature points (P1' to P4') shown in A of FIG. 13 are output. Similarly, when an image in which the logo shown in D of FIG. 13 is shown is input to the learning model M100, the relative coordinates of the feature points (P1' to P4') shown in D of FIG. 13 are output.

[0075] The learning of the learning model M100 in Specific Example 3 may be performed according to the following procedure. First, the learning unit 104 generates a homography matrix by randomly generating nine components of Equation 2. Subsequently, the learning unit 104 generates an image obtained by synthesizing the logo image subjected to projective transformation using the generated homography matrix with a background image in which the logo image does not exist. Subsequently, the learning unit 104 calculates the relative coordinates of four feature points in the logo image after projective transformation. Subsequently, the learning unit 104 generates learning data using the generated image as input data, the class information corresponding to the logo image, the position of the BBOX indicating the region where the logo image exists in the image, and the relative coordinates of the four feature points as output data. Note that the class information and the position of the BBOX may be specified by the user who generates the learning model. The learning unit 104 generates a large number of learning data by repeating the process of generating learning data.

[0076] Subsequently, the learning unit 104 learns the learning model M100 using the large number of generated learning data. As the loss function used for learning, for example, RMSLE using the mean squared error may be used, but it is not limited thereto.

[0077] Note that, as described in Specific Example 1, the determination unit 102 may be configured to determine only the transformation that is a linear transformation. In this case, in the description of Specific Example 3, the terms "projective transformation" and "projective transformation information" may be replaced with the terms "linear transformation" and "projective transformation information", respectively. Further, when learning the learning model M100, the learning unit 104 may generate a homography matrix or a matrix related to a linear transformation by randomly generating four components (h11 to h22 of Equation 2, or l11 to l22 of Equation 4), and generate a logo image subjected to linear transformation using the generated matrix. Points not specifically mentioned may be the same as the description of the learning procedure in Specific Example 3 described above.

[0078] As shown in FIG. 11, in the learning model M100 in Specific Example 3, the network N232 is connected to the FC layer of the network N220 instead of the network N100. Since the network N220 is a network for determining a BBOX, it is highly likely that some information for estimating the position of the BBOX is extracted in the FC layer of the network N220. Therefore, by connecting the network N232 to the FC layer of the network N220, a part of the process for estimating the position of the detection target in the image can be made common. As a result, the arguments of the network can be reduced compared to the learning model M100 of Specific Example 1, and the learning time can be shortened.

[0079] <Summary> According to the embodiments described above, by determining the projective transformation information from the image in which the detection target is photographed, it becomes possible to more appropriately determine whether the detection target is deformed.

[0080] The embodiments described above are for facilitating the understanding of the present invention and are not for limiting and interpreting the present invention. The flowcharts, sequences, each element included in the embodiments, and their arrangements, materials, conditions, shapes, sizes, etc. described in the embodiments are not limited to those exemplified and can be changed as appropriate. Also, it is possible to partially substitute or combine the configurations shown in different embodiments.

[0081] Further, since the linear transformation is an example of the projective transformation, the projective transformation information in the present embodiment may include the linear transformation information.

[0082] <Supplementary Note> This embodiment may be expressed as follows.

[0083] <Supplementary Note 1> A reception unit that receives an input of an image in which a detection target is photographed, A determination unit that determines information related to the projective transformation of the detection target by inputting the image into a learning model, An output unit that outputs information related to the projective transformation determined by the determination unit; An output device having

[0084] <Appendix 2> The learning model is a neural network including a first network that extracts a region candidate of an object depicted in the input image, and a second network that outputs information related to the projective transformation of the detection target from the region candidate of the object. The determination unit determines information related to the projective transformation by inputting the image into the neural network. The output device according to Appendix 1.

[0085] <Appendix 3> When the neural network inputs the image into the first network, the neural network outputs, as information related to the projective transformation, components of a homography matrix that is presumed to have been applied to the detection target before projective transformation, from the second network connected to the first network. The output device according to Appendix 2.

[0086] <Appendix 4> The second network includes at least one of a network that outputs components of a homography matrix related to rotation, a network that outputs components of a homography matrix related to scale transformation, and a network that outputs components of a homography matrix related to shear. When the neural network inputs the image into the first network, the neural network outputs, as information related to the projective transformation, components of a homography matrix related to at least one of rotation, scale transformation, and shear that is presumed to have been applied to the detection target before projective transformation, from the second network connected to the first network. The output device according to Appendix 2.

[0087] <Appendix 5> The neural network includes a third network connected to the first network for determining a bounding box that surrounds the detection target from the region candidates of the object, and a fourth network connected to the first network for determining the type of the detection target from the region candidates of the object. The determination unit determines the bounding box and the type of the detection target by inputting the image into the neural network. The output unit outputs the bounding box and the type of the detection target determined by the determination unit. The output device according to appendix 3 or 4.

[0088] <Appendix 6> When the neural network inputs the image into the first network, it outputs, from the second network, as information regarding the projective transformation, the coordinates of a plurality of feature points existing in the detection target after the projective transformation, which are relative coordinates from a predetermined reference point. The output device according to appendix 2.

[0089] <Appendix 7> The neural network includes a third network connected to the first network for determining a bounding box that surrounds the detection target from the region candidates of the object, and a fourth network connected to the first network for determining the type of the detection target from the region candidates of the object. The second network is connected to the third network. The determination unit determines the bounding box and the type of the detection target by inputting the image into the neural network. The output unit outputs the bounding box and the type of the detection target determined by the determination unit. The output device according to appendix 6.

[0090] <Appendix 8> A learning unit that trains the learning model using teacher data associating an image in which a detection target is photographed with information related to the projective transformation of the detection target. The output device according to any one of Appendices 1 to 7.

[0091] <Appendix 9> By inputting the image, the determination unit determines a bounding box surrounding the detection target and the type of the detection target. The output unit outputs the bounding box and the type of the detection target determined by the determination unit. The output device according to Appendix 1.

[0092] <Appendix 10> A step of receiving an input of an image in which a detection target is photographed, A step of determining information related to the projective transformation of the detection target by inputting the image into a learning model, A step of outputting the determined information related to the projective transformation, An output method performed by an output device, including the above steps.

[0093] <Appendix 11> To a computer, A step of receiving an input of an image in which a detection target is photographed, A step of determining information related to the projective transformation of the detection target by inputting the image into a learning model, A step of outputting the determined information related to the projective transformation, A program for causing the above steps to be executed.

Explanation of Reference Numerals

[0094] 1 Image determination system, 10 Information processing device, 11 Processor, 12 Storage device, 13 Network IF, 14 Input device, 15 Output device, 20 Terminal, 100 Storage unit, 101 Reception unit, 102 Determination unit, 103 Output unit, 104 Learning unit, N Communication network

Claims

1. A receiving unit that receives an input of an image in which a detection target is photographed; A determination unit that determines a bounding box, a type of the detection target, and information related to a projective transformation of the detection target by inputting the image into a learning model; An output unit; characterized by comprising: The learning model is a neural network including a first network that extracts a region candidate of an object appearing in the input image, and a second network that outputs information related to the projective transformation of the detection target from the region candidate of the object; The neural network further includes a third network connected to the first network for determining the bounding box surrounding the detection target from the region candidate of the object, and a fourth network connected to the first network for determining the type of the detection target from the region candidate of the object; The determination unit determines the bounding box, the type of the detection target, and the information related to the projective transformation by inputting the image into the neural network; The output unit outputs the bounding box, the type of the detection target, and the information related to the projective transformation determined by the determination unit; An output device.

2. When the neural network inputs the image into the first network, the neural network outputs, as information related to the projective transformation, components of a homography matrix estimated to be applied to the detection target before the projective transformation, from the second network connected to the first network; The output device according to claim 1.

3. The second network includes at least one of a network that outputs components of a homography matrix related to rotation, a network that outputs components of a homography matrix related to scale transformation, and a network that outputs components of a homography matrix related to shear; When the neural network inputs the image into the first network, the neural network outputs, as information related to the projective transformation, components of a homography matrix related to at least one of rotation, scale transformation, and shear estimated to be applied to the detection target before the projective transformation, from the second network connected to the first network; The output device according to claim 1.

4. When the neural network inputs the image into the first network, the second network outputs, as information regarding the projective transformation, the coordinates of a plurality of feature points existing in the detection target after the projective transformation, which are relative coordinates from a predetermined reference point. The output device according to claim 1.

5. The neural network includes a third network connected to the first network for determining a bounding box that surrounds the detection target from the region candidates of the object, and a fourth network connected to the first network for determining the type of the detection target from the region candidates of the object. The second network is connected to the third network. The determination unit determines the bounding box and the type of the detection target by inputting the image into the neural network. The output unit outputs the bounding box and the type of the detection target determined by the determination unit. The output device according to claim 4.

6. A learning unit that trains the learning model using teacher data associating an image in which a detection target is photographed with information related to the projective transformation of the detection target. The output device according to claim 1.

7. The determination unit determines a bounding box that surrounds the detection target and the type of the detection target by inputting the image. The output unit outputs the bounding box and the type of the detection target determined by the determination unit. The output device according to claim 1.

8. A step of receiving an input of an image in which a detection target is photographed. A step of determining a bounding box, the type of the detection target, and information related to the projective transformation of the detection target by inputting the image into a learning model. A step of outputting. Including. The learning model is a neural network including a first network that extracts region candidates of an object appearing in the input image, and a second network that outputs information related to the projective transformation of the detection target from the region candidates of the object. The neural network further includes a third network connected to the first network for determining the bounding box that surrounds the detection target from the region candidates of the object, and a fourth network connected to the first network for determining the type of the detection target from the region candidates of the object. The determining step determines the bounding box, the type of the detection target, and information related to the projective transformation by inputting the image into the neural network. The outputting step outputs the bounding box, the type of the detection target, and information related to the projective transformation determined in the determining step. An output method performed by an output device.

9. A computer receives an input of an image in which a detection target is photographed; determines a bounding box, the type of the detection target, and information related to the projective transformation of the detection target by inputting the image into a learning model; outputs; and causes to execute, wherein the learning model is a neural network including a first network that extracts a region candidate of an object shown in the input image, and a second network that outputs information related to the projective transformation of the detection target from the region candidate of the object. The neural network further includes a third network connected to the first network and determining the bounding box surrounding the detection target from the region candidate of the object, and a fourth network connected to the first network and determining the type of the detection target from the region candidate of the object. The determining step determines the bounding box, the type of the detection target, and information related to the projective transformation by inputting the image into the neural network. The outputting step outputs the bounding box, the type of the detection target, and information related to the projective transformation determined in the determining step. A program.

Citation Information

Patent Citations

  • Target identification method and device

    CN112597887A

  • Image processing device, image processing method and image processing program

    JP2013214155A

  • Object detection program and object detection device

    JP2020027405A

  • Machine learning models for direct homography regression for image rectification

    US11481683B1

  • Systems and methods for real time screen display coordinate and shape detection

    US20200336656A1