Detection method, detection device, and program

JP7901713B1Active Publication Date: 2026-08-06ZOSHINKAI PUBLISHERS
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ZOSHINKAI PUBLISHERS
Filing Date
2025-03-31
Publication Date
2026-08-06

AI Technical Summary

Benefits of technology

【0011】 本開示によれば、検出方法、検出装置およびプログラムは、画像の中から矩形状の対象物が写る対象領域を精度良く検出できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007901713000001_ABST
    Figure 0007901713000001_ABST
Patent Text Reader

Abstract

This invention provides a detection method that can accurately detect areas within an image that contain rectangular objects. [Solution] The detection method is performed by a computer to detect the area occupied by a rectangular object from an image. The object includes an L-shaped mark attached to its corner. The detection method comprises: step S2 of extracting multiple hexagonal figures from the image; step S3 of determining whether the characteristic conditions defining the features of the L-shaped mark are met for each of the multiple figures; and steps S4 to S10 of identifying the area based on the position of the figure that satisfies the characteristic conditions among the multiple figures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a detection method, a detection device, and a program.

Background Art

[0002] Conventionally, an educational system in which an instructor corrects an answer sheet or notes submitted by a learner is known. In recent years, due to the spread of the Internet, a learner transmits an image obtained by photographing a rectangular object such as an answer sheet or notes to an instructor. The instructor performs a correction operation using the image.

[0003] An image may include not only a target area (also referred to as a "foreground area") where an object appears but also a background area. When an instructor corrects an answer, the background area is unnecessary. Therefore, techniques for extracting a target area from an image have been developed. For example, Japanese Patent Application Laid-Open No. 2024-121891 (Patent Document 1) discloses a technique for extracting a region corresponding to a printed matter from a photographed image using an edge extraction algorithm.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] An object can be photographed in various environments. For example, a photographer can photograph an object placed on a table of the same color. In an image obtained by photographing in such an environment, the change in luminance is small at the boundary between the target area where the object appears and the background area. When the technique described in Patent Document 1 is applied to such an image, the target area cannot be accurately extracted from the image.

[0006] Furthermore, the photographer may photograph the object with the photographer's shadow cast upon it. Alternatively, the photographer may photograph the object with other objects (e.g., writing instruments and text) scattered around it. In images obtained by photography under such conditions, the brightness may vary significantly not only at the boundary between the object area and the background area, but also at the outlines of the shadow and other objects. When the technique described in Patent Document 1 is applied to such images, numerous edges are extracted, resulting in inaccurate extraction of the target area.

[0007] This disclosure has been made in view of the above circumstances, and its purpose is to provide a detection method, detection device, and program that can accurately detect a target area in an image in which a rectangular object is captured. [Means for solving the problem]

[0008] A detection method relating to one aspect of this disclosure is performed by a computer to detect a target area occupied by a rectangular object in an image. The object includes an L-shaped mark attached to its corner. The detection method comprises: extracting a plurality of hexagonal figures from the image; determining whether each of the plurality of figures satisfies the characteristic conditions that define the features of the L-shaped mark; and identifying the target area based on the position of the figure among the plurality of figures that satisfies the characteristic conditions.

[0009] A detection device relating to one aspect of this disclosure detects a target area occupied by a rectangular object in an image. The object includes an L-shaped mark attached to its corner. The detection device comprises an extraction unit that extracts a plurality of hexagonal figures from the image, a determination unit that determines whether or not a characteristic condition defining the features of the L-shaped mark is met for each of the plurality of figures, and an identification unit that identifies the target area based on the position of the figure that satisfies the characteristic condition among the plurality of figures.

[0010] A program relating to one aspect of this disclosure causes a computer to execute the detection method described above. [Effects of the Invention]

[0011] According to this disclosure, the detection method, detection device, and program can accurately detect a target area in an image that contains a rectangular object. [Brief explanation of the drawing]

[0012] [Figure 1] This figure shows an example of the configuration of an educational system to which the detection method according to the embodiment is applied. [Figure 2] This is a diagram showing an example of an image that includes an answer sheet. [Figure 3] This figure shows an example of the functional configuration of the first terminal according to the embodiment. [Figure 4] This figure shows an example of the processing flow of the first terminal according to the embodiment. [Figure 5] This figure shows an example of a hexagonal shape. [Figure 6] This diagram illustrates the preprocessing steps used to identify interpolation points. [Figure 7] This diagram illustrates how to identify the completion points. [Figure 8] This diagram shows the functional configuration of the first terminal according to Modification Example 1. [Figure 9] This figure shows an example of the processing flow of the first terminal according to Modification Example 1. [Figure 10] This figure shows another example of the processing flow of the first terminal according to Modification Example 1. [Modes for carrying out the invention]

[0013] Embodiments of the present invention will be described in detail with reference to the drawings. Note that identical or corresponding parts in the drawings are denoted by the same reference numerals, and their descriptions will not be repeated.

[0014] The detection method according to the present disclosure is applicable to systems in various fields that use an image obtained by photographing a rectangular object. Typically, the detection method according to the present disclosure is applicable to an educational system in which a grading operation is performed using an image obtained by photographing an object such as an answer sheet or a notebook. Hereinafter, an example in which the detection method according to the present disclosure is applied to an educational system will be described. However, the system to which the detection method according to the present disclosure is applied is not limited to an educational system. For example, the detection method according to the present disclosure may be applied to a system in which an accounting operation (or clerical operation) is performed using an image obtained by photographing a form, a system in which papers are photographed by digitization and digital processing is performed, and the like. Thus, the object typically includes a rectangular sheet of paper.

[0015] <System Configuration> FIG. 1 is a diagram showing an example of the configuration of an educational system to which the detection method according to the embodiment is applied. As shown in FIG. 1, the educational system 1 includes a first terminal 10, a server 20, and a second terminal 30. The first terminal 10, the server 20, and the second terminal 30 can communicate with each other via a network 40. The network 40 includes, for example, the Internet.

[0016] The first terminal 10 has a general-purpose computer architecture. The first terminal 10 includes, for example, a smartphone, a tablet, and a personal computer. The first terminal 10 is used by a learner. Specifically, the first terminal 10 generates image data (hereinafter simply referred to as "image 50") by photographing an answer sheet 2 on which an answer has been filled in by the learner in response to an operation by the learner, and transmits the image 50 to the server 20. The first terminal 10 is an example of the "detection device" of the present disclosure. The answer sheet 2 is rectangular and is an example of the "object" of the present disclosure.

[0017] The server 20 can be realized as one or more computers, one or more virtual machines constructed on a cloud environment, or a combination thereof. The server 20 stores the image 50 received from the first terminal 10.

[0018] The second terminal 30 has a general-purpose computer architecture. The second terminal 30 includes, for example, a personal computer, a tablet, and a smartphone. The second terminal 30 is used by an instructor. Specifically, the second terminal 30 accesses the server 20 according to an operation by the instructor and displays the image 50. Further, the second terminal 30 may transmit the correction result for the answer sheet 2 shown in the image 50 to the first terminal 10 according to an operation by the instructor.

[0019] As shown in FIG. 1, the first terminal 10 includes a processor 11, a memory 12, a storage 13, a communication interface (IF) 14, a camera 15, an input device 16, and a display 17.

[0020] The processor 11 is composed of a CPU (Central Processing Unit), an MPU (Micro Processing Unit), etc. The processor 11 realizes various processes by expanding and executing the program stored in the storage 13 in the memory 12.

[0021] The memory 12 is typically a volatile storage device such as a DRAM (Dynamic Random Access Memory), and stores programs read from the storage 13 and the like.

[0022] The storage 13 is typically a non-volatile magnetic storage device such as a hard disk drive. The storage 13 stores the application 18 executed by the processor 11.

[0023] The application 18 is downloaded from an app store not shown. The application 18 includes a set of instructions regarding detection of a target area where the answer sheet 2 from the image 50 is shown, correction of the image 50, and output of the image 50.

[0024] The communication interface 14 exchanges data with external devices (including the server 20 and the second terminal 30) via the network 40.

[0025] Camera 15 uses an image sensor such as a Charged-Coupled Device (CCD), Metal-oxide-semiconductor (MOS), or Complementary Metal-Oxide-Semiconductor (CMOS) to capture objects in the field of view. Camera 15 is a two-dimensional camera. Camera 15 may be built into the first terminal 10 or attached externally to the first terminal 10. Camera 15 captures the field of view where the answer sheet 2 is located in response to the learner's operation. Camera 15 outputs the image 50 obtained from the capture.

[0026] The input device 16 accepts input from the learner. The input device 16 includes, for example, at least one of a touchpad, mouse, and keyboard. The display 17 is, for example, a liquid crystal display. The input device 16 and the display 17 may be configured as an integrated touch panel.

[0027] <Example of a rectangular object> Figure 2 shows an example of an image showing an answer sheet. In image 50 shown in Figure 2, the brightness of the target area where the answer sheet 2 is visible is about the same as the brightness of the background. As described above, if the technology described in Patent Document 1 is applied to image 50, where the change in brightness at the boundary between the target area and the background area is small, the target area may not be detected with high accuracy.

[0028] To address these issues, the answer sheet 2 includes L-shaped marks 23 placed at the corners 21, as shown in Figure 2. The L-shaped marks 23 are pre-printed on the answer sheet 2. The L-shaped marks 23 are positioned so that the outer edge 25 of the bent portion 24 of the L-shaped marks 23 faces the corner 21. Typically, the L-shaped marks 23 are placed at each of the four corners of the answer sheet 2. However, taking into account the margins within the answer sheet 2, the L-shaped marks 23 may be placed at only three of the four corners of the answer sheet 2.

[0029] Furthermore, the width of the lines constituting the L-shaped mark 23 is preferably 1 mm or more. Also, the length of the lines constituting the L-shaped mark 23 (length from the bent portion 24 to the end) is preferably 5 mm or more.

[0030] <Functional Configuration of Terminal 1> Figure 3 is a diagram showing an example of the functional configuration of the first terminal according to the embodiment. As shown in Figure 3, the first terminal 10 has a functional configuration comprising an acquisition unit 100, a first detection unit 110, a correction unit 120, and an output unit 130. The acquisition unit 100, the first detection unit 110, the correction unit 120, and the output unit 130 are realized by the processor 11 shown in Figure 1 executing application 18.

[0031] The acquisition unit 100 acquires an image 50 showing the answer sheet 2. Specifically, the acquisition unit 100 displays a notification on the display 17 prompting the user to take a picture of the answer sheet 2 and activates the camera 15. In response to the operation on the input device 16, the acquisition unit 100 outputs a shooting command to the camera 15. As a result, the acquisition unit 100 acquires an image 50 showing the answer sheet 2 from the camera 15.

[0032] The first detection unit 110 detects the target area in the image 50 in which the answer sheet 2 is visible, based on the shape corresponding to the L-shaped mark 23 in the image 50 output from the camera 15. The first detection unit 110 may perform preprocessing on the image 50 to make the target area easier to detect. Preprocessing may include, for example, brightness conversion processing.

[0033] As shown in Figure 3, the first detection unit 110 includes an extraction unit 111, a determination unit 112, and a identification unit 113.

[0034] The extraction unit 111 extracts multiple hexagonal shapes 53 (see Figure 2) from the image 50. For each of the multiple shapes 53, the extraction unit 111 identifies the XY coordinates of six nodes (vertices).

[0035] The extraction unit 111 can extract multiple shapes 53 using a known image analysis algorithm. Such image analysis algorithms are provided, for example, by the image processing library OpenCV [https: / / opencv.org / ]. That is, the application 18 shown in Figure 1 may include the image processing library OpenCV.

[0036] The OpenCV image processing library includes the following four commands: The first command extracts clusters of pixels whose brightness difference or color difference is within a specified range. The second command extracts the contour of each cluster of pixels. The third command approximates the contour of each cluster of pixels with a polygon. The fourth command determines an image cluster with six nodes as shape 53.

[0037] As shown in Figure 2, the multiple figures 53 extracted by the extraction unit 111 may include not only objects that are truly hexagonal (including the L-shaped mark 23), but also objects of various shapes. This is due to factors such as the resolution of the image 50, the shooting conditions (camera shake), and the elevation angle of the camera 15 relative to the answer sheet 2. For example, letters and hemispherical objects may be extracted as hexagonal figures 53.

[0038] As described above, the width of the lines constituting the L-shaped mark 23 is preferably 1 mm or more. Furthermore, the length of the lines constituting the L-shaped mark 23 is preferably 5 mm or more. As a result, when the image 50 has a resolution of 12 million pixels, the number of pixels in the image 50 that capture the lines constituting the L-shaped mark 23 will be approximately 13.5 pixels for A4 size and approximately 9.5 pixels for A3 size relative to the width of the lines, and approximately 67.5 pixels for A4 size and approximately 47.5 pixels for A3 size relative to the length of the lines. Consequently, the situation in which the L-shaped mark 23 is not extracted as a figure 53 is suppressed. In other words, if the number of pixels that capture the lines constituting the L-shaped mark 23 is 9.5 or more along the width direction of the lines and 47.5 or more along the longitudinal direction of the lines, the situation in which the L-shaped mark 23 is not extracted as a figure 53 is suppressed.

[0039] The determination unit 112 determines whether the characteristic conditions defining the features of the L-shaped mark 23 are met for each of the multiple figures 53.

[0040] The identification unit 113 identifies the target area on which the answer sheet 2 is printed, based on the position of the figure among the multiple figures 53 that satisfies the characteristic conditions. Specifically, the identification unit 113 identifies the coordinates of the four vertices of the target area.

[0041] The correction unit 120 corrects the image 50. Normally, learners can photograph the answer sheet 2 by operating the first terminal 10 with the optical axis of the camera 15 tilted relative to the normal direction of the answer sheet 2. Therefore, in the image 50, the area in which the answer sheet 2 is captured may have a distorted shape from a rectangle. To eliminate such distortion of the area, the correction unit 120 performs a projection transformation on the image 50 so that the area becomes a rectangle. Specifically, the correction unit 120 calculates a transformation matrix to transform the coordinates of the four vertices of the area identified by the identification unit 113 into the coordinates of the four vertices of a rectangle having the same aspect ratio as the answer sheet 2. The correction unit 120 transforms the image 50 using the transformation matrix.

[0042] The output unit 130 outputs the projected image 50. Specifically, the output unit 130 saves the projected image 50 to the storage 13. Furthermore, the output unit 130 controls the communication interface 14 to send the projected image 50 to the server 20 in response to an operation on the input device 16.

[0043] <Processing flow of the first terminal> The processing flow at the first terminal 10 will be explained with reference to Figures 4 to 7. Figure 4 is a diagram showing an example of the processing flow at the first terminal according to the embodiment.

[0044] First, in step S1, the processor 11, which operates as the acquisition unit 100, acquires an image 50 showing the answer sheet 2.

[0045] In the next step S2, the processor 11, which operates as an extraction unit 111, extracts multiple hexagonal shapes 53 from the image 50.

[0046] Figure 5 shows an example of a hexagonal shape. Figure 5 shows the shape 53 corresponding to the L-shaped mark 23.

[0047] The hexagonal figure 53 contains nodes N(0) to N(5) arranged sequentially along its outline. Hereafter, unless otherwise distinguished, nodes N(0) to N(5) will be referred to simply as "node N," omitting their subscripts (0 to 5). In the example shown in Figure 5, the processor 11 assigns the subscripts (0 to 5) to node N counterclockwise along the outline of figure 53. The processor 11 may also assign the subscripts (0 to 5) to node N clockwise along the outline of figure 53. For each figure 53, the processor 11 identifies the XY coordinates of each node N(0) to N(5).

[0048] In the next step S3, the processor 11, which operates as a determination unit 112, determines whether the feature conditions that define the features of the L-shaped mark 23 are met for each of the multiple figures 53.

[0049] The feature conditions include a size condition relating to size, a first angle condition relating to the angle between two edges flanking node N, and a second angle condition relating to the angle between two edges flanking an edge. The first angle condition corresponds to the “first condition” of this disclosure. The second angle condition corresponds to the “second condition” of this disclosure. The size condition corresponds to the “third condition” of this disclosure.

[0050] The size condition is that the size of the figure 53 is within a reference range set according to the size of the L-shaped mark 23. The processor 11 calculates the area of ​​the figure 53 as the size of the figure 53 based on a known image analysis algorithm (for example, the image processing library OpenCV [https: / / opencv.org / ]). The reference range is determined according to the size of the L-shaped mark 23 and the resolution of the image 50. Specifically, the lower limit of the reference range represents the number of pixels that capture the L-shaped mark 23 on the answer sheet 2 placed in 50% of the camera 15's field of view. The upper limit of the reference range represents the number of pixels that capture the L-shaped mark 23 on the answer sheet 2 placed in 100% of the camera 15's field of view. For example, the reference range is 30 to 500 pixels.

[0051] The first angle condition is that the difference between each of the following angles θ(0) to θ(5) and 90 degrees or 270 degrees is less than or equal to the threshold α. Angle θ(0): The angle between the vector V(0) pointing from node N(0) to node N(1) and the vector V(1) pointing from node N(1) to node N(2). Angle θ(1): The angle between vector V(1) and vector V(2) pointing from node N(2) to node N(3). Angle θ(2): The angle between vector V(2) and vector V(3) pointing from node N(3) to node N(4). Angle θ(3): The angle between vector V(3) and vector V(4) pointing from node N(4) to node N(5). Angle θ(4): The angle between vector V(4) and vector V(5) pointing from node N(5) to node N(0). Angle θ(5): The angle between vector V(5) and vector V(0).

[0052] The first angle condition is based on the fact that each interior angle of the L-shaped mark 23 is 90 degrees or 270 degrees. The threshold α is predetermined according to the range of the assumed elevation angle of the camera 15 relative to the answer sheet 2. That is, the threshold α is predetermined according to the magnitude of the assumed distortion of the target area in which the answer sheet 2 is captured in the image 50. The threshold α is, for example, 30°. The threshold α is an example of the “first threshold” of this disclosure.

[0053] Processor 11 calculates vectors V(0) to V(5) based on the XY coordinates of nodes N(0) to N(5) (see Figure 5). Processor 11 calculates angles θ(0) to θ(5) according to the following formula using the dot product of two vectors that start or end at a certain node N. Hereafter, unless otherwise distinguished, angles θ(0) to θ(5) will each be referred to as "angle θ". θ(0)=arccos(V(0)·V(1) / |V(0)| / |V(1)|) θ(1)=arccos(V(1)·V(2) / |V(1)| / |V(2)|) θ(2)=arccos(V(2)·V(3) / |V(2)| / |V(3)|) θ(3)=arccos(V(3)·V(4) / |V(3)| / |V(4)|) θ(4)=arccos(V(4)·V(5) / |V(4)| / |V(5)|) θ(5)=arccos(V(5)·V(0) / |V(5)| / |V(0)|) The processor 11 determines that the first angle condition is met if the following equation is satisfied for each of the angles θ(0) to θ(5). 90-α≦θ≦90+α or 270-α≦θ≦270+α

[0054] Alternatively, processor 11 may calculate cos(θ(0)) to cos(θ(5)) according to the following formula. cos(θ(0))=V(0)·V(1) / |V(0)| / |V(1)| cos(θ(1))=V(1)·V(2) / |V(1)| / |V(2)| cos(θ(2))=V(2)·V(3) / |V(2)| / |V(3)| cos(θ(3))=V(3)·V(4) / |V(3)| / |V(4)| cos(θ(4))=V(4)·V(5) / |V(4)| / |V(5)| cos(θ(5))=V(5)·V(0) / |V(5)| / |V(0)| Then, the processor 11 determines that the first angle condition is satisfied if the following equation is satisfied for each of cos(θ(0)) to cos(θ(5)). -α1≦cos(θ)≦α1 The threshold α1 is determined according to the threshold α. For example, if the threshold α is 30°, then the threshold α1 is 1 / 2.

[0055] The second angle condition is that there exists a k such that the difference between each of the angles γ(k) and γ(mod(k+1,6)) from the following angles γ(0) to γ(5) and 0 degrees is less than or equal to the threshold β, and the difference between each of the remaining angles from γ(0) to γ(5) and 180 degrees is less than or equal to the threshold β. mod(k+1,6) represents the remainder when k+1 is divided by 6. Angle γ(0): The angle between vector V(0) and vector V(2). Angle γ(1): The angle between vector V(1) and vector V(3), Angle γ(2): The angle between vector V(2) and vector V(4), Angle γ(3): The angle between vector V(3) and vector V(5). Angle γ(4): The angle between vector V(4) and vector V(0). Angle γ(5): The angle between vector V(5) and vector V(1).

[0056] The second angle condition is based on the fact that two sides enclosing one of the two sides constituting the 270-degree interior angle of the L-shaped mark 23 are parallel, and that when a virtual point is moved in one direction along the contour, the virtual point moves in the same direction on those two sides. The threshold β is predetermined according to the range of the assumed elevation angle of the camera 15 relative to the answer sheet 2. That is, the threshold β is predetermined according to the assumed magnitude of distortion in the target area in the image 50 in which the answer sheet 2 is captured. The threshold β is, for example, 30°. The threshold β is an example of the “second threshold” of this disclosure.

[0057] The processor 11 calculates angles γ(0) to γ(5) according to the following equation, which uses the dot product of two vectors that enclose an edge between two adjacent nodes N. γ(0)=arccos(V(0)·V(2) / |V(0)| / |V(2)|) γ(1)=arccos(V(1)·V(3) / |V(1)| / |V(3)|) γ(2)=arccos(V(2)·V(4) / |V(2)| / |V(4)|) γ(3)=arccos(V(3)·V(5) / |V(3)| / |V(5)|) γ(4)=arccos(V(4)·V(0) / |V(4)| / |V(0)|) γ(5)=arccos(V(5)·V(1) / |V(5)| / |V(1)|)

[0058] In the example shown in Figure 5, vectors V(2) and V(4) are approximately parallel and point in the same direction. Also, vectors V(3) and (5) are approximately parallel and point in the same direction. In contrast, the combinations of vectors V(0) and V(2), V(1) and V(3), V(4) and V(0), and V(5) and V(1) are approximately parallel, but point in opposite directions. Therefore, the processor 11 determines that the above subcondition is satisfied at k=2, and thus determines that the second angle condition is met.

[0059] The processor 11 determines that the figure 53 that satisfies all of the size, first angle, and second angle conditions is the figure 53 that satisfies the feature conditions.

[0060] Returning to Figure 4, the processor 11 performs steps S4 to S11 to identify the target region based on the positions of the figures that satisfy the feature conditions. In step S4, the processor 11 determines whether four of the multiple figures 53 satisfy the feature conditions.

[0061] If the four figures 53 satisfy the feature condition (YES in step S4), the process moves to step S5. In step S5, the processor 11 identifies a node of interest corresponding to the outer edge 25 of the bent portion 24 in the L-shaped mark 23 (see Figure 2) for each of the four figures 53 that satisfy the feature condition.

[0062] As shown in Figure 5, in the figure 53 corresponding to the L-shaped mark 23, node N(mod(i+3,6)), three nodes after node N(i) with an interior angle of 270 degrees, corresponds to the outer edge 25 of the bent portion 24 in the L-shaped mark 23. As described above, the second angle condition is based on the fact that the two sides enclosing one of the two sides constituting the interior angle of 270 degrees in the L-shaped mark 23 are parallel, and when a virtual point is moved in one direction along the contour, the virtual point moves in the same direction on those two sides. Therefore, the relationship between k, which satisfies the subcondition of the second angle condition, and the subscript i of node N(i) with an interior angle of 270 degrees is as follows. i = k + 2 Therefore, the processor 11 determines node N(mod(k+5,6)) as the node of interest using k that satisfies the subcondition of the second angle condition. In the example shown in Figure 5, the above subcondition is satisfied when k=2. Therefore, the processor 11 determines node N(1) as the node of interest.

[0063] Returning to Figure 4, in step S6 following step S5, the processor 11 determines the four nodes of interest of the figure 53 that satisfy the feature conditions as the vertices of the target region on which the answer sheet 2 is reflected.

[0064] If none of the four figures 53 satisfy the characteristic conditions (NO in step S4), the process moves to step S7. In step S7, the processor 11 determines whether two or three of the figures 53 satisfy the characteristic conditions.

[0065] If two or three figures 53 satisfy the feature condition (YES in step S7), the process moves to step S8. For example, if L-shaped marks 23 are placed on only two or three of the four corners of answer sheet 2, the processor 11 may determine that two or three figures 53 satisfy the feature condition. Alternatively, if L-shaped marks 23 are placed on four or three corners of answer sheet 2, but shadows are cast near some of the L-shaped marks 23, the processor 11 may determine that two or three figures 53 satisfy the feature condition.

[0066] In step S8, the processor 11 identifies a node of interest for each of the two or three figures 53 that satisfy the feature condition, corresponding to the outer edge 25 of the bend 24 in the L-shaped mark 23. The method for identifying the node of interest is as described in step S5. The nodes of interest for the two or three figures 53 are considered to correspond to the L-shaped marks 23 placed in two or three corners of the answer sheet 2.

[0067] In step S9, following step S8, the processor 11 determines whether it is possible to identify the completion points corresponding to the remaining corners of the answer sheet 2. The processor 11 determines that if the three figures 53 satisfy the feature conditions, it is possible to identify the completion point corresponding to the remaining corner of the answer sheet 2. The processor 11 determines that if two figures 53 satisfy the feature conditions, it is possible to identify two completion points corresponding to the remaining two corners, in accordance with the fact that these two figures 53 correspond to L-shaped marks 23 placed on two non-adjacent corners of the answer sheet 2. The two non-adjacent corners are either a pair of the top-left and bottom-right corners, or a pair of the top-right and bottom-left corners.

[0068] As described above, the L-shaped mark 23 is positioned so that the outer edge 25 of the bent portion 24 faces the corner of the answer sheet 2. Therefore, as shown in Figure 2, the orientations of the L-shaped marks 23 placed at the four corners of the answer sheet 2 are different from each other. Thus, the processor 11 only needs to determine, based on the orientation of two figures 53 that satisfy the characteristic conditions, which of the two figures 53 corresponds to the L-shaped mark 23 placed at the top left, top right, bottom right, or bottom left corner of the answer sheet 2.

[0069] Specifically, for each of the two figures 53 that satisfy the feature conditions, the processor 11 identifies a node of interest, a first node adjacent to the node of interest clockwise along the outline of the figure 53, and a second node adjacent to the node of interest counterclockwise along the outline of the figure 53. For example, if node N(1) is identified as the node of interest, node N(0) is identified as the first node and node N(2) is identified as the second node. The processor 11 identifies that a figure 53 that satisfies the following placement condition (a) corresponds to the L-shaped mark 23 placed in the upper left corner of the answer sheet 2. The processor 11 identifies that a figure 53 that satisfies the following placement condition (b) corresponds to the L-shaped mark 23 placed in the upper right corner of the answer sheet 2. The processor 11 identifies that a figure 53 that satisfies the following placement condition (c) corresponds to the L-shaped mark 23 placed in the lower right corner of the answer sheet 2. The processor 11 identifies that the figure 53 that satisfies the following placement condition (d) corresponds to the L-shaped mark 23 attached to the lower left corner of the answer sheet 2. Placement condition (a): The X coordinate of the first node is greater than or equal to a reference value than the X coordinate of the node of interest, and the Y coordinate of the second node is greater than or equal to a reference value than the Y coordinate of the node of interest. Placement condition (b): The Y coordinate of the first node is greater than or equal to the Y coordinate of the node of interest by a certain threshold, and the X coordinate of the second node is less than or equal to the X coordinate of the node of interest by a certain threshold. Placement condition (c): The X coordinate of the first node is less than or equal to a certain threshold value than the X coordinate of the node of interest, and the Y coordinate of the second node is less than or equal to a certain threshold value than the Y coordinate of the node of interest. Placement condition (d): The Y coordinate of the first node is less than or equal to the Y coordinate of the node of interest by a certain threshold, and the X coordinate of the second node is greater than or equal to the X coordinate of the node of interest by a certain threshold. The reference value is set, for example, to half the length of the lines that make up the L-shaped mark 23. Note that placement conditions (a) to (d) assume that the origin (0,0) of the coordinate system of image 50 is the upper left corner. Therefore, placement conditions (a) to (d) can be appropriately changed depending on the coordinate system of image 50.

[0070] In this way, the processor 11 can determine which corners of the answer sheet 2—the upper left, upper right, lower right, and lower left—correspond to each of the two figures 53 that satisfy the characteristic conditions, and then determine whether it is possible to identify the completion points corresponding to the remaining corners.

[0071] If it is possible to identify the completion points (YES in step S9), the process moves to step S10. In step S10, the processor 11 identifies the completion points corresponding to the remaining corners of the answer sheet 2.

[0072] The method for identifying the completion points will be explained with reference to Figures 6 and 7. Figure 6 is a diagram illustrating the preprocessing for identifying the completion points. As shown in Figure 6, for each of the two or three figures 53 that satisfy the feature conditions, the processor 11 identifies a first extension line 56 of one of the two sides 54, 55 that enclose the node of interest TN (side 54), and a second extension line 57 of the other side (side 55). The first extension line 56 is parallel to side 54 and is a half-line with its endpoint being node TN1 on the opposite side of side 54 from the node of interest TN. The second extension line 57 is parallel to side 55 and is a half-line with its endpoint being node TN2 on the opposite side of side 55 from the node of interest TN.

[0073] The processor 11 identifies the intersection points of two lines selected from the first extension line 56 and the second extension line 57 of two or three figures 53 that satisfy the characteristic conditions as complementary points corresponding to the remaining corners of the answer sheet 2.

[0074] Figure 7 illustrates how to identify a completion point. Figure 7 shows an example of identifying one completion point from three figures 53. The processor 11 identifies one completion point from three figures 53 that satisfy the feature condition. For each of the multiple combinations of selecting two lines from the first extension line 56 and the second extension line 57 of the three figures 53 that satisfy the feature condition, the processor 11 identifies the intersection point of two lines. In the example shown in Figure 7, figures 53(1), 53(2), and 53(3) satisfy the feature condition. Therefore, the processor 11 identifies the intersection point of two lines for each of the multiple combinations of selecting two lines from the first extension line 56(1) to 56(3) and the second extension line 57(1) to 57(3) of each of the figures 53(1) to 53(3). In the example shown in Figure 7, the processor 11 identifies the intersection point P1 of the first extension line 56(3) and the second extension line 57(1), and the intersection point P2 of the first extension line 56(2) and the second extension line 57(3) within the image 50.

[0075] The processor 11 identifies the interpolation point corresponding to the L-shaped mark 23 from among multiple intersection points corresponding to multiple combinations. Specifically, the processor 11 identifies the intersection point located within image 50 and with the longest shortest distance to the three target nodes TN(1) to TN(3) of figures 53(1) to 53(3) as the interpolation point. In the example shown in Figure 7, both intersection points P1 and P2 are located within image 50. Intersection point P1 is closest to the target node TN(1) among the target nodes TN(1), TN(2), and TN(3). Intersection point P2 is closest to the target node TN(3) among the target nodes TN(1), TN(2), and TN(3). The distance between intersection point P1 and target node TN(1) is longer than the distance between intersection point P2 and target node TN(3). Therefore, the processor 11 identifies intersection point P1 as the interpolation point 58.

[0076] Alternatively, the processor 11 refers to the above placement conditions (a) to (d) to determine which of the top left, top right, bottom right, and bottom left corners of the answer sheet 2 corresponds to each of the three figures 53 that satisfy the feature conditions. Based on the determination, the processor 11 determines which of the top left, top right, bottom right, and bottom left corners of the answer sheet 2 should be used to identify the completion point. Specifically, the processor 11 determines that the completion point should be identified in a corner where there are no L-shaped marks 23 corresponding to the three figures 53 that satisfy the feature conditions. If the corner where the completion point should be identified is the top left corner of the answer sheet 2, the processor 11 determines the intersection of the first extension line 56 of the figure 53 corresponding to the L-shaped mark 23 placed in the bottom left corner of the answer sheet 2 and the second extension line 57 of the figure 53 corresponding to the L-shaped mark 23 placed in the top right corner of the answer sheet 2 as the completion point. If the corner where the completion point should be identified is the upper right corner of the answer sheet 2, the processor 11 determines the intersection of the first extension line 56 of the figure 53 corresponding to the L-shaped mark 23 placed in the upper left corner of the answer sheet 2 and the second extension line 57 of the figure 53 corresponding to the L-shaped mark 23 placed in the lower right corner of the answer sheet 2 as the completion point. If the corner where the completion point should be identified is the lower right corner of the answer sheet 2, the processor 11 determines the intersection of the first extension line 56 of the figure 53 corresponding to the L-shaped mark 23 placed in the upper right corner of the answer sheet 2 and the second extension line 57 of the figure 53 corresponding to the L-shaped mark 23 placed in the lower left corner of the answer sheet 2 as the completion point. If the corner where the completion point should be identified is the lower left corner of the answer sheet 2, the processor 11 determines the intersection of the first extension line 56 of the figure 53 corresponding to the L-shaped mark 23 placed in the lower right corner of the answer sheet 2 and the second extension line 57 of the figure 53 corresponding to the L-shaped mark 23 placed in the upper left corner of the answer sheet 2 as the completion point.

[0077] In the example shown in Figure 7, figures 53(1), 53(2), and 53(3) correspond to the L-shaped marks 23 placed in the upper right, lower right, and lower left corners of the answer sheet 2, respectively. Therefore, the processor 11 decides that it should identify a completion point for the upper left corner of the answer sheet 2. Accordingly, the processor 11 determines that the intersection point P1 of the first extension line 56(3) of figure 53(3) and the second extension line 57(1) of figure 53(1) is the completion point 58.

[0078] Alternatively, the processor 11 divides the image 50 equally into four 2x2 subregions 59(1) to 59(4). The processor 11 identifies the subregion from among the four subregions 59(1) to 59(4) in which the target node TN corresponding to the three figures 53 that satisfy the feature conditions does not exist. In the example shown in Figure 7, the target nodes TN(1), TN(2), and TN(3) do not exist in subregion 59(1). The processor 11 identifies the intersection P1 located in the identified subregion 59(1) as the interpolation point 58.

[0079] The processor 11 identifies two completion points from two figures 53 that satisfy the characteristic conditions. Specifically, the processor 11 determines one completion point 58 as the intersection of a first extension line 56 identified for one of the two figures 53 and a second extension line 57 identified for the other figure 53. Furthermore, the processor 11 determines another completion point 58 as the intersection of a second extension line 57 identified for one of the two figures 53 and a first extension line 56 identified for the other figure 53.

[0080] Returning to Figure 4, in step S11 following step S10, the processor 11 determines the notion node TN and completion point 58 of two or three figures 53 that satisfy the feature conditions as vertices of the target region on which the answer sheet 2 is projected. If three figures 53 satisfy the feature conditions, the processor 11 determines the notion node TN and one completion point 58 of those three figures 53 as vertices of the target region on which the answer sheet 2 is projected. If two figures 53 satisfy the feature conditions, the processor 11 determines the notion node TN and two completion points 58 of those two figures 53 as vertices of the target region on which the answer sheet 2 is projected.

[0081] After step S5 or step S11, the process moves to step S12. In step S12, the processor 11, which operates as the correction unit 120, performs a projection transformation on the image 50 so that the target area becomes a rectangle.

[0082] In step S13, following step S12, the processor 11 saves the projected image 50 to the storage 13. Furthermore, the processor 11 controls the communication interface 14 to send the projected image 50 to the server 20. After step S13, the processor 11 terminates processing.

[0083] If two or three figures 53 do not satisfy the characteristic conditions (NO in step S7), or if it is impossible to identify the completion points (NO in step S9), the processor 11 displays a message on the display 17 prompting the user to take another picture of the answer sheet 2, and then terminates the process.

[0084] <Example 1> Figure 8 shows the functional configuration of the first terminal according to Modification 1. The first terminal 10A according to Modification 1 has the same hardware configuration as the first terminal 10 shown in Figure 1. However, as shown in Figure 8, the first terminal 10A differs from the first terminal 10 in that it includes a second detection unit 140 and a selection unit 150 as functional components. The second detection unit 140 and the selection unit 150 are realized by the processor 11 executing application 18.

[0085] The second detection unit 140 uses an edge extraction algorithm to detect the target area in the image 50 where the answer sheet 2 is visible. Specifically, the second detection unit 140 uses an edge extraction algorithm to extract edges caused by the difference in brightness between the target area and the background area. The second detection unit 140 then detects the target area based on the extracted edges.

[0086] The second detection unit 140 may employ known techniques as edge extraction algorithms (for example, the techniques disclosed in Patent Document 1). Typically, the second detection unit 140 extracts edges from the image 50 using the Canny method and detects the region enclosed by the group of edges that approximates a rectangle as the target region.

[0087] The selection unit 150 selects the block from the first detection unit 110 and the second detection unit 140 that should be enabled.

[0088] For example, the selection unit 150 selects a block to be enabled based on the brightness distribution of the image 50. In the image 50, the greater the brightness difference between the target area where the answer sheet 2 is visible and the background area, the more accurately the boundary between the target area and the background area is extracted as an edge. Furthermore, the greater the brightness difference between the target area and the background area, the greater the variance of the brightness distribution in the image 50. Therefore, the selection unit 150 enables the second detection unit 140 when the variance of the brightness distribution in the image 50 exceeds a predetermined threshold. This allows the second detection unit 140 to accurately detect the target area based on the edges extracted from the image 50. On the other hand, the selection unit 150 enables the first detection unit 110 when the variance of the brightness distribution in the image 50 is below a predetermined threshold. This allows the first detection unit 110 to accurately detect the target area from the image 50, where the brightness difference between the target area and the background area is small, based on the figure 53 corresponding to the L-shaped mark 23.

[0089] Alternatively, the selection unit 150 may enable the second detection unit 140 by default, and enable the first detection unit 110 if the detection of the target area by the second detection unit 140 fails. The selection unit 150 determines, for example, whether the detection of the target area by the second detection unit 140 was successful, based on the size of the target area detected by the second detection unit 140. Specifically, the selection unit 150 determines that the detection of the target area by the second detection unit 140 was successful if the size of the target area detected by the second detection unit 140 exceeds a reference value (for example, 45% of the total size of the image 50). The selection unit 150 determines that the detection of the target area by the second detection unit 140 failed if the target area is not detected by the second detection unit 140, or if the size of the target area detected by the second detection unit 140 is less than or equal to the reference value.

[0090] Figure 9 shows an example of the processing flow of the first terminal according to Modification 1. The flow shown in Figure 9 includes the same steps S1, S12, and S13 as the flow shown in Figure 4. Therefore, a detailed explanation of these steps will be omitted below.

[0091] As shown in Figure 9, after step S1, the process moves to step S21. In step S21, the processor 11, which operates as the selection unit 150, selects a detection method based on the brightness distribution of the image 50. Specifically, the processor 11 selects a first detection method if the variance of the brightness distribution of the image 50 is below a predetermined threshold. The processor 11 selects a second detection method if the variance of the brightness distribution of the image 50 exceeds a predetermined threshold.

[0092] If the second detection method is selected, in step S22, the processor 11, which operates as the second detection unit 140, extracts edges from the image 50 using an edge extraction algorithm and detects the target area in which the answer sheet 2 is visible based on the extracted edges.

[0093] On the other hand, if the first detection method is selected, in step S23, the processor 11, which operates as the first detection unit 110, detects the target area based on the figure 53 corresponding to the L-shaped mark 23. That is, the processor 11 executes steps S2 to S11 shown in Figure 4.

[0094] After step S22 or step S23, the processor 11 performs steps S12 and S13.

[0095] Figure 10 shows another example of the processing flow of the first terminal according to Modification 1. The flow shown in Figure 10 includes the same steps S1, S12, and S13 as the flow shown in Figure 4. Therefore, a detailed explanation of these steps will be omitted below.

[0096] As shown in Figure 10, after step S1, the process moves to step S22. In step S22, the processor 11, which operates as the selection unit 150, selects the second detection method by default. Therefore, the processor 11, which operates as the second detection unit 140, extracts edges from the image 50 using an edge extraction algorithm and detects the target area in which the answer sheet 2 is visible based on the extracted edges.

[0097] In step S31, following step S22, the processor 11, which operates as a selection unit 150, determines whether the detection of the target area in step S22 was successful or not.

[0098] If detection of the target area fails (NO in step S31), in step S23, the processor 11 selects the first detection method. Therefore, the processor 11, operating as the first detection unit 110, detects the target area based on the figure 53 corresponding to the L-shaped mark 23. That is, the processor 11 executes steps S2 to S11 shown in Figure 4.

[0099] If the detection of the target area is successful (YES in step S31), or after step S23, the processor 11 performs steps S12 and S13.

[0100] According to the first terminal 10A of the modified example 1, the processor 11 can appropriately select a method for detecting the target area in which the answer sheet 2 is visible, depending on the state of the image 50. As a result, the target area is detected with high accuracy.

[0101] <Modification 2> The second detection unit 140 may detect the target area using a saliency map obtained from the image 50. The saliency map is generated using known saliency detection techniques and represents the degree to which a person is paying attention to an area. The second detection unit 140 only needs to detect a block of rectangular pixels with a relatively high degree of attention in the saliency map as the target area.

[0102] Alternatively, the second detection unit 140 may detect the target region using an edge extraction algorithm and a saliency map. For example, the second detection unit 140 detects a region as the target region that contains a pixel cluster of relatively high importance and is surrounded by a group of edges that approximate a rectangle.

[0103] <Variation 3> In the above explanation, the feature conditions were assumed to include size conditions, first angle conditions, and second angle conditions. However, the size condition may be excluded from the feature conditions.

[0104] Alternatively, the second angle condition may be excluded from the feature conditions. Alternatively, the size condition and the second angle condition may be excluded from the feature conditions. In these cases, the processor 11 calculates cos(θ(0))~cos(θ(5)) according to the above formula, and also calculates sin(θ(0))~sin(θ(5)) according to the following formula using the cross product of two vectors that start or end at a certain node N. sin(θ(0))=V(0)×V(1) / |V(0)| / |V(1)| sin(θ(1))=V(1)×V(2) / |V(1)| / |V(2)| sin(θ(2))=V(2)×V(3) / |V(2)| / |V(3)| sin(θ(3))=V(3)×V(4) / |V(3)| / |V(4)| sin(θ(4))=V(4)×V(5) / |V(4)| / |V(5)| sin(θ(5))=V(5)×V(0) / |V(5)| / |V(0)| Then, based on cos(θ(0))~cos(θ(5)) and sin(θ(0))~cos(θ(5)), the processor 11 identifies angles θ(i) from θ(0) to θ(5) such that the difference from 270 degrees is less than or equal to the threshold α. Specifically, the processor 11 identifies angles θ(i) that satisfy the following equation. -α1≦cos(θ)≦α1 -1 ≤ sin(θ) ≤ -α² The threshold α2 is determined according to the threshold α. For example, if the threshold α is 30°, then the threshold α2 is 3 1 / 2 It is / 2.

[0105] Then, the processor 11 can determine node N(mod(i+3,6)) from among nodes N(0) to N(5) as the node TN of interest, which corresponds to the outer edge 25 of the bent portion 24 in the L-shaped mark 23. Note that mod(i+3,6) represents the remainder when i+3 is divided by 6.

[0106] <Modification 4> In the above description, the first terminal 10 is assumed to include a first detection unit 110 and a correction unit 120. However, the first detection unit 110 and the correction unit 120 may be provided in the server 20 or the second terminal 30. In this case, the output unit 130 of the first terminal 10 controls the communication interface 14 to transmit the image 50 obtained by the camera 15 to the server 20. The first detection unit 110 provided in the server 20 or the second terminal 30 then detects the target area from the image 50 output from the first terminal 10 based on the figure 53 corresponding to the L-shaped mark 23. The server 20 or the second terminal 30 equipped with the first detection unit 110 corresponds to the “detection device” in this disclosure.

[0107] <Modification 5> The processor 11, operating as the first detection unit 110, may determine that the detection of the target area has failed if the size of the target area detected based on the figure 53 corresponding to the L-shaped mark 23 is less than or equal to a reference value (for example, 40.5% of the total size of the image 50). In this case, the processor 11 displays a message on the display 17 prompting the user to take another picture of the answer sheet 2.

[0108] While embodiments of the present invention have been described, the embodiments disclosed herein should be considered in all respects to be illustrative and not restrictive. The scope of the present invention is defined by the claims, and all modifications within the meaning and scope equivalent to the claims are intended to be included. [Explanation of Symbols]

[0109] 1 Educational system, 2 Answer sheet, 10, 10A First terminal, 11 Processor, 12 Memory, 13 Storage, 14 Communication interface, 15 Camera, 16 Input device, 17 Display, 18 Application, 20 Server, 21 Corner, 23 L-shaped mark, 24 Bent part, 25 Outer edge, 30 Second terminal, 40 Network, 50 Image, 53 Figure, 54, 55 Sides, 56 First extension line, 57 Second extension line, 58 Complementary point, 100 Acquisition unit, 110 First detection unit, 111 Extraction unit, 112 Judgment unit, 113 Identification unit, 120 Correction unit, 130 Output unit, 140 Second detection unit, 150 Selection unit, P1, P2 Intersection point, TN Node of interest.

Claims

1. A computer-based detection method for detecting a region occupied by a rectangular object within an image, wherein the object includes an L-shaped mark attached to its corner, and the detection method is: From the aforementioned image, extract multiple hexagonal shapes, For each of the aforementioned multiple figures, it is determined whether or not the characteristic conditions defining the features of the L-shaped mark are met, The process includes identifying the target region based on the position of the figure among the plurality of figures that satisfies the characteristic conditions, Each of the aforementioned plurality of figures includes nodes N(0) to N(5) arranged in order along the outline of the hexagon, The aforementioned determination is, The angle θ(0) between the vector V(0) pointing from node N(0) to node N(1) and the vector V(1) pointing from node N(1) to node N(2), The angle θ(1) between the aforementioned vector V(1) and the vector V(2) pointing from node N(2) to node N(3), The angle θ(2) between the aforementioned vector V(2) and the vector V(3) pointing from node N(3) to node N(4), The angle θ(3) between the aforementioned vector V(3) and the vector V(4) pointing from node N(4) to node N(5), The angle θ(4) between the vector V(4) and the vector V(5) pointing from node N(5) to node N(0), This includes calculating the angle θ(5) between the vector V(5) and the vector V(0), The detection method wherein the characteristic condition includes a first condition that the difference between each of the angles θ(0) to θ(5) and 90 degrees or 270 degrees is less than or equal to a first threshold.

2. The aforementioned determination is, The angle γ(0) between the aforementioned vector V(0) and the aforementioned vector V(2), The angle γ(1) between the aforementioned vector V(1) and the aforementioned vector V(3), The angle γ(2) between the aforementioned vector V(2) and the aforementioned vector V(4), The angle γ(3) between the aforementioned vector V(3) and the aforementioned vector V(5), The angle γ(4) between the vector V(4) and the vector V(0), This includes calculating the angle γ(5) between the vector V(5) and the vector V(1), The detection method according to claim 1, wherein the characteristic condition includes a second condition that there exists a k that satisfies the sub-condition that the difference between each of the angles γ(k) and γ(mod(k+1,6)) among the angles γ(0) to γ(5) and 0 degrees is less than or equal to a second threshold, and the difference between each of the remaining angles among the angles γ(0) to γ(5) and 180 degrees is less than or equal to the second threshold, and mod(k+1,6) represents the remainder when k+1 is divided by 6.

3. The L-shaped mark is positioned such that the outer edge of the bent portion of the L-shaped mark faces the corner. The aforementioned determination means that Identifying k that satisfies the aforementioned sub-conditions, The detection method according to claim 2, comprising determining node N(mod(k+5,6)) among the nodes N(0) to N(5) as a node of interest corresponding to the outer edge of the bend in the L-shaped mark, wherein mod(k+5,6) represents the remainder when k+5 is divided by 6.

4. The L-shaped mark is positioned such that the outer edge of the bent portion of the L-shaped mark faces the corner. The aforementioned determination means that To identify the angle θ(i) among the angles θ(0) to θ(5) such that the difference from 270 degrees is less than or equal to the first threshold, The detection method according to claim 1, comprising determining node N(mod(i+3,6)) among the nodes N(0) to N(5) as a node of interest corresponding to the outer edge of the bend in the L-shaped mark, wherein mod(i+3,6) represents the remainder when i+3 is divided by 6.

5. The detection method according to any one of claims 1 to 4, wherein the characteristic condition further includes a third condition that the size of the target figure is within a reference range set according to the size of the L-shaped mark.

6. The detection method according to claim 3 or 4, wherein, in response to the determination that four of the plurality of figures satisfy the characteristic conditions, the identification includes determining the nodes of interest of the four figures as vertices of the target region.

7. In accordance with the determination that two or three of the aforementioned multiple figures satisfy the characteristic conditions, the identification is as follows: For each of the two or three figures, identify the first extension of one of the two sides enclosing the node of interest, and the second extension of the other of the two sides. Identifying the intersection of two lines selected from the first and second extensions of the two or three figures as complementary points corresponding to the corners of the object, The detection method according to claim 3 or 4, comprising determining the nodes of interest and the interpolation points of the two or three figures as vertices of the target region.

8. A detection device for detecting a target area occupied by a rectangular object within an image, wherein the object includes an L-shaped mark attached to its corner, and the detection device An extraction unit that extracts multiple hexagonal shapes from the aforementioned image, A determination unit that determines whether or not the characteristic conditions defining the characteristics of the L-shaped mark are met for each of the aforementioned plurality of figures, The system includes a specification unit that identifies the target region based on the position of the figure among the plurality of figures that satisfies the characteristic conditions, Each of the aforementioned plurality of figures includes nodes N(0) to N(5) arranged in order along the outline of the hexagon, The unit that makes the determination said, The angle θ(0) between the vector V(0) pointing from node N(0) to node N(1) and the vector V(1) pointing from node N(1) to node N(2), The angle θ(1) between the aforementioned vector V(1) and the vector V(2) pointing from node N(2) to node N(3), The angle θ(2) between the aforementioned vector V(2) and the vector V(3) pointing from node N(3) to node N(4), The angle θ(3) between the aforementioned vector V(3) and the vector V(4) pointing from node N(4) to node N(5), The angle θ(4) between the vector V(4) and the vector V(5) pointing from node N(5) to node N(0), The angle θ(5) between the aforementioned vector V(5) and the aforementioned vector V(0) is calculated, The detection device, wherein the characteristic conditions include a first condition that the difference between each of the angles θ(0) to θ(5) and 90 degrees or 270 degrees is less than or equal to a first threshold.

9. A program for causing a computer to execute the detection method described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Data input device, data input system, display data analyszer and medium

    JP2001224557A

  • Specific image position estimating device, specific image position estimating method, specific image position estimating program, recording medium computer-readable with specific image position estimating program stored, and medium

    JP2005293409A

  • Image processing method, image region detecting method, image processing program, image region detection program, image processing apparatus, and image region detecting apparatus

    JP2008283649A

  • Reader, image forming apparatus, position detection method, and program

    JP2019103122A

  • Image processing device, image processing method, program and recording medium

    JP2024121891A