Image processing methods, methods for generating training data
The image processing method addresses the issue of inconsistent training data by applying affine transformations to incorporate background features, enhancing machine learning accuracy in object recognition.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2026-04-02
AI Technical Summary
Existing methods for generating training data for machine learning do not adequately incorporate image information from regions excluding the target object, leading to inconsistencies in brightness distribution and proportion of images, which affects learning accuracy.
An image processing method that applies affine transformations to generate new images by moving pixel groups from one region to another, dividing regions to reflect features beyond the target object, and using these transformed images as training data.
Facilitates the generation of training data that includes features from the background and context, enhancing the learning process by reflecting luminance distribution and other image information, thereby improving object recognition accuracy.
Smart Images

Figure 0007839726000001 
Figure 0007839726000002 
Figure 0007839726000003
Abstract
Description
Technical Field
[0006] ,
[0001] The present disclosure relates to an image processing method and a method for generating teaching data.
Background Art
[0002] Machine learning using teaching data to recognize a desired object in an image is known. Increasing the teaching data using data augmentation is one method of improving the learning accuracy of machine learning. In data augmentation, for example, adjustment of the brightness of an image, rotation of the image, and enlargement of the image are performed to obtain new teaching data. <000001b><000001c><000001d>For example, Patent Document 1 below discloses a technique of generating a trimming image obtained by cutting out a predetermined region including an object from an image of the object, and pasting the trimmed image on a background image after rotation. <000001e>
Prior Art Documents
Patent Documents
Patent Document 1
[0005]
Summary of the Invention
Problems to be Solved by the Invention
[0006]
[0007] Patent Document 1 does not disclose how to set the background image. For example, when generating the second teaching data from the first teaching data, it is unclear whether the image information of the region excluding the trimming image from the image of the object is reflected in the second teaching data. In this case, the brightness distribution, for example, is different between the second teaching data and the first teaching data. For example, if the brightness value of the background image is fixed, it is assumed that not only the brightness distribution but also the proportion of images in the second training data will differ from the proportion of images in the first training data. Regarding the data augmentation method for obtaining the second training data, there is room for improvement, for example, with regard to image information in areas excluding cropped images.
[0007] This disclosure was made in view of the above-mentioned issues and aims to contribute to the generation of training data using image features other than those of the target object. [Means for solving the problem]
[0008] A first aspect of the image processing method according to this disclosure comprises a transformation step of applying a first affine transformation to a first image having a plurality of pixels and occupying a first region to obtain a second image having the plurality of pixels and occupying a second region that is congruent to and inconsistent with the first region, and moving the plurality of pixels having the second image to the first region to obtain a third image.
[0009] The second region is divided into a third region and a fourth region, with the third region located within the first region and the fourth region located outside the first region.
[0010] The first region is divided into the third region and the fifth region, and the fifth region is located outside the second region.
[0011] The movement step includes a first step of setting a first source pixel group, which is a group of pixels in which a predetermined first number greater than 1 of the plurality of pixels included in the fourth region are linked together, and a second step of applying a second affine transformation to the first source pixel group to determine a first destination pixel group to be arranged in the fifth region.
[0012] The first affine transformation is either a rotation or a translation, or both. The second affine transformation is either a rotation or a translation, or both, or a reflection transformation.
[0013] A second aspect of the image processing method according to this disclosure is the first aspect thereof, wherein the first source pixel group is inscribed in the fourth region, and the first destination pixel group is inscribed in the fifth region.
[0014] A third aspect of the image processing method according to the present disclosure is the first or second aspect thereof, wherein the movement step includes a third step of setting a second source pixel group which is a group of pixels in which a predetermined second number of the plurality of pixels included in the fourth region are linked together, and a fourth step of applying a third affine transform to the second source pixel group to determine a second destination pixel group to be arranged in the fifth region. The second number is less than or equal to the first number. The third affine transform is either rotation or translation, or both, or a reflection transform.
[0015] A fourth aspect of the image processing method according to this disclosure is a third aspect thereof, wherein the second number is 1. The third and fourth steps are repeatedly performed until all of the plurality of pixels in the second image are moved to or duplicated in the first region.
[0016] A fifth aspect of the image processing method according to this disclosure is a third aspect thereof, wherein the first and second steps are performed multiple times prior to the third step.
[0017] For example, a segmented figure that is expanded by connecting pixels within the fourth region, starting from a position away from the third region, is set as the first group of source pixels.
[0018] For example, the first group of destination pixels is placed in a segmented region that is expanded by connecting pixels within the fifth region, starting from a position away from the third region.
[0019] For example, the fourth region and the fifth region are mirror images of each other, and the second affine transformation is a mirror transformation.
[0020] For example, the fourth region and the fifth region are in a mirror image relationship, and the third affine transformation is a mirror image transformation.
[0021] The method for generating teacher data according to the present disclosure is a method for generating teacher data to be used for machine learning for recognizing an object in an image. This method is the image processing method according to the present disclosure Noi A step of adopting the plurality of third images obtained by any of the misalignments as the teacher data, a step of imaging the object prior to the image processing method to set the first image, and when obtaining each of the third images, from the pixel group located in the third region in both the first image and the second image, a step of extracting the pixel group corresponding to the object in either the first image or the second image.
Advantages of the Invention
[0022] According to the first aspect of the image processing method according to the present disclosure, teacher data for an object annotated within a third region can be easily obtained while incorporating features that do not solely depend on the luminance distribution of images other than the object.
[0023] According to the second, third, and fifth aspects of the image processing method according to the present disclosure, teacher data in which features of images other than the object are more reflected can be obtained.
[0024] According to the fourth aspect of the image processing method according to the present disclosure, among the features of images other than the object, information regarding the luminance distribution is likely to be reflected in the teacher data.
[0025] For example, the teacher data obtained by the method for generating teacher data according to the present disclosure is used for machine learning for recognizing the object.
Brief Description of the Drawings
[0026] [Figure 1] It is a diagram illustrating a first image. [Figure 2] It is a diagram showing the relationship between a second image and a first region. [Figure 3] This is a diagram illustrating an example of the third image. [Figure 4] This figure illustrates another example of the third image. [Figure 5] This flowchart illustrates the process of creating training images. [Figure 6] This flowchart illustrates the process of moving the first pixel group. [Figure 7] This is a flowchart illustrating the details of the splitting process. [Figure 8] This diagram illustrates the first, second, third, and fourth domains. [Figure 9] This diagram illustrates the first, second, third, and fourth domains. [Figure 10] This diagram illustrates the first, second, third, and fourth domains. [Figure 11] This diagram illustrates the first, second, third, and fourth domains. [Figure 12] This diagram illustrates the first, second, third, and fourth domains. [Figure 13] This diagram illustrates the first, second, third, and fourth domains. [Figure 14] This diagram illustrates the first, second, third, and fourth domains. [Figure 15] This diagram illustrates the first, second, third, and fourth domains. [Figure 16] This diagram illustrates the first, second, third, and fourth domains. [Figure 17] This diagram illustrates the first, second, third, and fourth domains. [Figure 18] This flowchart illustrates the process of moving the second pixel group. [Figure 19] This flowchart shows one other example of the content of the first pixel group movement process. [Figure 20] This flowchart illustrates the details of the partitioning process performed in step S35c. [Figure 21]This flowchart illustrates the contents of the first placement process. [Figure 22] This flowchart illustrates the details of the partitioning process performed in step S305. [Figure 23] This is a conceptual diagram illustrating the partitioning process performed in step S35c. [Figure 24] This is a conceptual diagram illustrating the partitioning process performed in step S35c. [Figure 25] This is a conceptual diagram illustrating the partitioning process performed in step S35c. [Figure 26] This is a conceptual diagram illustrating the partitioning process performed in step S35c. [Figure 27] This is a conceptual diagram illustrating the partitioning process performed in step S35c. [Figure 28] This is a conceptual diagram illustrating the partitioning process performed in step S35c. [Figure 29] This is a conceptual diagram illustrating the partitioning process performed in step S35c. [Figure 30] This is a conceptual diagram illustrating the partitioning process performed in step S305. [Figure 31] This is a conceptual diagram illustrating the partitioning process performed in step S305. [Figure 32] This is a conceptual diagram illustrating the partitioning process performed in step S305. [Figure 33] This is a conceptual diagram illustrating the partitioning process performed in step S305. [Figure 34] This is a conceptual diagram illustrating the partitioning process performed in step S305. [Figure 35] This is a conceptual diagram illustrating the partitioning process performed in step S305. [Figure 36] This flowchart shows a second example of the content of the first pixel group movement process. [Figure 37] This flowchart illustrates the details of the partitioning process performed in step S35d. [Figure 38] This is a block diagram illustrating the generation and use of training images. [Figure 39] This is a diagram illustrating the display image. [Modes for carrying out the invention]
[0027] The embodiments of this disclosure will be described below with reference to the attached drawings. The components described in each embodiment are illustrative only, and the scope of this disclosure is not limited to illustrative purposes. The drawings are for illustrative purposes only. In the drawings, the dimensions and number of parts may be exaggerated or simplified as needed to facilitate understanding. In the drawings, parts with similar configurations and functions are denoted by the same reference numerals, and redundant explanations are omitted where appropriate.
[0028] In this specification, expressions indicating relative or absolute positional relationships (e.g., "rotation"), unless otherwise specified, describe not only the exact positional relationship but also a state including tolerances, as well as a state that is relatively displaced with respect to angle or distance within a range that yields a similar degree of function. Expressions indicating that two or more things are equivalent (e.g., "congruent"), unless otherwise specified, describe not only a state that is quantitatively exactly equivalent but also a state in which tolerances or differences that yield a similar degree of function exist.
[0029] Unless otherwise specified, expressions describing shape (e.g., "square") shall not only represent the geometrically precise shape but also the range within which a similar effect can be achieved.
[0030] Expressions such as "to possess," "to be equipped with," "to have," "to include," or "to have" a single component are not exclusive expressions that exclude the existence of other components.
[0031] The term "connected" generally refers to a state in which two elements are in contact with each other, unless otherwise specified.
[0032] <1. Overview> Figure 1 illustrates a first image F1, which is an image obtained by imaging a predetermined object and has multiple pixels. For example, the predetermined object is a chuck pin. Multiple chuck pins are provided on a base that holds, for example, a semiconductor wafer (not shown), and support the periphery of the semiconductor wafer. In Figure 1, the chuck pin is schematically shown as a collection of three cylinders and appears as a group of pixels 100 that constitute a part of the first image F1.
[0033] The first image F1 occupies the first region R1. For example, the first region R1 has an outline Q. In the first image F1, pixels are arranged in a matrix along two non-parallel directions. Here, these two directions are orthogonal, and the outline Q is approximately rectangular. For example, if each pixel in the first image F1 is square in shape, the outline Q is rectangular. For example, if each pixel in the first image F1 is circular in shape, the outline Q is rectangular with semicircles along each side, and this shape is included in the "approximately rectangular" described above.
[0034] To avoid complexity in the illustration, in Figure 1, the code indicating the first region R1 and the code indicating the first image F1 are collectively denoted as "F1(R1)". This notation does not imply that the first region R1 and the first image F1 are identical.
[0035] Figure 2 shows the relationship between the second image F2 and the first region R1. To avoid complexity in the illustration, in Figure 2, the code representing the first region R1 and the code representing the outline Q are collectively denoted as "Q(R1)". This notation does not imply identity between the first region R1 and the outline Q. The second image F2 is obtained by applying the first affine transformation to the first image F1 and is an image with multiple pixels. The second image F2 occupies the second region R2. To avoid complexity in the illustration, in Figure 2, the code representing the second region R2 and the code representing the second image F2 are collectively denoted as "F2(R2)". This notation does not imply identity between the second region R2 and the second image F2.
[0036] The first affine transformation is either a rotation or a translation, or both. The second region R2 is congruent to and inconsistent with the first region R1. If the angle of rotation in the first affine transformation is 2nπ radians (where n is an integer), it is accompanied by a translation at a non-zero distance. If the translation distance in the first affine transformation is zero, it is accompanied by a rotation at an angle other than 2nπ radians (where n is an integer).
[0037] The second region R2 is divided into the third region R3 and the fourth region R4. The fourth region R4 is a collective term for one to four fourth regions R41, R42, R43, and R44. The third region R3 is located within the first region R1, while the fourth region R4 is located outside the first region R1. The third region R3 is an overlapping region between the first region R1 and the second region R2, but the pixels in the first image F1 and the pixels in the second image F2 do not necessarily overlap within the third region R3.
[0038] The first region R1 is divided into the third region R3 and the fifth region R5. The fifth region R5 consists of one to four regions, in this case the four regions R51, R52, R53, and R54. The fifth region R5 is located outside the second region R2.
[0039] Pixel group 100 is located in the third region R3 in both the first image F1 and the second image F2. Pixel group 100 is annotated by region extraction on either the first image F1 or the second image F2.
[0040] Machine learning is envisioned to recognize an object (a chuck pin in the above example) in an image. This machine learning is provided with images as training data (hereinafter referred to as "training images"). Image F1, the first image, can be used as one of the training images.
[0041] Figure 3 illustrates the third image 3A. The third image 3A represents the process of moving or duplicating multiple pixels from the second image F2 to the first region R1 (this process is also tentatively referred to simply as the "movement process"). toIt can also be said that the third image 3A is obtained by data augmenting the first image F1. To avoid the complexity of the illustration, in Figure 3, the code indicating the first region R1 and the code indicating the outer boundary Q are collectively denoted as "Q(R1)". This notation does not imply that the first region R1 and the outer boundary Q are identical.
[0042] The movement process includes, for example, a first step and a second step. The first step sets a source pixel group Gk (where k is a positive integer) contained in the fourth region R4. The source pixel group Gk is a group of pixels of number Mk (where k is the same as k in the source pixel group Gk: Mk is an integer greater than 1) that are concatenated. For example, the number Mk is the square of an integer greater than or equal to 2 (i.e., Mk ≥ 4), and the pixels in the source pixel group Gk are arranged in a square shape. In this disclosure, "arranged in a square shape" means arranged in a matrix along two mutually orthogonal directions.
[0043] The second step involves applying a second affine transformation to the source pixel group Gk to obtain the destination pixel group Hk (where k is the same as k in the source pixel group Gk). The second affine transformation is one of several affine transformations: rotation, translation, or both, or a reflection transformation. If the angle of rotation in the second affine transformation is 2nπ radians (where n is an integer), it is accompanied by a translation at a non-zero distance. If the translation distance in the second affine transformation is zero, it is accompanied by a rotation at an angle other than 2nπ radians (where n is an integer).
[0044] The following example illustrates the case where the sum of the rotation angle in the first affine transformation and the rotation angle in the second affine transformation is mπ / 2 (where m is an integer) radians. In this case, and when the pixels in the source pixel group Gk are arranged in a square shape, the pixels in the destination pixel group Hk are also arranged in a square shape.
[0045] The destination pixel group Hk is located within the fifth region R5. In the example above, the destination pixel group Hk is located inside one of the fifth regions R51, R52, R53, or R54.
[0046] Images 31a, 32a, 33a, and 34a occupy the fifth region R51, R52, R53, and R54, respectively. In Figure 3, the code indicating the fifth region R51 and the code indicating image 31a are collectively denoted as "31a(R51)". This notation does not imply identity between the fifth region R51 and image 31a. In Figure 3, the code indicating the fifth region R52 and the code indicating image 32a are collectively denoted as "32a(R52)". This notation does not imply identity between the fifth region R52 and image 32a. In Figure 3, the code indicating the fifth region R53 and the code indicating image 33a are collectively denoted as "33a(R53)". This notation does not imply identity between the fifth region R53 and image 33a. In Figure 3, the code indicating region 5R54 and the code indicating image 34a are collectively denoted as "34a(R54)". This notation does not imply that region 5R54 and image 34a are identical.
[0047] Image 30a occupies the third region R3 and coincides with the area between the second image F2 and the third image F3. In Figure 3, the code indicating the third region R3 and the code indicating image 30a are collectively denoted as "30a(R3)". This notation does not imply identity between the third region R3 and image 30a.
[0048] The information contained in the source pixel group Gk is reflected in the destination pixel group Hk. Both the first and second affine transformations are either rotation or translation, or both, and the number Mk is greater than 1. Not only the brightness and color tone of each pixel constituting the source pixel group Gk, but also the positional relationship between those pixels is reflected in the destination pixel group Hk. The third image 3A differs from the first image F1, but largely reflects the information contained in the first image F1. The third image 3A is suitable as a training image to be used in machine learning for recognizing objects in images. For example, the third image 3A is used as a training image together with the first image F1.
[0049] <2. Example of arrangement of destination pixel group Hk> For example, the outline of the source pixel group Gk and the outline of the destination pixel group Hk are congruent. Below, with reference to Figure 3, an example is given where the pixels in the source pixel group Gk are arranged in a square shape, the pixels in the destination pixel group Hk are also arranged in a square shape, and the integer k is between 1 and 4 (inclusive).
[0050] Images 31a, 32a, 33a, and 34a each have pixel set sets 310a, 320a, 330a, and 340a, respectively. Pixel set 310a is a set of pixel sets 311a, 312a, 313a, 314a, 315a, 316a, and 317a, which are examples of any of the destination pixel sets H1, H2, H3, and H4. Pixel set 320a is a set of pixel sets 321a, 322a, 323a, and 324a, which are examples of any of the destination pixel sets H1, H2, H3, and H4. Pixel set 330a is a set of pixel sets 331a, 332a, 333a, and 334a, which are examples of any of the destination pixel sets H1, H2, H3, and H4. The pixel group set 340a is the set of pixel groups 341a, 342a, 343a, 344a, 345a, and 346a, which are examples of any of the destination pixel groups H1, H2, H3, and H4.
[0051] As the number Mk is updated and the pair of the first and second steps is repeatedly executed, a destination pixel group Hk of different sizes is obtained. Referring to Figure 3, the number of pixels in pixel groups 311a and 331a is the same, number M1, which is more than the number of pixels in each of the other pixel groups 312a, 313a, 314a, 315a, 316a, 317a, 321a, 322a, 323a, 324a, 332a, 333a, 334a, 341a, 342a, 343a, 344a, 345a, and 346a. As the pair of the first and second steps is executed, pixel group 311a is placed in the fifth region R51, and pixel group 331a is placed in the fifth region R53.
[0052] The pixel groups 312a, 321a, and 332a each have the same number of pixels, which is M2. M2 is smaller than M1, and greater than the number of pixels in each of the pixel groups 313a, 314a, 315a, 316a, 317a, 322a, 323a, 324a, 333a, 334a, 341a, 342a, 343a, 344a, 345a, and 346a. Pixel group 312a is located in the fifth region R51, pixel group 321a is located in the fifth region R52, and pixel group 332a is located in the fifth region R53.
[0053] The pixel groups 313a, 314a, 322a, 333a, 341a, and 342a each have the same number of pixels, which is M3. M3 is smaller than M2, but greater than the number of pixels in each of the pixel groups 315a, 316a, 317a, 323a, 324a, 334a, 343a, 344a, 345a, and 346a. Pixel groups 313a and 314a are located in the fifth region R51, pixel group 322a is located in the fifth region R52, pixel group 333a is located in the fifth region R53, and pixel groups 341a and 342a are located in the fifth region R54.
[0054] Each of the pixel groups 315a, 316a, 317a, 323a, 324a, 334a, 343a, 344a, 345a, and 346a has the same number of pixels, which is M4 and less than M3. Pixel groups 315a, 316a, and 317a are located in the fifth region R51, pixel groups 323a and 324a are located in the fifth region R52, pixel group 334a is located in the fifth region R53, and pixel groups 343a, 344a, 345a, and 346a are located in the fifth region R54.
[0055] In this way, pairs of the first and second steps are repeatedly executed to obtain pixel sets 310a, 320a, 330a, and 340a. For example, the minimum value of the number Mk is set when pairs of the first and second steps are repeatedly executed. Referring to Figure 3, the number M4 is set to the minimum value of the number Mk.
[0056] The larger the number Mk, the more the third image 3A reflects the information contained in the first image F1, but the smaller the number of destination pixel groups Hk obtained. Since Mk ≥ 2, it is not necessarily the case that the entire fifth region R5 is occupied by destination pixel groups Hk. Referring to Figure 3, images 31a, 32a, 33a, and 34a are examples of cases where there are pixels other than the pixel groups 310a, 320a, 330a, and 340a, respectively.
[0057] In the fifth region R5, the areas not occupied by the destination pixel groups H1, H2, H3, and H4 are arranged by applying a second affine transformation to each of the pixels in the fourth region R4 other than the source pixel groups G1, G2, G3, and G4. With this arrangement, the fifth region R5 is completely occupied by the pixels that were located in the fourth region R4.
[0058] For pixels in image 31a other than pixel set 310a, pixels in image 32a other than pixel set 320a, pixels in image 33a other than pixel set 330a, and pixels in image 34a other than pixel set 340a, pixels obtained by applying a second affine transformation to each of the pixels in the fourth region R41, R42, R43, and R44 other than the source pixel sets G1, G2, G3, and G4 are selected.
[0059] For example, in image 31a, pixels other than the pixel group set 310a are 、 In the fourth region R41, for each of the pixels other than the source pixel group G1, G2, G3, and G4 do Pixels to which a second affine transformation has been applied are used. For example, in image 32a, pixels other than the pixel group set 320a are used. 、 In the fourth region R42, for each of the pixels other than the source pixel group G1, G2, G3, and G4 do Pixels to which a second affine transformation has been applied are used. For example, in image 33a, pixels other than the pixel group set 330a are used. 、 In the fourth region R43, for each of the pixels other than the source pixel group G1, G2, G3, and G4 do Pixels to which a second affine transformation has been applied are used. For example, in image 34a, pixels other than the pixel group set 340a are used. 、In the fourth region R44, for each of the pixels other than the source pixel group G1, G2, G3, and G4 do Pixels that have undergone a second affine transformation are used.
[0060] The number Mk may be a fixed value of 2 or more. For example, referring to Figure 3, the number M1 is set to the minimum value of the number Mk, the pixel group set 310a has only the pixel group 311a which is an example of the destination pixel group H1, the pixel group set 330a has only the pixel group 331a which is an example of the destination pixel group H1, and the pixel group sets 320a and 340a are empty sets that do not have the destination pixel group H1. In this case, the pixels in image 31a other than the pixel group 311a, the pixels in image 32a, the pixels in image 33a other than the pixel group 331a, and the pixels in image 34a each have a corresponding pixel in the fourth region R41, R42, R43, R44 other than the source pixel group G1. do Pixels that have undergone a second affine transformation are used.
[0061] Repeatedly executing the first and second processes while updating the number Mk without fixing it contributes to the third image F3 largely reflecting the information contained in the first image F1.
[0062] It is also possible to place pixels obtained by applying a second affine transformation to each of the pixels in the fourth region R4 into the fifth region R5. In other words, this is the case when the source pixel group Gk has only one pixel and Mk is fixed at 1. In this case, compared to the case where the number Mk is 2 or more, it is undesirable from the standpoint that the positional relationships between the pixels constituting the source pixel group Gk are not easily reflected in the destination pixel group Hk.
[0063] Figure 4 illustrates the third image 3B. The third image 3B is obtained when the first and second affine transformations differ in rotation angle and displacement from those illustrated in Figures 1, 2, and 3.
[0064] Images 31b, 32b, 33b, and 34b occupy the fifth region R51, R52, R53, and R54, respectively. 4In this diagram, the code indicating the fifth region R51 and the code indicating image 31b are collectively denoted as "31b(R51)". This notation does not imply that the fifth region R51 and image 31b are identical. 4 In this diagram, the code indicating the fifth region R52 and the code indicating image 32b are collectively denoted as "32b(R52)". This notation does not imply that the fifth region R52 and image 32b are identical. 4 In this context, the code indicating region 5R53 and the code indicating image 33b are collectively denoted as "33b(R53)". This notation does not imply identity between region 5R53 and image 33b. 4 In this context, the code indicating region 5R54 and the code indicating image 34b are collectively denoted as "34b(R54)". This notation does not imply identity between region 5R54 and image 34b.
[0065] Image 30b occupies the third region R3 and coincides with the area between the second image F2 and the third image F3. Figure 4 In this context, the code indicating the third region R3 and the code indicating image 30b are collectively denoted as "30b(R3)". This notation does not imply identity between the third region R3 and image 30b.
[0066] Images 31b, 32b, 33b, and 34b each have pixel set sets 310b, 320b, 330b, and 340b, respectively. Pixel set 310b is a set of pixel groups 311b, 312b, 313b, 314b, 315b, and 316b, which are examples of any of the destination pixel group Hk. Pixel set 320b is a set of pixel groups 321b and 322b, which are examples of any of the destination pixel group Hk. Pixel set 330b is a set of pixel groups 331b, 332b, 333b, and 334b, which are examples of any of the destination pixel group Hk. ,335b This is a set. Pixel group set 340b is an example of any of the destination pixel group Hk, pixel groups 341b, 342 b It is a set.
[0067] As the number Mk is updated and the first and second processes are repeated, pixel group 310b is placed in the fifth region R51, pixel group 320b in the fifth region R52, pixel group 330b in the fifth region R53, and pixel group 340b in the fifth region R54.
[0068] For pixels in image 31b other than pixel group set 310b, pixels in image 32b other than pixel group set 320b, pixels in image 33b other than pixel group set 330b, and pixels in image 34b other than pixel group set 340b, each of the pixels in the fourth region R41, R42, R43, and R44 other than the source pixel group Gk is do Pixels that have undergone a second affine transformation are used.
[0069] For example, in image 31b, pixels other than the pixel set 310b are 、 In the fourth region R41, for each of the pixels other than the source pixel group Gk do Pixels to which a second affine transformation has been applied are used. For example, in image 32b, pixels other than the pixel group set 320b are used. 、 In the fourth region R42, for each of the pixels other than the source pixel group Gk do Pixels to which a second affine transformation has been applied are used. For example, in image 33b, pixels other than the pixel group set 330b are used. 、 In the fourth region R43, for each of the pixels other than the source pixel group Gk do Pixels to which a second affine transformation has been applied are used. For example, in image 34b, pixels other than the pixel group set 340b are used. 、 In the fourth region R44, for each of the pixels other than the source pixel group Gk do Pixels that have undergone a second affine transformation are used.
[0070] For example, the third images 3A and 3B are selected as training images. In this case, the first image F1 may or may not be selected as a training image.
[0071] Figure 5 is a flowchart illustrating the process of creating training images. In Figure 5, this process is labeled "Creation of training images." This process consists of steps S1, S2, S3, S4, and S5, which are executed in this order.
[0072] Step S1 performs the setting of the first image F1. Specifically, the image obtained by imaging an object, such as a chuck pin, is set as the first image F1. For example, the chuck pin appears as pixel group 100 in the third region R3 of the first image F1.
[0073] Step S2 performs the generation of the second image F2. Specifically, step S2 generates the second image F2 from the first image F1 by the first affine transformation (see Figure 2).
[0074] Step S3 executes the first pixel group movement process. The first pixel group movement process involves setting the source pixel group Gk and obtaining the destination pixel group Hk by a second affine transformation. Referring to Figure 3, the execution of step S3 yields the pixel group sets 310a, 320a, 330a, and 340a. Referring to Figure 4, the execution of step S3 yields the pixel group sets 310b, 320b, 330b, and 340b.
[0075] Step S4 executes the second pixel group movement process. In the second pixel group movement process, the pixels are placed in the fifth region R5, in the area not occupied by the destination pixel group Hk in the first pixel group movement process, by applying a second affine transformation to each of the pixels in the fourth region R4 other than the source pixel group Gk. In the example described above, after the execution of step S4, the third images 3A and 3B are obtained.
[0076] Step S5 involves saving the third image. The saved third image can be used as a training image.
[0077] Annotation to obtain pixel group 100 is performed in either step S1 or S2 when obtaining each of the third images. However, this annotation extracts the pixel group corresponding to the object from the pixel group located in the third region R3 in both the first image F1 and the second image F2. Since the pixels in the first image F1 and the pixels in the second image F2 do not necessarily overlap in the third region R3, region extraction is performed from the overlapping pixel groups.
[0078] For example, when multiple first affine transformations are pre-set, for each third image of The third region R3 is predetermined. From the images of the object (described later as captured image F0), the image in which the object is captured at a position corresponding to the determined third region R3 is adopted as the first image F1.
[0079] <3. Example of setting the source pixel group Gk> Figure 6 is a flowchart illustrating the contents of the first pixel group movement process performed in step S3. Step S3 includes steps S31, S32, S33, S34, S35a, S35b, S36, S37, S38, S39, and S30.
[0080] Figure 7 shows steps S35a, S3 5 This is a flowchart illustrating the contents of the division process performed in step b. Figures 8, 9, and 10 sequentially illustrate the situation in which step S35a is performed. Figures 11 to 17 sequentially illustrate the situation in which step S35b is performed. As will be described in detail later, both steps S35a and S35b can be said to be steps that obtain a set of partitioned shapes that are candidates for the source pixel group Gk from the fourth region R4.
[0081] After the second image F2 is generated in step S2 (see Figure 5), step S31 is executed. Step S31 recognizes the third region R3, the fourth region R4, and the fifth region R5. This recognition is performed by comparing the first image F1 with the second image F2. For example, the first region R1 is recognized as the region occupied by the first image F1, the outline Q is extracted from the first region R1, and the third region R3 and the fourth region R4 are recognized by comparing the outline Q with the second region R2. The fifth region R5 is recognized by comparing the first region R1 with the third region R3. In Figure 6, step S31 is abbreviated as "Recognition of the third to fifth regions".
[0082] After step S31 is executed, step S32 is executed. Step S32 specifies one of the fourth regions R4. In Figure 2, the fourth regions R41, R42, R43, and R44 are shown as examples of the fourth region R4. Step S32 selects and specifies, for example, multiple of these regions without overlap. The fourth region R4 specified by step S32 may be referred to as "fourth region R4z" below.
[0083] After steps S33, S35a, or steps S33, S34, S35b, S36, S37, S38 described later are performed, step S39 determines whether all of the fourth region R4 have been specified. If the result of the determination in step S39 is negative, step S32 is performed again, and any of the fourth region R4 that have not yet been specified are selected and specified. This can be seen as an update of the fourth region R4z.
[0084] In step S33, the number of convex angles in the fourth region R4z whose vertices do not lie on the outer boundary Q (hereinafter and in the drawings, simply referred to as "convex angles") is determined. The fourth regions R41, R42, R43, and R44 exemplified in Figure 2 each have one convex angle (single). In the following, Figures 8 to 17 will be used to explain the cases of one and multiple convex angles, and therefore, the first region R1, second region R2, third region R3, and fourth region R4 will be exemplified in a different configuration than that shown in Figures 1 to 3. For the sake of simplicity, the following examples will show that the pixel shape is square, and both the first region R1 and the second region R2 are rectangular.
[0085] In Figures 8 to 17, the first region R1 is divided into the third region R3 and the fifth region R55 and R56, and the second region R2 is divided into the third region R3 and the fourth region R45 and R46.
[0086] The first region R1 is rectangular as described above, and this rectangle has vertices P51, P52, P53, and P54 located in this order in a counterclockwise direction. This rectangle has an edge L512 connecting points P51 and P52, an edge L523 connecting points P52 and P53, an edge L534 connecting points P53 and P54, and an edge L541 connecting points P54 and P51.
[0087] The second region R2 is rectangular as described above, and this rectangle has vertices P41, P42, P43, and P44 located in this order in a counterclockwise direction. This rectangle has an edge L412 connecting points P41 and P42, an edge L423 connecting points P42 and P43, an edge L434 connecting points P43 and P44, and an edge L441 connecting points P44 and P41.
[0088] Edges L441 and L541 intersect at point P61, edges L412 and L523 intersect at point P62, edges L423 and L523 intersect at point P63, and edges L423 and L534 intersect at point P64.
[0089] The third region R3 is a hexagon enclosed by points P41, P62, P63, P64, P54, and P61. The fourth region 45 has only one convex angle, the one with point P42 as its vertex. The fourth region 46 has only two convex angles, the one with point P43 as its vertex and the one with point P44 as its vertex. The angle with point P54 as its vertex is a concave angle, and the angles with points P61 and P64 as their vertices are all on the outer boundary Q, so none of these angles qualify as "convex angles" in the sense described above.
[0090] If the fourth domain R45 is specified in step S32, the result of the determination in step S33 is "singular". If the fourth domain R46 is specified in step S32, the result of the determination in step S33 is "multiple".
[0091] If the result of the determination in step S33 is "singular", step S35a is executed. Step S35a is a partitioning process, which has steps S351, S352, S353, S354, S355, and S356 (see Figure 7). This partitioning process is used in either step S35a or S35b.
[0092] The specific details of step S35a will be explained using Figures 7 to 11, taking the case where the fourth region R45 is specified in step S32 as an example.
[0093] In step S35a, step S351 is executed. Step S351 specifies the vertex of the convex angle as the starting point for the process of obtaining the segmented figure. Referring to Figure 8, point P42 is the vertex and is the starting point.
[0094] After step S351 is executed, step S352 is executed. Step S352 enlarges the divided shape. For example, a square is adopted as the divided shape. For example, the initial value of the divided shape adopted in step S352 is a square with 4 pixels. The amount by which the divided shape is enlarged in step S352 is set to one pixel in each of the two directions parallel to each of the adjacent sides of the square as the divided shape.
[0095] Figure 8 illustrates the state in which the square is expanded starting from point P42 by step S352, resulting in the formation of the segmented figure T40.
[0096] After step S352 is executed, step S353 is executed. Step S353 determines whether the segmented figure has come into contact with the third region R3. If this determination is negative, step S354 is executed. In step S354, it is determined whether the segmented figure has come into contact with the vertices of other convex angles. If this determination is negative, step S352 is executed again, and the segmented figure is further expanded.
[0097] The segmented figure T40 shown in Figure 8 does not touch the edge that forms the boundary of the third region R3, so the result of the decision in step S353 is negative. The fourth region R45 does not have any other convex angles, so the result of the decision in step S354 is negative. Step S352 is performed again, and the segmented figure is expanded from segmented figure T40.
[0098] Figure 9 illustrates a situation where steps S353, S354, and S352 are repeatedly executed, resulting in the acquisition of a sectioned figure T41, which is enlarged starting from point P42. The sectioned figure T41 contacts edge L523 between points P62 and P63, and also contacts the third region R3. Therefore, if step S353 is executed first after the sectioned figure T41 is obtained, the judgment result of step S353 is positive.
[0099] If the result of step S353 is positive, step S354 is not executed, nor is step S352 executed, and the sectioned figure T41 is not further enlarged. This is because the sectioned figure obtained by further enlarging sectioned figure T41 would extend beyond the fourth region R45, which is undesirable.
[0100] The judgment in step S353 is negative, and the judgment in step S354 is positive Examples of situations where this is a target will be given later.
[0101] If the result of the judgment in step S353 is positive, or if the result of the judgment in step S354 is positive, then step S355 is executed. Step S355 determines whether the size of the obtained segmented figure is less than or equal to a predetermined value. This predetermined value corresponds to the minimum value of the number Mk.
[0102] The segmented shape is a candidate for the source pixel group Gk. If step S35a is performed, the fourth region R4z has one convex angle, and the segmented shape is adopted as the source pixel group Gk. The smaller this predetermined value is, the more source pixel groups Gk are obtained, which contributes to the third image F3 reflecting more of the features of the first image F1.
[0103] The decision made in step S354 determines whether step S352 is executed further, which contributes to increasing the size of the segmented figure and, consequently, the pixel group included in the source pixel group Gk. Due to this contribution, the third image F3 reflects more of the features of the first image F1.
[0104] The explanation continues with an example where the size of the divided figure T41 is not less than or equal to a predetermined value. In this case, the judgment in step S355 is negative, and step S356 is executed. Step S356 updates the starting point. For example, either point P412 or P423 is adopted as the updated starting point. Point P412 is located in the part where the outline of the divided figure T41 contacts edge L412, and is furthest from the original starting point P42. Point P423 is located in the part where the outline of the divided figure T41 contacts edge L423, and is furthest from the original starting point P42.
[0105] Figure 10 shows the situation in which the segmented shapes T42, T43, and T44 are obtained. The segmented shapes T42, T43, and T44 are obtained by, for example, the following process.
[0106] If the starting point is updated to point P412 in step S356, the segmented figure is enlarged by executing steps S352 and S353, and segmented figure T42 is obtained. Segmented figure T42 is an enlarged square starting from point P412, touches side L523 between points P62 and P63, and touches the third region R3.
[0107] If a segmented figure T42 is obtained, the judgment result in step S353 is positive, and the size of the segmented figure T42 is not less than or equal to the predetermined value in step S355, mosquito This scenario is anticipated. In this case, step S356 is executed again, and the starting point is updated to point P423.
[0108] The segmented figure starting from point P423 is enlarged by performing steps S352 and S353 to obtain segmented figure T43. Segmented figure T43 is a square that has been enlarged starting from point P423, and it is in contact with the edge L523 between points P62 and P63, and is in contact with the third region R3.
[0109] If a segmented figure T43 is obtained, the judgment result in step S353 is positive, and the size of the segmented figure T43 is not less than or equal to the predetermined value in step S355, mosquito This is a possible scenario. In this case, a divided figure T44 is obtained based on a divided figure T43 in the same manner as the process used to obtain a divided figure T43 based on a divided figure T41. The divided figure T44 is in contact with the edge L523 between points P62 and P63, and is in contact with the third region R3.
[0110] If the size of the divided figure T44 is less than or equal to the predetermined value specified in step S355, the process returns to step S39 without further determining the divided figure.
[0111] If the fourth region R45 is specified in step S32 without specifying the fourth region R46, and steps S35a and S39 are executed, the judgment in step S39 becomes negative and step S32 is executed again. At this time, the execution of step S32 specifies the fourth region R46. The judgment result of step S33, which is executed afterward, is "multiple", and step S39 is executed via steps S34, S35b, S36, S37, and S38.
[0112] Figure 1 1 Figure 16 illustrates the case where the fourth region R46 is specified in step S32. After steps S34, S35b, and S36 are executed, in step S37 it is determined whether all convex angles in the fourth region R4z specified in step S32 have been selected. If the result of the determination in step S37 is negative, step S34 is executed again, and any convex angles in the fourth region R4z that have not yet been specified are selected and specified. This can be seen as an update of the convex angles in the fourth region R4z.
[0113] Step S35b is a division process. The specific details of step S35b will be explained using Figures 7 and 11-13, taking as an example the case where a convex angle with point P44 as the vertex is specified in step S34. The specific details of step S35b will be explained using Figures 7 and 14-16, taking as an example the case where a convex angle with point P43 as the vertex is specified in step S34.
[0114] Referring to Figure 7, step S351 is executed in step S35b. Step S351 specifies the vertex of the convex corner as the starting point for the process of obtaining the segmented figure. Referring to Figure 11, point P44 is the vertex and is the starting point. Figure 11 illustrates the state in which the square is expanded starting from point P44 by step S352, and the segmented figure V41 is obtained.
[0115] For the segmented figure V41, the judgment result is negative in both steps S353 and S354, and step S352 is executed again.
[0116] Figure 12 illustrates a situation where steps S353, S354, and S352 are repeatedly executed, resulting in the acquisition of a sectioned figure V42 that is enlarged starting from point P44. The sectioned figure V42 contacts edge L541 between points P54 and P61, and contacts the third region R3. Therefore, if step S353 is executed first after the sectioned figure V42 is obtained, the judgment result of step S353 is positive.
[0117] Steps S355 and S356 are executed in the same manner as in step S35a, and points P441 and P443 become the new starting points. Point P441 is located in the part where the outline of the section figure V42 contacts edge L441, and is furthest from the original starting point P44. Point P443 is located in the part where the outline of the section figure V42 contacts edge L434, and is furthest from the original starting point P44.
[0118] Figure 13 shows the situation in which section figures V43 and V46 are obtained. Section figures V43 and V46 are obtained, for example, by the following process. If the starting point is updated to point P441 in step S356, the section figure is enlarged by executing steps S352, S353, and S354, and section figure V43 is obtained. If the starting point is updated to point P443 in step S356, the section figure is enlarged by executing steps S352, S353, and S354, and section figure V46 is obtained.
[0119] For the segmented figure V43, the judgment result in step S353 is positive, and step S355 is executed. In contrast, for the segmented figure V46, the judgment result in step S353 is negative, and the judgment in step S354 is positive, so step S355 is executed. This is because the segmented figure V46 does not contact point P54 but contacts point P43, and the segmented figure obtained by further expanding segmented figure V46 would extend beyond the fourth region R45, which is undesirable.
[0120] The generation of smaller divisional shapes than divisional shapes V43 and V46 shown in Figure 13 is omitted from the description. After step S355 is executed and a positive judgment result is obtained, the process returns to the first pixel group movement process (see Figure 6), and step S36 is executed.
[0121] Step S36 stores a set of segmented shapes (tentatively called a "group of segmented shapes") obtained in step S35b for each convex angle specified in step S34. The stored group of segmented shapes for each convex angle become candidates for the source pixel group Gk.
[0122] In step S37, it is determined whether all convex angles have been selected. For example, if in step S34 a convex angle with point P44 as its vertex is specified but no convex angle with point P43 as its vertex is specified, then in step S34, which is accessed via step S37, the convex angle with point P43 as its vertex will be specified.
[0123] Figure 14 illustrates the state after a convex angle with point P43 as its vertex is specified in step S34, and then steps S352, S353, and S354 are executed via step S351 in step S35b, resulting in the acquisition of segmented figure V49. Segmented figure V49 is further expanded starting from point P43.
[0124] Figure 15 illustrates a state in which the segmented figure V49 is enlarged to obtain segmented figure V47. Segmented figure V47 is in contact with point P54, and the judgment result in step S353 is positive, so step S355 is followed by step S356. Point P443 is located in the part where the outline of segmented figure V47 is in contact with edge L434, and is the furthest from the original starting point P43. After point P443 is specified as the starting point in step S356, steps S352, S353, and S354 are executed.
[0125] Figure 16 shows the situation where a segmented figure V48 is obtained, starting from point P443. The segmented figure V48 touches vertex P44 without touching the third region R3. If segmented figure V48 is obtained, the judgment result in step S353 is negative, and the judgment in step S354 is positive, so step S355 is executed. Segmented figure V4 8 Since it touches point P44 without touching the third region R3, it is a segmented figure V4. 8 This is because the resulting segmented figure, obtained by further enlargement, extends beyond the fourth region R46, which is undesirable.
[0126] The generation of smaller divisional shapes than the divisional shapes V47 and V48 shown in Figure 16 is omitted from the description. After step S355 is executed and a positive judgment result is obtained, the process returns to the first pixel group movement process (see Figure 6), and step S36 is executed. In step S36, the group of divisional shapes obtained based on the convex angle including vertex P43 is stored.
[0127] In the examples shown in Figures 12 to 16, both the convex angle containing vertex P44 and the convex angle containing vertex P43 were selected as starting points. Therefore, the result of the decision in step S37 is positive, and step S38 is executed.
[0128] In step S38, the group of segmented figures obtained based on the convex angle containing vertex P44 and the group of segmented figures obtained based on the convex angle containing vertex P43 are compared, and one of them is selected as the source pixel group Gk.
[0129] This comparison is based on either the size and / or number of segmented shapes included in the segmented shape group (abbreviated as "AND / OR" in Figure 6). A larger size of the segmented shapes contributes to a greater reflection of the pixel relationships in the first image F1 in the third image F3. If the size of the segmented shapes is the same, a larger number of segmented shapes of that size contributes to a greater reflection of the pixel relationships in the first image F1 in the third image F3. Below, an example is given in step S38 where the segmented shape group obtained based on the convex angle including vertex P44 is adopted as the source pixel group Gk.
[0130] As a result of the processes described in Figures 8 to 16, a positive judgment is obtained in step S39, and step S30 is executed. Step S30 performs a second affine transformation. This process can be seen as rotating and moving the segmented figures obtained in steps S35a and S38 and placing them in the fifth region R5. From this perspective, step S30 is labeled "First Placement Process" in Figure 6.
[0131] Figure 17 shows the situation after step S30 is executed. Pixel groups D41, D42, D43, D44, E42, E43, and E46 are exemplified as the destination pixel group Hk. Pixel groups D41, D42, D43, and D44 are placed in the fifth region R56, and pixel groups E42, E43, and E46 are placed in the fifth region R55.
[0132] Pixel groups D41, D42, D43, and D44 are obtained by applying a second affine transformation to the segmented shapes T41, T42, T43, and T44 (see Figure 10), respectively. Pixel groups E42, E43, and E46 are obtained by applying a second affine transformation to the segmented shapes V42, V43, and V46 (see Figure 13), respectively.
[0133] For example, by the first affine transformation, the fifth region R55 is congruent to the fourth region 46, and the fifth region R56 is congruent to the fourth region 45. The pixel group D41, D42, D43, D44 obtained by performing the second affine transformation on the piecewise figures T41, T42, T43, T44 that were contained in the fourth region R45, is contained in the fifth region R56. The pixel group E42, E43, E46 obtained by performing the second affine transformation on the piecewise figures V42, V43, V46 that were contained in the fourth region R46, is contained in the fifth region R55.
[0134] The fact that the source pixel group Gk is inscribed in the fourth region R4, and the destination pixel group Hk is inscribed in the fifth region R5, contributes to enlarging the segmented shape, and consequently, to the relationship between pixels in the first image F1 being more clearly reflected in the third image F3.
[0135] In the example described above, the piecewise figure T41, which is expanded starting from vertex P42 and results in a positive judgment in step S353, is tangent to edge L523 and inscribed in the fourth region R45. The pixel group D41 obtained by performing a second affine transformation on the piecewise figure T41 is positioned to include vertex P53, is tangent to edge L423, and is inscribed in the fifth region R56.
[0136] Starting from vertex P44, the segmented figure V42, which is expanded to make the judgment result in step S353 positive, is tangent to edge L541 and inscribed in the fourth region R46. The pixel group E42 obtained by performing a second affine transformation on segmented figure V42 is placed at a position including vertex P51, is tangent to edge L441, and is inscribed in the fifth region R55.
[0137] Figure 18 is a flowchart illustrating the contents of the second pixel group movement process performed in step S4. 4 isThe process includes steps S41, S42, and S43, which are executed in this order. Step S41 selects a group of pixels in the fourth region R4 that were not used in the first placement process executed in step S30. For example, if we consider the fourth region R45, this group of pixels is the group of pixels present in the fourth region R45 other than the segmented figures T41, T42, T43, T44 (see Figure 10), and if we consider the fourth region R46, this group of pixels present in the fourth region R46 other than the segmented figures V42, V43, V46 (see Figure 13).
[0138] Step S42 divides the pixel group selected in step S41 into individual pixels. In terms of the first step, this corresponds to the case where the number Mk is set to a value of 1. Step S43 applies a second affine transformation to each of the pixels divided in step S42, and places the pixels in the region of the fifth area R5 where a pixel group has not yet been determined. In Figure 18, step S43 is labeled "Second Placement Process". This placement of pixels corresponds to determining pixels in the region of the fifth area R5 where no pixels have yet been placed.
[0139] The pixels determined by step S43 are, for example, pixels located in the fifth region R56 other than pixel groups D41, D42, D43, and D44 when considering the fifth region R56, and pixels located in the fifth region R55 other than pixel groups E42, E43, and E46 when considering the fifth region R55. Placed These are pixels (see Figure 17).
[0140] Once step S43 is completed, step S4 also ends, and the process returns to step S5.
[0141] Since the second image F2 is not used as a training image, the source pixel group Gk may be deleted or retained if the destination pixel group Hk is obtained. As long as the second affine transformation is followed, the source pixel group Gk may be moved to obtain the destination pixel group Hk, or the source pixel group Gk may be duplicated to obtain the destination pixel group Hk.
[0142] <4. Other examples of first pixel group movement processing> Figure 19 is a flowchart of a first alternative example of the contents of the first pixel group movement process performed in step S3 (see Figure 5). Step S3 in this first alternative example includes steps S31, S32, S33, and S39 shown in Figure 6, as well as steps S30 and S34, which are limited as described later, and steps S35c and S360.
[0143] The content of the processes performed in steps S31, S32, S33, and S39 in Figure 19, and the order of these processes, are the same as those in Figure 6. Step S35c in Figure 19 is performed after step S34 and before step S39, similar to step S35b in Figure 6. Step S360 is performed when the judgment in step S39 is a positive result. After step S360 is performed, step S30 is performed. After step S30 is performed, the process returns to step S4 (see Figure 5), similar to the flowchart in Figure 6.
[0144] In the first example, in step S34, the convex angle that yields the largest number of section figures among the convex angles of the fourth region R4 specified in step S32 is specified. By obtaining section figures based on this specified convex angle, steps S35b, S36, S37, and S38 (see Figure 6) are omitted.
[0145] In Figure 19, step S34 uses the distance along the outer boundary of the fourth region R4 from the vertex of the convex angle as a criterion for specifying one of the convex angles of the fourth region R4 specified in step S32. For a single convex angle, the aforementioned distance is the shortest distance from the vertex of the convex angle to the outer boundary Q along the edge that intersects the outer boundary Q with the vertex of the convex angle.
[0146] The convex angle with the longest distance is specified in step S34. If the above distances for different convex angles are equal, any of the convex angles may be specified in step S34.
[0147] In the case illustrated in Figures 8 to 17, in the fourth region R46, the convex angle having vertex P44 is flanked by edges L441 and L434. Edge L434 does not intersect with the outer boundary Q, while edge L441 intersects with the outer boundary Q at point P61. The aforementioned distance for the convex angle having vertex P44 is the distance from vertex P44 to point P61 along edge L441. This can also be considered as the distance from vertex P44 along edge L441 to the third region R3.
[0148] In the fourth region R46, the convex angle with vertex P43 is flanked by edges L434 and L423. Edge L434 does not intersect with the outer boundary Q, while edge L423 intersects with the outer boundary Q at points P63 and P64. The distance from vertex P43 to point P64 along edge L423 is shorter than the distance from vertex P43 to point P63 along edge L423. The aforementioned distance for the convex angle with vertex P43 is the distance from vertex P43 to point P64 along edge L423. This can also be considered as the distance from vertex P43 along edge L423 to the third region R3.
[0149] The distance from vertex P44 along edge L441 to point P61 is longer than the distance from vertex P43 along edge L423 to point P64. Therefore, in step S34, the convex angle having vertex P44 is specified.
[0150] Step S35c corresponds to steps S35a and S35b (see Figures 6 and 7). Step S35c performs a division process based on one convex angle of the fourth region R4 specified in step S32. This single convex angle is the single convex angle when the determination result in step S33 is "single", and when the determination result in step S33 is "multiple", it is the convex angle specified as described above in step S34. Figure 20 is a flowchart illustrating the contents of the division process performed in step S35c.
[0151] Figure 21 is a flowchart illustrating the contents of the first placement process performed in step S30 used in Figure 19. Step S30 includes steps S302, S303, S304, S305, and S309. Figure 22 is a flowchart illustrating the contents of step S305.
[0152] <4-1. Example of the contents of the splitting process (step S35c)> Figures 23 to 29 are conceptual diagrams illustrating the partitioning process performed in step S35c. Figures 23 to 29 show the fourth region R4 and a portion of the third region R3. The fourth region R4 shown in Figures 23 to 29 corresponds to the fourth region R46 exemplified in Figures 8 to 17. Each pixel in the fourth region R4 is drawn as a square. The following explanation of step S35c will use Figures 23 to 29.
[0153] Step S35c corresponds to steps S351a, S354a, S352a, S353a, S353b, S355a, S36a , It has S356a, S357a, S358a, and S359a.
[0154] Step S351a sets the initial values for the start and end points that determine the divisional shape in step S35c. Specifically, if the determination result in step S33 is "single", the vertices of the single convex angle are designated as the start and end points of the divisional shape. If the determination result in step S33 is "multiple", the vertices of the convex angles specified as described above are designated as the start and end points of the divisional shape in step S34.
[0155] The fourth region R4 contains vertices P4a and P4b. Vertices P4a and P4b correspond to vertices P44 and P43, respectively, as illustrated in Figures 8 through 17.
[0156] Since the fourth region R4 has multiple convex angles, each with a different vertex, in step S3, step S33 is executed, followed by step S34. The fourth region R4 and the third region R3 are in contact at points P4c and P4d, and at points P4c and P4d, the outline Q (not shown) and the edges of the fourth region R4 intersect. Points P4c and P4d correspond to vertices P61 and P64, respectively, as illustrated in Figures 8 to 17.
[0157] The distance between vertices P4a and P4c (10 pixels in Figures 23 to 29) is longer than the distance between vertices P4b and P4d (6 pixels in Figures 23 to 29). Therefore, in step S34, vertex P4a is specified, and in step S351a, vertex P4a is specified as the start point P4as and the end point P4ae.
[0158] After step S351a is performed, step S354a is performed. In step S354a, it is determined whether the endpoint can move one unit away from the starting point. This movement is limited to the region where pixels exist. Step S354a corresponds to step S354 in Figure 7. In Figures 23 to 29, this one unit corresponds to one pixel in each of the two adjacent directions, and the direction of movement is the diagonal direction of the square exhibited by the pixels. In the example shown in Figure 23, even if the endpoint P4ae moves from the starting point P4as towards the lower right in the figure, the endpoint P4ae is in the fourth region R4, so the result of the determination in step S354a is positive.
[0159] If the result of the decision in step S354a is positive, the process executed in step S352a moves the end point one unit away from the starting point. If the result of the decision in step S354a is negative, step S355a is executed.
[0160] Figure 24 shows the state after step S352a has been performed, starting from the state in Figure 23. In the figure, the hatched squares represent sectioned figures defined by the starting point P4as and the ending point P4ae.
[0161] The segmented shape is enlarged by connecting pixels within the fourth region R4. This enlargement starts at a position away from the third region R3. The segmented shape can be set as the source pixel group Gk.
[0162] After step S352a is executed, step S353a is executed. In step S353a, it is determined whether the segmented figure defined by the start and end points contains at least one pixel in the third region R3 (abbreviated as "third region pixel" in Figure 20). If the result of this determination is negative, step S354a is executed again. In the state illustrated in Figure 24, the result of the determination in step S353a is negative, and the endpoint P4ae moves by executing steps S354a and S352a again. This movement expands the segmented figure to the size of four pixels (see Figure 25).
[0163] Figure 26 shows the state after steps S354a and S352a have been repeatedly executed and the endpoint P4ae has reached the boundary between the third region R3 and the fourth region R4. After this state, steps S353a, S354a, and S352a are executed, causing the endpoint P4ae to move to the third region R3, resulting in the state shown in Figure 27.
[0164] In the state shown in Figure 27, the segmented figure defined by the starting point P4as and the ending point P4ae contains one pixel from the third region R3. Therefore, the judgment result in step S353a is positive, and step S353b is executed. The segmented figure should consist only of pixels included in the fourth region R4. Therefore, a process is needed to cancel out the process executed in step S352a immediately preceding the positive judgment result obtained in step S353a. This canceling process is executed in step S353b.
[0165] Specifically, the process in step S353b moves the endpoint one unit closer to the starting point. By executing step S353b from the state shown in Figure 27, the state shown in Figure 26 is obtained again.
[0166] After step S353b is executed, step S355a is executed. In step S355a, it is determined whether the size of the segmented figure obtained in step S353b is less than or equal to a predetermined value, similar to step S355 illustrated in Figure 7.
[0167] If the result of the judgment in step S355a is negative, the divided figure used in the judgment is stored in the divided figure list (not shown) so that it can be used in the first arrangement process (see step S30 in Figure 19). If the result of the judgment in step S355a is positive, the divided figure used in the judgment is not used in the first arrangement process, so step S36a is not executed.
[0168] After a positive judgment result is obtained in step S355a, or after the judgment result in step S355a is negative and step S36a is executed, step S356a is executed. A provisional starting point is set by the execution of step S356a. A provisional starting point is a candidate starting point that will become the starting point of the divided figure under predetermined conditions. The provisional starting point is located away from the starting point towards the ending point in each of the two directions in which pixels are adjacent to each other, in terms of the components of the distance from the starting point to the ending point.
[0169] Figure 28 shows the state after step S356a is executed. In Figure 28, the segmented shape V4a, which is stored in the segmented shape list by step S36a, is represented by a group of pixels with cross-hatching. For example, the segmented shape list is represented as Lv={V4a}.
[0170] Figure 28 also illustrates the temporary starting points P4a1s and P4a2s. The temporary starting points P4a1s and P4a2s are determined based on the starting point P4as and the ending point P4ae. Specifically, the temporary starting points P4a1s and P4a2s are set at a distance (in this case, 5 pixels) from the starting point P4as to the ending point P4ae in each of the two directions in which pixels are adjacent.
[0171] A provisional starting point is adopted if one of the following conditions (α) or (β) and condition (γ) are satisfied: (α) The provisional starting point is located in the fourth region R4; (β) The provisional starting point is located on the outer edge of the fourth region R4 but not on the outer edge of the third region R3; (γ) The provisional starting point is different from the previous starting point. Referring to Figure 28, both provisional starting points P4a1s and P4a2s satisfy conditions (β) and (γ).
[0172] After step S356a is executed, step S357a is executed. Execution of step S357a results in the registration of provisional starting points that satisfy predetermined conditions (corresponding to the above-mentioned conditions (α), (β), and (γ)) as starting points in an ordered starting point list (not shown). In the above example, both provisional starting points P4a1s and P4a2s are registered as starting points in the starting point list. For example, the starting point list can be represented as Ls={P4a1s,P4a2s}.
[0173] After step S357a is executed, step S358a is executed. In step S358a, it is determined whether a start point is registered in the start point list. If this determination is positive, step S359a is executed, and the start and end points are set again. The start point registered at the top of the start point list is set as the new start and end points. The execution of step S359a is a resetting of the start and end points, and can also be said to be an update of them.
[0174] After step S359a is executed, step S354a is executed. When step S359a is executed, the start point used to reset the start and end points is removed from the start point list.
[0175] In the example above, when step S358a is executed, the starting point list is Ls={P4a1s,P4a2s}. When step S359a is executed, the starting point P4a1s, which was registered at the top of the starting point list, is reset as both the starting and ending point, and the starting point list is updated to Ls={P4a2s}. In this way, the starting point list retains starting points that were not used in setting up the divided shape.
[0176] Based on the starting point P4a1s, steps S354a, S352a, S353a, and S353b are executed to obtain the segmented figure V4a1 (see Figure 29). The segmented figure V4a1 has a size of 9 pixels, and assuming that the predetermined value in step S355a is 4 pixels, for example, step S36a is executed. After the execution of step S36a, the segmented figure list is represented as Lv={V4a,V4a1}.
[0177] Furthermore, step S356a sets temporary starting points P4a3s and P4a4s. Temporary starting point P4a3s satisfies conditions (β) and (γ). Temporary starting point P4a4s satisfies conditions (α) and (γ). Step S357a adds temporary starting points P4a3s and P4a4s to the starting point list. As a result, the starting point list becomes Ls={P4a2s,P4a3s,P4a4s}.
[0178] The result of the decision in step S358a, which is executed immediately after this, is positive, and step S359a is executed. In step S359a, the starting point P4a2s, which was registered at the top of the starting point list, is adopted as the new starting and ending point, and the starting point list becomes Ls={P4a3s,P4a4s}.
[0179] Steps S354a, S352a, and S353a are executed based on the starting point P4a2s, and the segmented figure V4a2 is obtained (see Figure 29). For the endpoint P4a2e obtained before a positive judgment result is obtained in step S353a, the judgment result in step S354a becomes negative. This situation corresponds to the case where a positive judgment result is obtained in step S354 as illustrated in Figure 7.
[0180] After the segmented figure V4a2 is obtained, step S36a is executed, and the segmented figure list is represented as Lv={V4a,V4a1,V4a2}. Furthermore, step S356a is executed to set temporary starting points P4a5s and P4a6s (temporary starting point P4a6s coincides with vertex P4b). Temporary starting point P4a5s satisfies conditions (α) and (γ), and temporary starting point P4a6s satisfies conditions (β) and (γ). After executing step S357a, the starting point list becomes Ls={P4a3s,P4a4s,P4a5s,P4a6s}.
[0181] After this, step S359a is executed in the order registered in the starting point list. However, in the example above, the divisional figures obtained based on the starting points P4a3s and P4a4s result in a positive judgment in step S355a, and the provisional starting points obtained thereafter do not satisfy either condition (α) or (β).
[0182] Based on the starting point P4a6s, a subdivision figure is not actually obtained, and its size is treated as zero for convenience. As a result, the judgment result in step S355a is positive, and the provisional starting point set in step S356a coincides with the starting point P4a6s, so condition (γ) is not satisfied.
[0183] Starting point P4a 5 Based on s, the resulting segmented figure results in a positive judgment in step S355a, and the subsequent provisional starting point also does not satisfy the predetermined conditions in step S357a.
[0184] As illustrated in Figure 29, even when there are no more starting points registered in the starting point list and the result of the judgment in step S358a becomes negative, the division figure list remains Lv={V4a,V4a1,V4a2}. Since the result of the judgment in step S358a is negative, the process returns to step S39 (see Figure 19).
[0185] The segmented shapes V4a, V4a1, and V4a2 illustrated in Figure 29 correspond to the segmented shapes V42, V43, and V46 illustrated in Figure 13, respectively. The segmented shapes V47 and V48, based on the convex angle including vertex P43, illustrated in Figure 16, cannot be obtained in the first pixel group movement process shown in Figure 19.
[0186] Referring to Figure 19, steps S32, S33, and S35c are executed (and step S34 is also executed if the specified fourth region R4 has multiple convex angles) until all fourth regions R4 are specified based on the determination in step S39. After all fourth regions R4 have been specified based on the determination in step S39, step S360 is executed.
[0187] The process performed in step S360 sorts the compartmentalized shapes registered in the compartmentalized shape list obtained for each of the specified fourth regions R4, grouping them together based on their size. Referring to the examples in Figures 8 to 17, the compartmentalized shape list is Lv={V42,V46,V43,T41,T43,T42,T44}.
[0188] After step S360 is executed, step S30 shown in Figure 21 is executed. Step S30 performs the first placement process. In step S30, steps S302, S303, S305, and S309 are executed in this order.
[0189] In step S302, the fifth region R5 is specified. Steps S303 and S304 are executed for the fifth region R5 specified in step S302 in the same manner as steps S33 and S34 described with reference to Figure 19.
[0190] In step S302, the convex angle that yields the largest number of large partitioned regions is identified among the convex angles of the fifth region R5 specified in step S302. The partitioned figures registered in the partitioned figure list are placed in the partitioned regions as the destination pixel group Hk after undergoing a second affine transformation. The partitioned regions are expanded by concatenating pixels within the fifth region R5. This expansion starts at a position away from the third region R3. Obtaining partitioned regions based on the identified convex angle as described above contributes to placing large partitioned figures in the partitioned regions.
[0191] Step S304 uses the distance along the outer boundary of the fifth region R5 from the vertex of the convex angle as a criterion for specifying one of the convex angles of the fifth region R5 specified in step S302. For a single convex angle, the aforementioned distance is the shortest distance from the vertex of the convex angle along the edge that intersects the second region R2 with respect to the second region R2.
[0192] The convex angle with the longest distance is specified in step S304. If the above distances for different convex angles are equal, any of the convex angles may be specified in step S304.
[0193] In the case illustrated in Figures 8 to 17, in the fifth region R55, the convex angle having vertex P51 is flanked by edges L541 and L512. Edge L512 does not intersect with the second region R2, while edge L541 intersects with the second region R2 at point P61. The aforementioned distance for the convex angle having vertex P51 is the distance from vertex P51 to point P61 along edge L541. This can also be considered as the distance from vertex P51 along edge L541 to the third region R3.
[0194] In the fifth region R55, the convex angle with vertex P52 is flanked by edges L512 and L523. Edge L512 does not intersect with the second region R2, while edge L523 intersects the second region between points P62 and P63. The distance from vertex P52 to point P62 along edge L523 is shorter than the distance from vertex P52 to any point where edge L523 intersects with the second region R2. The aforementioned distance for the convex angle with vertex P52 is the distance from vertex P52 to point P62 along edge L523. This can also be considered as the distance from vertex P52 to the third region R3 along edge L523.
[0195] The distance from vertex P51 along edge L541 to point P61 is longer than the distance from vertex P52 along edge L523 to point P62. Therefore, step S304 specifies the convex angle having vertex P51.
[0196] Step S305 performs a division process based on one of the convex angles of the fifth region R5 specified in step S302. This single convex angle is the single convex angle when the determination result in step S303 is "single", and when the determination result in step S303 is "multiple", it is the convex angles specified as described above in step S304.
[0197] After step S305 is executed, step S309 determines whether all fifth regions R5 have been designated. As long as this determination is negative, the process returns to step S302. When there are no more fifth regions R5 that have not undergone partitioning in step S305, the determination result of step S309 becomes positive, and the process returns to step S4 (see Figure 5).
[0198] Referring to the examples in Figures 8 to 17, the fifth region R55, R56 is in step S30 2 These are specified sequentially by S309. For example, the fifth region R55 is specified before the fifth region R56 is specified.
[0199] <4-2. Example of the contents of the splitting process (step S305)> Figures 30 to 35 are conceptual diagrams illustrating the partitioning process performed in step S305 (see Figure 22). Figures 30 to 35 show the fifth region R5 and a portion of the third region R3. The fifth region R5 shown in Figures 30 to 35 corresponds to the fifth region R55 exemplified in Figures 8 to 17. Each pixel in the fifth region R5 is drawn as a square. The following explanation of step S305 will use Figures 30 to 35.
[0200] Step S305 is a step of steps S351b, S354b, S352b, S353c, S353d, S355b , S36b , S356b, S356c, S357b, S358b , It has S359b.
[0201] Step S351b sets the initial values for the start and end points that determine the division region in step S305. Specifically, if the determination result in step S303 is "single", the vertex of that single convex angle is designated as the start and end points of the division region. If the determination result in step S303 is "multiple", the vertices of the convex angles identified as described above in step S304 are designated as the start and end points of the division region.
[0202] The fifth region, R5, contains vertices P5a and P5b. Vertices P5a and P5b correspond to vertices P51 and P52, respectively, as illustrated in Figures 8 through 17.
[0203] Since the fifth region R5 has multiple convex angles, each with a different vertex, in step S30, step S303 is executed first, followed by step S304. The fifth region R5 and the third region R3 are in contact at points P5c and P5d, and at points P5c and P5d, the edges of the second region R2 (not shown) and the edges of the fifth region R5 intersect. Points P5c and P5d correspond to vertices P61 and P62, respectively, as illustrated in Figures 8 to 17.
[0204] The distance between vertices P5a and P5c (10 pixels in Figures 30 to 35) is longer than the distance between vertices P5b and P5d (6 pixels in Figures 30 to 35). Therefore, vertex P5a is specified in step S304, and vertex P5a is specified as the start point P5as and the end point P5ae in step S351b.
[0205] After step S351b is executed, step S354b is executed. In step S354b, it is determined whether the endpoint can move one unit away from the starting point. This movement is limited to the region where pixels exist. In Figures 30 to 35, this one unit corresponds to one pixel in each of the two adjacent directions, and the direction of movement is the diagonal direction of the square formed by the pixels. In the example shown in Figure 30, even if the endpoint P5ae moves from the starting point P5as towards the lower left in the figure, the endpoint P5ae is in the fifth region R5, so the result of the determination in step S354b is positive.
[0206] If the result of the decision in step S354b is positive, the process executed in step S352b moves the end point one unit away from the starting point. If the result of the decision in step S354b is negative, step S355b is executed.
[0207] Figure 31 shows the state after step S352b has been executed from the state shown in Figure 30. In the figure, the hatched squares indicate the divided region defined by the starting point P5as and the ending point P5ae.
[0208] After step S352b is executed, step S353c is executed. In step S353c, it is determined whether the divided region defined by the start and end points contains at least one pixel (abbreviated as "third region pixel" in Figure 22) within the third region R3. If the result of this determination is negative, step S354b is executed again. In the state illustrated in Figure 31, step S35 3c The judgment result is negative, and the endpoint P5ae moves by executing steps S354b and S352b again. This movement expands the divided region to the size of four pixels.
[0209] Figure 32 shows the state after steps S354b and S352b have been repeatedly executed, and the endpoint P5ae has reached the boundary between the third region R3 and the fifth region R5. After this state, steps S353c, S354b, and S352b are executed, causing the endpoint P5ae to move to the third region R3, resulting in the state shown in Figure 33.
[0210] In the state shown in Figure 33, the segmented figure defined by the starting point P5as and the ending point P5ae includes one pixel in the third region R3. Therefore, the judgment result in step S353c is positive, and step S353d is executed. The segmented region is the area in which the segmented figure is placed and should consist only of pixels included in the fifth region R5. Therefore, a process is needed to cancel out the process executed in step S352b immediately preceding the positive judgment result obtained in step S353c. This canceling process is executed in step S353d.
[0211] Specifically, the process in step S353d moves the endpoint one unit closer to the starting point. By executing step S353d from the state shown in Figure 33, the state shown in Figure 32 is obtained again. In Figure 34, the divided region V5a obtained by the execution of step S353d is shown by a group of pixels that have been cross-hatched.
[0212] After step S353d is executed, step S355b is executed. In step S355b, it is determined whether there are any partitioned figures smaller than or equal to the size of the partitioned area. Step S35 5b The partitioned area referred to is the partitioned area obtained by step S353d, which was executed immediately before. Step S35 5b The term "divided shape" refers to a divided shape that is registered in the divided shape list immediately before step S353d is executed. In the example above, step S355b determines whether there are any divided shapes in the divided shape list Lv={V4a,V4a1,V4a2} that are smaller than or equal to the size of the divided area V5a.
[0213] If the result of the judgment in step S355b is positive, step S36b is executed. In step S36b, a second affine transformation is applied to the divisional region used for the judgment, and divisional figures are placed therein. At this time, the largest divisional figure among the divisional figures that result in a positive judgment in step S355b is placed in the divisional region. In this placement, for example, the vertices of the divisional figure are placed at the starting points that define the divisional region.
[0214] In the example above, the partitioned area V5a has a size equivalent to 25 pixels. All of the partitioned shapes V4a, V4a1, and V4a2 stored in the partitioned shape list are 25 pixels or less in size, and step S3 55 The result of judgment b is positive.
[0215] Of the divided shapes V4a, V4a1, and V4a2, the largest divided shape V4a is placed in the divided area V5a. (List of divided shapes) to In step S360 (see Figure 19), the compartmentalized shapes are registered in order from largest to smallest. For example, the compartmentalized shapes to be placed in the compartmentalized area are selected from the beginning of the compartmentalized shape list.
[0216] In step S36b, the partitioned shape placed in the partitioned area is removed from the partitioned shape list. Following the example above, partitioned shape V4a is placed in partitioned area V5a using the second affine transformation, and the partitioned shape list is updated to Lv={V4a1,V4a2}.
[0217] After step S36b is executed, step S356c is executed. If the result of the determination in step S355b is negative, step S356b is executed. A provisional starting point is set by the execution of steps S356b and S356c. A provisional starting point is a candidate starting point that will become the starting point of the divided area under predetermined conditions. The setting of the provisional starting point differs between step S356b and step S356c.
[0218] The temporary starting point set in step S356b is located away from the starting point towards the ending point by a component of the distance from the starting point to the ending point in each of the two directions in which the pixels are adjacent.
[0219] Figure 34 shows the state after step S356b is executed. In Figure 34, temporary start points P5a1s and P5a2s are shown as examples. Temporary start points P5a1s and P5a2s are determined based on the start point P5as and the end point P5ae. Specifically, temporary start points P5a1s and P5a2s are set at a distance (in this case, 5 pixels) from the start point P5as to the end point P5ae in each of the two directions in which pixels are adjacent.
[0220] A provisional starting point is adopted if one of the following conditions (δ) or (ε) and condition (ζ) are satisfied: (δ) The provisional starting point is located in the fifth region R5; (ε) The provisional starting point is located on the outer edge of the fifth region R5 but not on the outer edge of the third region R3; (ζ) The provisional starting point is different from the previous starting point. Referring to Figure 34, both provisional starting points P5a1s and P5a2s satisfy conditions (ε) and (ζ).
[0221] After step S356b is executed, step S357b is executed. Execution of step S357b registers provisional start points that satisfy predetermined conditions (corresponding to the above-mentioned conditions (δ), (ε), and (ζ)) as start points, in an ordered manner, in the start point list (not shown). In the above example, both provisional start points P5a1s and P5a2s are registered as start points in the start point list. For example, the start point list can be represented as Ls={P5a1s,P5a2s}.
[0222] After step S357b is executed, step S358b is executed. In step S358b, it is determined whether a start point is registered in the start point list. If this determination is positive, step S359b is executed, and the start and end points are set again. The start point registered at the top of the start point list is set as the new start and end points. The execution of step S359b is a resetting of the start and end points, and can also be said to be an update of them.
[0223] After step S359b is executed, step S354b is executed. When step S359b is executed, the start point used to reset the start and end points is removed from the start point list.
[0224] In the example above, when step S358b is executed, the start point list is Ls={P5a1s,P5a2s}. When step S359b is executed, the start point P5a1s, which was registered at the top of the start point list, is reset as both the start and end point, and the start point list is updated to Ls={P5a2s}. In this way, the start point list retains start points that were not used to set the partitioned area.
[0225] The temporary starting points set in step S356c are located at the components of the placed divisional figure in each of the two directions in which the pixels are adjacent, and are located away from the starting point. In the example above, the size of the divisional region V5a and the size of the divisional figure V4a are the same, and in step S356c, the temporary starting points P5a1s and P5a2s are set in the same way as in step S356b.
[0226] For example, consider the case where step S353d is executed to obtain a partitioned region V5a, and the partitioned shape list is Lv={V4a1,V4a2}. In this case, step S36b is executed to place partitioned shape V4a1 in partitioned region V5a by a second affine transformation. The size of partitioned shape V4a1 is four pixels in each of the two directions in which pixels are adjacent. Now consider the case where the vertices of partitioned shape V4a1 coincide with the starting point P5as and partitioned shape V4a1 is placed in partitioned region V5a by a second affine transformation. Under this assumption, the temporary starting point P5a1s is one pixel closer to the starting point P5as than the temporary starting point P5a1s set in step S356b; and the temporary starting point P5a2s is one pixel closer to the starting point P5as than the temporary starting point P5a2s set in step S356b.
[0227] In this case, we consider the scenario where the vertices of the segmented figure V4a1 coincide with the endpoint P5ae, and the segmented figure V4a1 is placed in the segmented region V5a by a second affine transformation. Under this scenario, the temporary starting points P5a1s and P5a2s coincide with the temporary starting points P5a1s and P5a2s set in S356b, respectively. In this case, a group of pixels with a width of one pixel remains in the segmented region V5a, arranged in an L-shape between the placed segmented figure V4a1 and the starting point P5as. It is difficult to place a segmented figure in the position where these pixels remain.
[0228] The fact that the vertices of the divisional shapes placed within a divisional area coincide with the starting points that define that divisional area contributes to placing large divisional shapes within the divisional area.
[0229] Steps S357b, S358b, and S359 are performed after step S356c has been executed. b In this case, steps S357b, S358b, and S359 are executed after step S356b is performed. b The same process is performed.
[0230] After step S354b is executed with the starting point P5a1s reset as both the starting and ending point, steps S352b, S353c, and S353d are executed to obtain the ending point P5a1e and the divided region V5a1 (see Figure 35). Similarly, after step S354b is executed with the starting point P5a2s reset as both the starting and ending point, step S352b , S353c , S353d is executed to obtain the endpoint P5a2e and the partitioned region V5a2 (see Figure 35).
[0231] Each time step S357b is executed, a starting point is added to the starting point list, and each time step S359b is executed, a starting point is removed from the starting point list. Similar to the division process performed in step S35c, the starting points registered in the starting point list are updated, and eventually the judgment result of step S358b becomes negative. If the judgment result of step S358b is negative, the process returns to step S309 (see Figure 21), and the process returns to step S4 (see Figure 5), where the second pixel group movement process is executed.
[0232] <4-3. Use of Mirror Transformation> When the fourth region R4 and the fifth region R5 are congruent and mirror images of each other, a reflection transformation is used in the second affine transformation. The specific process is explained below using the first pixel group movement process performed in step S3. Referring to the examples in Figures 8 to 17, it is assumed that the fourth region R45 and the fifth region R56 are mirror images of each other, and the fourth region R46 and the fifth region R55 are mirror images of each other. This reflection can also be considered as line symmetry, and the axis of symmetry in this line symmetry corresponds to the line connecting points P61 and P63.
[0233] Figure 36 is a flowchart showing a second example of the contents of the first pixel group movement process. Step S3 includes steps S31, S32, S33, and S39 shown in Figure 6, as well as step S34 and steps S35d and S301 described in Figure 19.
[0234] The content of the processes performed in steps S31, S32, S33, and S39 in Figure 36, and the order of these processes, are the same as those in Figure 6. The content of the process performed in step S34 in Figure 36 and its relationship to the process performed in step S33 are the same as those in Figure 19.
[0235] Step S301 is performed after step S31 has been executed and the third region R3, fourth region R4, and fifth region R5 have been recognized, but before the fourth region R4 is specified in step S32.
[0236] The process performed in step S301 generates a coordinate transformation matrix from the fourth region R4 to the fifth region R5. A reflection transformation is used in this coordinate transformation matrix.
[0237] After step S301 is executed, steps S32, S33, and S34 are executed as described above. Step S35d performs a division process based on one convex angle of the fourth region R4 specified in step S32. This single convex angle is the single convex angle when the determination result in step S33 is "single", and when the determination result in step S33 is "multiple", it is the convex angle specified as described above in step S34.
[0238] Figure 37 is a flowchart illustrating the contents of the partitioning process performed in step S35d. This partitioning process has the same configuration as the partitioning process performed in step S35c (see Figure 20), but with step S36a replaced by step S302.
[0239] The process performed in step S302 moves the divided figure to the fifth region using the coordinate transformation matrix generated in step S301 (see Figure 36). Instead of this movement, duplication may be performed.
[0240] Following the examples shown in Figures 8 to 17, by moving the segmented figures T41, T42, T43, T44, V42, V43, and V46 through mirror transformation, pixel groups D41, D42, D43, D44, E42, E43, and E46 are obtained, respectively.
[0241] However, the arrangement of pixels in pixel group D41 is a mirror image of the arrangement of pixels in segmented figure T41. The same applies to segmented figures T42, T43, T44, V42, V43, and V46.
[0242] < 5 General explanation> For example, the process until the destination pixel group Hk is obtained is re-interpreted as follows, taking the pixel groups D41, D42, D43, D44, E42, E43, E46 described above as examples.
[0243] A conversion step is executed, which is a step of performing a first affine transformation on the first image F1 occupying the first region R1 to obtain a second image F2 occupying the second region R2.
[0244] The movement step includes a first step and a second step. The first step sets the division figure T41 as the source pixel group Ga, which is a pixel group in which a number Ma greater than 1 is connected and is included in the fourth region R45. The second step performs a second affine transformation on the division figure T41 to obtain the pixel group D41 arranged in the fifth region R56 as the destination pixel group Ha.
[0245] The first step sets the division figure V42 as the source pixel group Gb, which is a pixel group in which a number Mb greater than 1 is connected and is included in the fourth region R46. The second step performs a second affine transformation on the division figure V42 to obtain the pixel group E42 arranged in the fifth region R55 as the destination pixel group Hb.
[0246] By executing the first step and the second step, a teacher image of an object annotated in the third region R3 (in the example according to FIGS. 1 to 4, a spin chuck) can be easily obtained while incorporating features that do not remain only in the luminance distribution of images other than the object.
[0247] The movement step includes a third step and a fourth step. Looking at the fourth region R45, in the third step, the division figures T42, T43, T44 are set as the source pixel group Gc, which is a pixel group in which a number Mc equal to or less than the number Ma is connected. When the number Mc is equal to the number Ma, the third step can also be regarded as a repetition of the first step. When the number Mc is smaller than the number Ma, the third step can also be regarded as a repetition of the first step with the number Mk updated. In the example described above, Ma > Mc.
[0248] Regarding the fourth region R46, in the third step, the divided figures V43 and V46 are set as the source pixel group Gd which is a pixel group in which the number Md of pixels connected is less than or equal to the number Mb. When the number Md is equal to the number Mb, the third step can be regarded as a repetition of the first step. When the number Md is smaller than the number Mb, the third step can be regarded as a repetition of the first step with the number Mk updated. In the example described above, Mb > Md.
[0249] In the fourth step, a third affine transformation is applied to the source pixel group obtained in the third step to obtain a destination pixel group. When the third step is regarded as a repetition of the first step, the fourth step can be regarded as a repetition of the second step. Similar to the second affine transformation, the third affine transformation is an affine transformation, which is either rotation or translation or both, or a mirror transformation.
[0250] Regarding the fifth region R56, in the fourth step, a third affine transformation is applied to the divided figure T43 (see FIG. 10) set in the third step, and the pixel group D43 is obtained as the destination pixel group Hk. By repeatedly reducing the number Mc and repeating the third and fourth steps, the divided figures T42 and T44 are set, and the pixel groups D42 and D44 are obtained.
[0251] By executing the third and fourth steps, a teacher image in which the features of an image other than the annotation target are more reflected is obtained.
[0252] The process of setting the divided figures T41 and T43 and obtaining the pixel groups D41 and D43 can be regarded as a repetition of the first and second steps with the number Ma reduced. The process of setting the divided figures T42 and T44 and obtaining the pixel groups D42 and D44 can be regarded as a repetition of the third and fourth steps with the number Mc of よ Small appropriately updated. For example, from this perspective, the first and second steps are executed multiple times prior to the third step.
[0253] Looking at the fifth region R55, the fourth step involves applying a third affine transformation to the segmented shape V46 (see Figure 13) set in the third step to obtain the pixel group E46 as the destination pixel group Hk. By updating the number Md to a smaller value and repeating the third and fourth steps, segmented shape V43 is set and the pixel group E43 is obtained.
[0254] The process of setting segmented shapes V42 and V46 and determining pixel groups E42 and E46 can be seen as a repetition of the first and second processes with a smaller number of elements (Mb). The process of setting segmented shape V43 and determining pixel group E43 can be seen as the third and fourth processes with an even smaller number of elements (Md). For example, from this perspective, the first and second processes are executed multiple times prior to the third process.
[0255] Repeating the first and second steps, or repeating the third and fourth steps. in This allows for the creation of training images that better reflect the features of images other than the annotated target.
[0256] Second placement process (Step S in Figure 18) 43 (Reference) corresponds to the case where the number of pixel groups set in the third step is 1. Due to the third affine transformation performed in the second placement process, information about the brightness distribution, among the image features other than those to be annotated, is more easily reflected in the training image. For example, the third and fourth steps are repeatedly performed until all the pixels of the second image F2 are moved or duplicated in the first region R1.
[0257] In the first example, the segmented shapes V4a, V4a1, and V4a2 are expanded by connecting pixels within the fourth region R4, starting from a position away from the third region R3. These segmented shapes are set as the source pixel group Gk. The segmented regions V5a, V5a1, and V5a2 are expanded by connecting pixels within the fifth region R5, starting from a position away from the third region R3. The destination pixel group Hk is placed in these segmented regions.
[0258] In the second example, the fourth region R4 and the fifth region R5 are mirror images of each other. In this case, the second affine transformation is, for example, a mirror transformation. Similarly, when the fourth region R4 and the fifth region R5 are mirror images of each other, a mirror transformation can be adopted as the third affine transformation.
[0259] <6. Generation and Use of Training Images> figure 38 This is a block diagram illustrating the generation and use of training images. The imaging device 600 captures the configuration 90 to obtain the captured image F0. The configuration 90 includes, for example, a spin chuck 91. The captured image F0 is output from the imaging device 600. For example, a known camera is used as the imaging device 600.
[0260] The captured image F0 is input to the image generator 610. The image generator 610 generates at least one training image Ji (where i is an integer greater than or equal to 1). The training image Ji is output from the image generator 610.
[0261] The image generator 610 has a function schematically illustrated by the hardware image input interface 611. This function corresponds to setting the first image, which is illustrated as step S1 in Figure 5.
[0262] The image input interface 611 outputs a first image F1 selected from the captured image F0. For example, the first image F1 is selected from the captured image F0 based on the condition that at least one spin chuck 91 in the third region R3 can be annotated as a pixel group 100. This condition is set by determining at least one first affine transform together with the input of the captured image F0, or prior to the input of the captured image F0.
[0263] The image generator 610 has a function schematically illustrated by the hardware-based first affine transformation processing unit 612. This function corresponds to the transformation process, and specifically, it is a function that applies a first affine transformation to the first image F1 to obtain the second image F2. This function corresponds to the generation of the second image illustrated as step S2 in Figure 5. The first affine transformation processing unit 612 outputs the second image F2.
[0264] The image generator 610 has a function schematically illustrated by the hardware region recognition unit 613. Specifically, this function is the recognition of the third region R3, the fourth region R4, and the fifth region R5, which is illustrated as step S31 in Figure 6. For this function to be executed, the first image F1 and the second image F2 are used (see Figures 2, 8 to 17). The region recognition unit 613 outputs the third region R3, the fourth region R4, and the fifth region R5.
[0265] The image generator 610 has a function schematically illustrated by the hardware-based segmentation processing unit 614. This function corresponds to either or both of the first and third steps, and corresponds to the process from when step S31 is executed until when step S30 is executed in Figure 6. This function may also include the function corresponding to the steps illustrated in steps S41 and S42 in Figure 18. The fourth region R4 is input to the segmentation processing unit 614, and the segmentation processing unit 614 generates the source pixel group Gk. The source pixel group Gk is output from the segmentation processing unit 614.
[0266] The image generator 610 has a function schematically exemplified by the hardware-based second affine transformation unit 615. This function corresponds to either or both of the second and fourth steps, and corresponds to the first placement process exemplified as step S30 in Figure 6. This function may also include the second placement process exemplified in step S43 in Figure 18. For the function of the second affine transformation unit 615 to be executed, the fifth region R5 and the source pixel group Gk are used (see Figures 3, 4, and 17). The second affine transformation unit 615 generates the destination pixel group Hk and outputs it.
[0267] The image generator 610 has functions schematically illustrated by a synthesizing unit 616 which is hardware. The function is a function of synthesizing a destination pixel group Hk and a third region R3 to generate a third image F3. The function is included in, for example, either or both of the second step and the fourth step, or is associated with either or both of these, and is included in or is associated with the first arrangement process exemplified as step S30 in FIG. 6. The function may be included in or be associated with the second arrangement process exemplified as step S43 in FIG. 18.
[0268] The image generator 610 has functions schematically illustrated by a storage unit 617 which is hardware. Specifically, the function is storing the third image F3, which is exemplified as step S5 in FIG. 5. The third image F3 stored in the storage unit 617 is read out from the storage unit 617 as a teacher image Ji.
[0269] The auxiliary device 620 has a function of adding an auxiliary image to the captured image F0. For example, the auxiliary device 620 generates a display image F9 which is an image with a label identifying the pixel group 100 added to the captured image F0. The display image F9 is output from the auxiliary device 620.
[0270] The auxiliary device 620 has functions schematically illustrated by a learning unit 621 which is hardware. The function is a function of generating a learned model W using the teacher image Ji. As described above, in addition to the teacher image Ji, the first image F1 can be used as a teacher image as described above. The mode in which the first image F1 is adopted as a teacher image is exemplified by a dashed arrow in the figure 38 and is illustrated by a dashed arrow.
[0271] The learned model W is output from the learning unit 621. The learned model W includes a program and parameters having a function of selecting the pixel group 100 from image data.
[0272] The auxiliary device 620 has a function schematically illustrated by the hardware element identification unit 622. This function selects a pixel group 100 from the captured image F0 using a trained model W and adds a marker to identify it in the captured image F0. The marker is added to the captured image F0 to generate the display image F9. The display image F9 is output from the element identification unit 622 and, consequently, from the auxiliary device 620.
[0273] Both the image generator 610 and the auxiliary device 620 are implemented using, for example, a processor and memory. The processor consists of, for example, one or more central processing units (CPUs). The memory consists of a volatile storage medium, exemplified by RAM (Random Access Memory), or a non-volatile storage medium, exemplified by a hard disk drive (HDD) or a solid state drive (SSD). The functions of the storage unit 617 are carried out by, for example, this memory.
[0274] The memory stores, for example, programs and various information. The processor realizes the various functions and processes described above by, for example, reading and executing programs stored in memory. RAM, for example, is used as a workspace at this time and stores temporarily generated or acquired information. At least some of the functions schematically exemplified as hardware in the image generator 610 and auxiliary device 620 may be realized by hardware such as dedicated electronic circuits.
[0275] Display image F9 is input to the display unit 630. The display unit 630 displays display image F9. 38 In this example, a marker 631 that identifies the pixel group 100 is shown.
[0276] The function of the image generator 610 can be described as a method for generating training images Ji, which are training data used for machine learning to recognize an object (a chuck pin in the example above) in the captured image F0. This machine learning creates a trained model W using the training images Ji (or even the first image F1), and is exemplified as a function of the auxiliary device 620.
[0277] As described with respect to the function of the memory unit 617, the third image F3 is adopted as the training image Ji, which is the training data. The third image F3 is obtained by the image processing method exemplified in steps S2, S3, and S4 of Figure 5.
[0278] Prior to the image processing method, as explained regarding the function of the image input interface 611 and as illustrated in step S1 of Figure 5, the first image F1 is set from the captured image F0. The captured image F0 is the object (Figure 38 In this example, the spin chuck 91) is imaged to obtain the result.
[0279] The pixel group 100 is annotated by region extraction. The pixel group 100 is extracted from the pixel group located in the third region R3 in both the first image F1 and the second image F2, specifically from the pixel group located in the third region R3. This region extraction is performed when obtaining each of the third images F3.
[0280] From this, it can be said that the method for generating training data, which is a training image Ji, includes the steps of: adopting the third image F3 as the training image Ji; capturing an object prior to generating the second image F2 and the third image F3 to set the first image F1; and extracting a region of the pixel group 100 in either the first image F1 or the second image F2 each time the third image F3 is obtained.
[0281] figure 39The figure shows another example of the display image F9 displayed by the display unit 630. For example, the imager 600 and the display unit 630 are mounted on the same device, and the above-mentioned markings are visible while imaging the object to be annotated. An example of such a device is smart glasses.
[0282] figure 39 Image 8 is shown as an example of display image F9. In Image 8, the label 80 is shown. In Image 8, the object to be annotated is captured, and a frame 81 surrounding the object is displayed. The label 83 explicitly indicates that the object is a spin chuck as "Spin Chuck". The label 83 is connected to the frame 81 via an arrow 82. The frame 81, arrow 82, and label 83 are included in the label 80.
[0283] <7. Transformation> <7-1. When the 4th region R4 and the 5th region R5 are not congruent> For example, as shown in Figure 17, the fifth region R55 is congruent to the fourth region 46, and the fifth region R56 is congruent to the fourth region 45. However, this disclosure is not limited to the case where the fourth region R4 and the fifth region R5 are congruent. For example, if the fifth region R55 does not include a region the size of the piecewise figure V42, the second affine transformation may not be applied to it, and instead, the second affine transformation may be applied to a smaller region, such as the piecewise figure V46, to arrange the pixel group E46 along vertex P51 and edges L512 and L541.
[0284] <7-2. When the 4th region R4 and the 5th region R5 are congruent> When the first affine transformation is a rotation without translation, and the center of that rotation is the center of the first region R1, then the fourth region R4 and the fifth region R5 are congruent. When such a first affine transformation is adopted, the second affine transformation is also a rotation without translation.
[0285] Let's assume that the second region R2 shown in Figures 8 to 17 is obtained by a first affine transformation involving only rotation around the center of the first region R1. In this case, the pixel groups D41, D42, D43, D44, E42, E43, and E46 are obtained by a second affine transformation involving only rotation around the center of the first region R1 for each of the segmented figures T41, T42, T43, T44, V42, V43, and V46, respectively.
[0286] <7-3.B> The source pixel group Gk and the destination pixel group Hk can be, for example, circular or polygonal shapes.
[0287] It goes without saying that all or part of each of the above embodiments and various modifications can be combined as appropriate and in a non-contradictory manner. [Explanation of symbols]
[0288] F1 Image 1 F2 2nd image F3,3A,3B 3rd image G1, G2, G3, G4, Ga , Gb, Gc, Gd, Gk moving source pixel group H1,H2,H3,H4,Ha,Hb,Hk Destination pixel group Image of a teacher M1,M2,M3,M4,Ma,Mb,Mc,Md,Mk Quantity R1 1st area R2 2nd area R3 3rd area R4,R41,R42,R43,R44,R45,R46,R4z 4th area R5,R51,R52,R53,R54,R55,R56 5th area S1, S2, S3, S4, S5, S30, S31, S32, S33, S34, S35a, S35b, S35c, S35d, S36, S36a, S36b, S37, S38, S39, S41, S42, S43, S301, S302, S303, S304, S305, S309, S351, S351a, S351b, S352, S352a, S352b, S353, S353a, S353b, S353c, S353d ,S 354, S354a, S354b, S355, S355a, S355b, S356, S356a, S356b, S356c, S357a, S357b, S358a, S358b, S359a, S359b, S360 step T40, T41, T42, T43, T44, V41, V42, V43, V46, V47, V48, V49, V4a, V4a1, V4a2 divided figures V5a, V5a1, V5a2 divided areas.
Claims
1. A transformation step of applying a first affine transformation to a first image having multiple pixels and occupying a first region to obtain a second image having the same multiple pixels and occupying a second region that is congruent to and inconsistent with the first region, A moving step to obtain a third image by moving or duplicating the plurality of pixels in the second image to the first region. Equipped with, The second region is divided into a third region and a fourth region, the third region is located within the first region, and the fourth region is located outside the first region. The first region is divided into the third region and the fifth region, and the fifth region is located outside the second region. The aforementioned transfer process is, A first step of setting a first source pixel group which is a group of pixels in which a predetermined first number greater than 1 of the plurality of pixels included in the fourth region are linked together, A second step involves applying a second affine transformation to the first source pixel group to determine the first destination pixel group to be arranged within the fifth region. It has, An image processing method in which the first affine transformation is either a rotation or a translation, or both, and the second affine transformation is either a rotation or a translation, or a mirror transformation.
2. The image processing method according to claim 1, wherein the first group of source pixels is inscribed in the fourth region, and the first group of destination pixels is inscribed in the fifth region.
3. The aforementioned transfer process is, A third step of setting a second group of source pixels, which is a group of pixels in which a predetermined second number of the plurality of pixels included in the fourth region are linked together, A fourth step involves applying a third affine transformation to the second source pixel group to determine the second destination pixel group to be arranged within the fifth region. It has, The second number is less than or equal to the first number, The image processing method according to claim 1 or 2, wherein the third affine transformation is either a rotation or a translation, or both, or a mirror transformation.
4. The image processing method according to claim 3, wherein the second number is 1, and the third and fourth steps are repeatedly performed until all of the plurality of pixels in the second image are moved to or duplicated in the first region.
5. The image processing method according to claim 3, wherein the first and second steps are performed multiple times prior to the third step.
6. The image processing method according to claim 1, wherein a segmented figure, which is enlarged by connecting pixels within the fourth region starting from a position away from the third region, is set as the first group of source pixels for movement.
7. The image processing method according to claim 1 or claim 6, wherein the first group of moving pixels is arranged in a segmented region that is expanded by connecting pixels within the fifth region, starting from a position away from the third region.
8. The image processing method according to claim 1 or claim 2, wherein the fourth region and the fifth region are mirror images of each other, and the second affine transform is a mirror transformation.
9. The image processing method according to claim 3, wherein the fourth region and the fifth region are mirror images of each other, and the third affine transform is a mirror transformation.
10. A method for generating training data to be used for machine learning to recognize objects in images, A step of adopting a plurality of the third images obtained by the image processing method described in claim 1 as training data, Prior to the image processing method, a step of capturing an image of the object and setting a first image, In obtaining each of the third images, the process involves extracting a region in either the first or second image from the pixel group located in the third region, specifically the pixel group corresponding to the object. A method for generating training data, comprising the following features.
Citation Information
Patent Citations
Teacher data generation device, teacher data generation method and computer program
JP2022137611A
Interactive system for automatically synthesizing a content-aware fill
US20190287224A1
Learning data creation system and learning data creation method
WO2021176605A1