Training data generation system, machine learning model training method, and training data generation method

The image synthesis system addresses the challenge of generating learning data that accurately represents environmental conditions by combining background, object, and occlusion images using synthesis conditions linked to environmental information, resulting in improved learning effects and model performance.

JP7674211B2Active Publication Date: 2025-05-09HITACHI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021156151
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-24
Publication Date
2025-05-09
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

Existing methods for generating learning data for machine learning models, particularly in deep learning, face challenges in efficiently creating data that accurately represents various environmental conditions, leading to reduced performance of machine learning models due to occlusions and environmental variations.

Method used

An image synthesis system that generates composite images by combining background images, object images, and occlusion images using synthesis conditions linked to environmental information, thereby creating learning data that is adapted to specific environments and accounts for occlusions.

Benefits of technology

The proposed solution enables the easy generation of learning data with improved learning effects, taking into account environmental conditions and occlusions, which enhances the performance and accuracy of machine learning models in image analysis tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007674211000001
    Figure 0007674211000001
  • Figure 0007674211000002
    Figure 0007674211000002
  • Figure 0007674211000003
    Figure 0007674211000003
Patent Text Reader

Abstract

To provide a learning data generation system which easily generates learning data that improves a learning effect in consideration of an environment, a learning method of a machine learning model and a learning data generation method.SOLUTION: A machine learning system 1 includes an image composition unit. The image composition unit generates a composite image from a first image, a second image and a background image by using a composition condition with the background image, the first image, correct answer information linked to the first image, the second image and the composition condition linked to environmental information as input, and generates learning data from the composite image and the correct answer information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to machine learning and a technique for generating training data (teacher data) used in machine learning, and more particularly to a technique for easily generating training data suitable for use in deep learning (DL). [Background technology]

[0002] In recent years, technology that uses deep learning to generate a machine learning model and perform processes such as image analysis, image generation, voice recognition, language processing, etc. As the machine learning model, for example, a well-known configuration such as a deep neural network (DNN) is used.

[0003] Taking image analysis as an example, known processes include classification, object detection, skeleton detection, and segmentation. Classification refers to labeling objects in an image. Object detection refers to labeling objects in an image and detecting their positions (XY coordinates). Skeleton detection refers to labeling specific parts of an object in an image and detecting its position. Segmentation refers to assigning label information to each pixel in an image and color-coding the areas.

[0004] To generate a machine learning model that performs processes such as classification, object detection, skeleton detection, and segmentation, it is necessary to train the model using training data. Conventionally, training a detection model for object detection required training data consisting of images and correct labels. In this case, if correct labels were assigned manually to each piece of data, the increase in creation costs would be a problem, so a method for automatically generating training data was desired.

[0005] Patent Document 1 relates in particular to a technology for generating learning data to be used in deep learning, and discloses a learning data generation device that includes an area determination unit that determines a target area on a background image, an image synthesis unit that generates a synthetic image by pasting an image of an object in the target area, and a correct label creation unit that creates a correct label for the synthetic image based on data related to the image of the object. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] JP 2020-149086 A Summary of the Invention [Problem to be solved by the invention]

[0007] In the invention of Patent Document 1, learning data is automatically generated according to a generation instruction for learning data. In addition, correct answer labels are added to all objects on the image of the learning data. Therefore, learning data to which appropriate correct answer labels are added can be easily generated.

[0008] However, in actual image analysis, occlusion may occur in an image of a target object. Occlusion generally refers to a state in which an object in the foreground hides an object behind it. Occlusion affects the results of image analysis, but the pattern of occurrence generally depends on the environment in which the target object is placed. Here, the environment includes various conditions of the space in which the target object is placed, such as the size, the type, number, and size of the target object and objects other than the target object, the relationship between the target object and objects other than the target object, the background, brightness, and other conditions.

[0009] In this regard, in Patent Document 1, the synthesis conditions are determined according to the conditions specified by the operator, but it is time-consuming to determine the cutout images and synthesis conditions for each environment. That is, it is necessary to select a combination of background, object, and other than object (hereinafter, for convenience, these may be called "parts") so as to approximate the actual image in accordance with the actual environment. It is also necessary to determine synthesis conditions for synthesizing parts so as to approximate the actual image. It is a burden to manually determine such parts and synthesis conditions, and if they are determined randomly, the learning data will not match the actual environment, and there is a risk of the performance of the machine learning model decreasing.

[0010] Therefore, an object of the present invention is to easily generate learning data that takes the environment into account and improves the learning effect. [Means for solving the problem]

[0011] A preferred aspect of the present invention is a learning data generation system that includes an image synthesis unit that receives as input a background image, a first image and correct answer information linked to the first image, a second image, and synthesis conditions linked to environmental information, generates a synthetic image from the first image, the second image, and the background image using the synthesis conditions, and generates learning data from the synthetic image and the correct answer information.

[0012] Another preferred aspect of the present invention is a method for learning a machine learning model, comprising: performing machine learning on a machine learning model using learning data generated by the above-described learning data generation system.

[0013] Another preferred aspect of the present invention is a learning data generation method that executes a first step of preparing synthesis conditions linked to environmental information; a second step of retrieving a background image, an object image, and a non-object image extracted from image data; a third step of generating a synthetic image by synthesizing the background image, the object image, and the non-object image using synthesis conditions linked to environmental conditions that match the image data; and a fourth step of generating learning data using the synthetic image and answer information attached to the object image. Effect of the Invention

[0014] According to the present invention, it is possible to easily generate learning data with improved learning effect that takes the environment into consideration. Problems, configurations and effects other than those described above will become apparent from the following description of the embodiments. [Brief description of the drawings]

[0015] [Figure 1] FIG. 1 is an overall block diagram of a machine learning system according to an embodiment. [Diagram 2] FIG. 4 is a conceptual diagram showing the process of a preprocessing unit. [Diagram 3] 1A and 1B are conceptual diagrams illustrating images generated from a still image. [Figure 4] FIG. 13 is a schematic diagram showing an example of a GUI displayed when performing preprocessing. [Diagram 5] FIG. 4 is a conceptual diagram showing the process of an image synthesis unit. [Figure 6] FIG. 13 is a schematic diagram showing an example of a GUI displayed when editing synthesis conditions. [Figure 7] 11 is a conceptual diagram showing an example in which environmental information of video data is inherited by related background images, object images, correct answer information, occlusion images, synthesis conditions, and learning data. [Figure 8] FIG. 13 is an overall block diagram of a machine learning system according to another embodiment. [Figure 9] 1 is a table illustrating synthesis conditions classified in a hierarchical structure. [Figure 10] FIG. 13 is an overall block diagram of a machine learning system according to another embodiment. [Figure 11]FIG. 1 is a conceptual diagram illustrating a process for automatically estimating environmental information. [Figure 12] FIG. 13 is a schematic diagram showing an example of a GUI displayed when extracting environmental information. [Figure 13] A conceptual diagram of a process for automatically selecting synthesis conditions appropriate for the environment. [Figure 14] FIG. 13 is an explanatory diagram of a GUI for creating an object with correct answer information in the case of skeleton detection. [Figure 15] FIG. 13 is an explanatory diagram of a GUI for image synthesis. [Figure 16] FIG. 11 is an explanatory diagram showing the concept of segmentation processing according to another embodiment. [Figure 17] FIG. 1 is an explanatory diagram showing the concept of image synthesis of learning data capable of adapting to environmental changes. [Figure 18] 11A and 11B are conceptual diagrams illustrating the effect when correct answer information is also added to an occlusion image. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0016] Hereinafter, examples of the present invention will be described in detail with reference to the drawings. Note that the following description is for explaining one embodiment of the present invention and does not limit the scope of the present invention. Therefore, a person skilled in the art can adopt an embodiment in which each or all of these elements are replaced with an equivalent, and such an embodiment is also included in the scope of the present invention.

[0017] In the configurations of the embodiments described below, the same parts or parts having similar functions are denoted by the same reference numerals in different drawings, and duplicated explanations may be omitted.

[0018] When there are multiple elements having the same or similar functions, they may be described by using the same reference numerals with different subscripts. However, when there is no need to distinguish between multiple elements, the subscripts may be omitted.

[0019] The designations "first," "second," "third," and the like in this specification are given to identify components, and do not necessarily limit the number, order, or content. Furthermore, numbers for identifying components are used in different contexts, and a number used in one context does not necessarily indicate the same configuration in another context. Furthermore, a component identified by a certain number does not prevent the component from also having the function of a component identified by another number.

[0020] In order to facilitate understanding of the invention, the position, size, shape, range, etc. of each component shown in the drawings, etc. may not represent the actual position, size, shape, range, etc. Therefore, the present invention is not necessarily limited to the position, size, shape, range, etc. disclosed in the drawings, etc.

[0021] All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety.

[0022] As used herein, elements referred to in the singular are intended to include the plural unless the context clearly indicates otherwise.

[0023] An example of a learning data generation system described in the embodiment is one that automatically generates learning data for a machine learning model for performing image analysis. This system generates learning data adapted to the environment in which the object to be analyzed is placed. Here, the learning data refers to data consisting of a pair of a question and an answer that is used to train the machine learning model. In the case of image analysis, the question is generally image data.

[0024] In order to improve the analysis accuracy of machine learning models, learning data that includes obstacles (occlusions) that can cause accuracy to deteriorate is necessary. Specific examples of occlusions include glasses, masks, and hats worn by the target person. Another example would be a person looking into a magnifying glass. Another example would be a situation where multiple people overlap. However, the patterns of occlusion occurrence vary depending on the conditions in each environment, and differences in each environment must be taken into consideration.

[0025] In one example of the system configuration of the embodiment, a background image, an object image, and an image of an object other than the object are extracted from an image of an object placed in a specified environment, and correct answer information is linked to the object image. A composite image is then generated from the background image, the object image, and the image of the object other than the object, and learning data is generated from the composite image and the correct answer information linked to the object image. Here, the object is the thing that is the subject of analysis in the case of image analysis.

[0026] When generating a composite image, the images are composited using a composite condition for an object image associated with a predetermined environment and an image of an object other than the object.

[0027] As already mentioned, the environment refers to various conditions in the space in which the object is placed. Environmental information is information that expresses a circle, and in the system, for example, the space in which the object is placed can be specified by an environmental information tag that is classified and defined by the user. Alternatively, the environmental information can be specified by treating the space in which the object is placed as image data and classifying and defining the device that acquired the image data (e.g., a video camera). Alternatively, the environmental information can be specified by treating the space in which the object is placed as image data and classifying or grouping the feature quantities obtained from the image data. EXAMPLES

[0028] <1. Overall configuration of machine learning system> FIG. 1 is an overall block diagram of a machine learning system according to an embodiment. In the following, an example is shown in which a system is configured to generate learning data and perform learning using a single information processing device. However, any part may be a separate information processing device connected via a network. Alternatively, the parts do not need to be connected via a network as long as data can be sent using a portable recording medium. In short, it is sufficient that necessary data can be moved between each element.

[0029] The machine learning system 1 can be configured with an information processing device such as a general server. Like a general information processing device, the machine learning system 1 includes a central processing unit (CPU) 101, an input / output interface (I / O) 102, a memory 200, and a database (DB) 300. Buses connecting each part are omitted.

[0030] A variety of known input / output devices, such as a keyboard and a display, can be connected to the input / output interface 102. An interface to a network for connecting to other information processing devices may also be provided.

[0031] The memory 200 generally uses a relatively high-speed storage device such as a semiconductor memory, and stores the programs executed by the CPU 101. The database 300 uses a relatively large-capacity storage device such as a magnetic disk device, and mainly stores data. However, the division of roles between the memory 200 and the database 300 is not particularly limited, and varies depending on the capacity of the storage device used.

[0032] In this embodiment, a preprocessing unit 201, an image synthesis unit 202, and a machine learning unit 203 are stored as programs in the memory 200. In this embodiment, a predetermined function is realized by the CPU 101 executing the programs stored in the memory 200. Note that, as long as the same function can be realized, dedicated hardware may be used instead of software.

[0033] The database 300 stores video data 301, background images 302, object images 303, correct answer information 304, occlusion images 305, and learning data 306. The role of each part will be explained later.

[0034] An example of a specific process of the machine learning system 1 will be described below. Here, for example, skeleton detection will be described as an example in which a person is extracted as an object from a video image, and the coordinates of any part of the person, such as joints, eyes, and ears, are extracted and labeled. Of course, the embodiment can be applied to other image analysis processes. When performing image analysis such as skeleton detection using a machine learning model, the quality and quantity of learning data for learning the machine learning model are important. In the following embodiment, an example of easily generating high-quality learning data will be described.

[0035] <2. Pre-processing section> 2 is a conceptual diagram showing the processing of the preprocessing unit 201. When the preprocessing unit 201 receives a command instructing preprocessing via, for example, the input / output interface 102, it starts the following processing.

[0036] The pre-processing unit 201 first reads out video data 301 from the database 300. The video data 301 is acquired, for example, from a video camera that captures an object.

[0037] The pre-processing unit 201 reads the video data 301 and extracts, for example, one frame from the video data 301 to obtain a still image (S2001). Multiple still images can be obtained from the video data 301, and the number of still images to be obtained is arbitrary, but the following description will be given of the processing for one still image.

[0038] The pre-processing unit 201 generates a background image 302 without an object (person) or occlusion, a background image 2100 without an object (person) and including occlusion, and a background image 2200 without occlusion and including an object (person). Note that although occlusion generally refers to an object that overlaps with an object, in this specification it is a general concept that refers to any object specified other than the object.

[0039] 3 is a diagram for explaining the concept of each image generated from a still image 2000. A background image 2100 including occlusion includes an occlusion image (here, a magnifying glass) 305 in addition to the background (here, a mountain), and a background image 2200 including an object includes an object image 303 (here, a person). In the simplest example of such processing, a person (user) uses an appropriate GUI (Graphical User Interface) to specify objects and occlusions while viewing a still image 2000, and separates or erases them from the background. Alternatively, they may be automatically separated using a known image recognition technique. If the objects and occlusions are predefined, they can be extracted with a certain degree of accuracy. Alternatively, they may be automatically extracted, and then the user may check and correct them.

[0040] Next, the pre-processing unit 201 cuts out the occlusion image 305 and the object image 303 from the background image 2100 including the occlusion and the background image 2200 including the object (person). To do this, it is sufficient to simply take the difference from the background image 302. Note that, although there is one type of object and occlusion in the above example, there may be multiple types.

[0041] Next, the pre-processing unit 201 assigns an individual correct answer value to the object image 303 (S2003). In the most common method, the user assigns the correct answer using an appropriate GUI. In the case of skeleton detection, parts such as the eyes, elbow joints, and wrists are specified on the image of the object image 303 and their coordinates are recorded. In this way, correct answer information 304 is obtained. In the case of skeleton detection, the correct answer information 304 includes information for identifying the object, for example, a "person", the object image 303, and the coordinates of each part on the object image. In the case of classification, the set of information for identifying the object and the object image 303 is regarded as the correct answer information 304. In the case of segmentation, the set of the segmented area and the category classification is regarded as the correct answer information 304. In the case of other image analyses, known correct answer information corresponding to the analysis is used.

[0042] FIG. 4 shows an example of a GUI that is displayed on a display via the input / output interface 102 when carrying out these pre-processing operations.

[0043] The pre-processing unit 201 records the background image 302, the object image 303, the correct answer information 304, and the occlusion image 305 created as described above in the database 300. The object image 303 and the correct answer information 304 are paired. At this time, an environmental information tag indicating environmental information is added to these generated data. The environmental information may be arbitrarily set by the user and added.

[0044] It is also possible to automatically add environmental information. For example, a tag indicating the environment in which the video camera that captured the video data 301 was placed (or an ID that identifies the video camera) may be automatically added to the background image 302, the object image 303, the correct answer information 304, and the occlusion image 305 generated based on the video data 301.

[0045] For example, if the environmental information of the video camera that shot the video data 301 is "first factory, front-end process room, first camera," the environmental information is tagged as "first factory," or "first factory, front-end process room," or "first factory, front-end process room, first camera." In other words, the environmental information can be defined arbitrarily. The environmental information may be hierarchically structured as described above.

[0046] Generally, multiple types of still images 2000 can be obtained from video data 301, and multiple types of background images 302, object images 303, correct answer information 304, and occlusion images 305 can also be obtained, so these multiple images and information are linked to common environmental information and grouped. Through the above processing, parts for synthesizing learning data are classified by environmental information and prepared in database 300.

[0047] <3. Image synthesis section> 5 is a conceptual diagram showing the process of the image synthesis unit 202. When the image synthesis unit 202 receives a command instructing image synthesis via, for example, the input / output interface 102, it starts the following process.

[0048] The image synthesis unit 202 reads out a background image 302, an object image 303, correct answer information 304, and an occlusion image 305 from the database 300. At this time, for example, it is assumed that an image of a video camera installed in "first factory, first work room" is used to perform image analysis in a machine learning model. In this case, in order to generate learning data to be used for learning the machine learning model, for example, environmental information such as "first factory, pre-process room" is specified via the input / output interface 102. This makes it possible to call up data tagged with environmental information such as "first factory, pre-process room, first camera" or "first factory, pre-process room, second camera." Alternatively, when environmental information such as "first factory, pre-process room, first camera" is specified, data tagged with environmental information such as "first factory, pre-process room, first camera" can be called up.

[0049] Next, the image synthesis unit 202 synthesizes the background image 302, the object image 303, the correct answer information 304, and the occlusion image 305 to generate a synthetic image in which the object and the occlusion are pasted onto the background image. At this time, the user can specify synthesis conditions 501 via the input / output interface 102, for example.

[0050] Here, the synthesis condition 501 can be a template including the following information: -Type and number of objects and occlusions to be composited Specifying the position (coordinates) of an object or occlusion relative to the background image Variation of object and occlusion from a specified position (coordinate) relative to the background image Layering of objects and occlusions - Size of objects and occlusions - Orientation (angle) of object or occlusion -Brightness of the composite image Composite image resolution Special effects on composite images

[0051] In this embodiment, the synthesis conditions that have been specified once are saved as a template so that they can be reused. Specifically, for example, when data tagged with the environmental information of "First factory, front-end process room" is called up to synthesize an image, a user who is familiar with the environment of "First factory, front-end process room" can specify the synthesis conditions. The specified synthesis conditions are saved as a template with the environmental information tag of "First factory, front-end process room" attached.

[0052] For example, as synthesis conditions 501 when generating a composite image using the background image 302, object image (person) 303, and occlusion image (loupe) 305 shown in Fig. 3, the order of overlay is specified, for example, background image 302 as layer 1 (bottom layer), object image 303 as layer 2, and occlusion image 305 as layer 3 (top layer). In addition, the coordinates at which the object image 303 is placed are (Xmin, Ymin), and the coordinates at which the occlusion image 305 is placed are (X'min, Y'min), with variations of (δX, δY) and (δX', δY'), respectively, are set. The variation condition may be set by specifying the area to be pasted and randomly placing the images within that area.

[0053] 6 shows an example of a GUI that is displayed on the display via the input / output interface 102 when the user edits the synthesis conditions. The user can set the synthesis conditions by moving the object image 303 and the occlusion image 305 with the correct answer information 304 on the screen and specifying the order of overlay. Note that these parts (background image, object image, occlusion image) can be used in a mixed manner, generated from multiple different still images with the same environment information tag.

[0054] Once such synthesis conditions 501 are set as a template, they are recorded in the memory of the image synthesis unit 202 with an environmental information tag indicating environmental information, such as "first factory, front-end process room." In this way, when synthesizing images using a background image 302, an object image 303, correct answer information 304, and an occlusion image 305 having the same or similar environmental information (for example, "first factory, front-end process room" or "second factory, front-end process room"), it is highly possible that the same synthesis conditions can be reused. Note that when the synthesis conditions 501 are stored as a template, the learning accuracy and efficiency of learning data generated using the synthesis conditions may be evaluated using real images (not synthetic images) to select the synthesis conditions to be saved.

[0055] In order to enable reuse of the composition conditions, it is desirable to link the generated composition conditions to environmental information as described above. The environmental information to be linked may be set arbitrarily by the user, or the same environmental information as that of the parts, etc., whose images were composed using the composition conditions may be automatically set. In the above example, the composition conditions are automatically given the same environmental information tag of "First factory, front-end process room" as the part to which the composition conditions are applied.

[0056] Thus, image synthesis unit 202 generates a pair of synthetic image 502 and correct label 503. In this example of skeleton detection, correct label 503 is the position of object image 303 in the synthetic image and the position coordinates of parts within it, such as the eye, elbow joint, and wrist. When synthetic image 502 is generated, object image 303 is specified, and the position where the object image is pasted is known, so that correct label 503 can be generated from this and correct information 304. A pair of synthetic image 502 and correct label 503 constitutes one piece of learning data 306.

[0057] The image synthesis unit 202 records the generated learning data 306 in the database 300. When recording, the same environmental information as the parts and synthesis conditions used when generating the learning data may be added. In this way, it is possible to identify the data for the learning model to be used in what environment. That is, in one example, the parts, synthesis conditions, and the created learning data are linked by the environmental information. Instead of tagging each data, the data may be stored in different databases or files for each environmental condition.

[0058] FIG. 7 shows an example in which the environmental information of video data 301 is inherited by the related still image 2000, background image 302, object image 303, correct answer information 304, occlusion image 305, synthesis condition 501, and learning data 306.

[0059] From video data 301 identified by the environmental information tag A, multiple types of still images identified by the environmental information tag A are obtained, and from these, using the preprocessing unit 201, multiple types of background images 302, object images 303, correct answer information 304, and occlusion images 305 identified by the environmental information tag A are obtained. When synthesizing these to generate a composite image, common synthesis conditions identified by the environmental information tag A can be applied.

[0060] That is, after the image synthesis unit 202 generates synthesis conditions for synthesizing a part having an environmental information tag A, the generated synthesis conditions 501 identified by the environmental information tag A can be reused to synthesize parts having the same environmental information tag A. Alternatively, if the environmental information is classified in advance as shown in Fig. 9 later, it can be reused to synthesize parts having similar environmental information tags.

[0061] Learning data 306 consisting of a set of a synthetic image 502 and a correct answer label 503 can generate many types of learning data with different positional relationships between objects according to the above-mentioned Xδ and Yδ variation conditions. Also, by randomly selecting parts from multiple types of object images 303 and occlusion images 305, many types of learning data with different states of the same object can be generated. This is training data with the synthetic image 502 as the question and the correct answer label 503 as the answer, and can be used as learning data for a machine learning model.

[0062] <4. Machine Learning Department> The machine learning unit 203 performs learning of a machine learning model (not shown) using the generated learning data 306. Since a known technique can be applied to the learning process of the machine learning model itself, a description thereof will be omitted.

[0063] When performing learning, for example, by searching for tags indicating environmental information via the input / output interface 102, learning data 306 according to the environment can be selected. For example, for human skeleton detection processing in "first factory, front-end process room", it is preferable to perform machine learning using learning data tagged with the environmental information tag "first factory, front-end process room". Alternatively, even if there is no data with the same environmental information, learning data tagged with "first factory, front-end process room" may be more suitable for learning for the purpose of "second factory, front-end process room" than learning data tagged with the environmental information tag "first employee dormitory, kitchen". In this way, it is possible to select learning data that takes environmental information into consideration.

[0064] According to this embodiment, when learning data suitable for image analysis performed by a machine learning model is automatically obtained, it is possible to reuse synthesis conditions for image synthesis. EXAMPLES

[0065] Fig. 8 is an overall block diagram of a machine learning system according to another embodiment. Differences from the embodiment in Fig. 1 will be described.

[0066] After the image synthesis unit 202 creates a template of the synthesis conditions 501, or after synthesizing an image using the template, the synthesis condition management unit 204 attaches a tag indicating environmental information to the synthesis conditions 501 and stores it in the database 300 as synthesis condition data 307.

[0067] In one example, the synthesis condition management unit 204 attaches an environmental information tag indicating environmental information designated by the user via the input / output interface 102 to the synthesis condition 501 .

[0068] In another example, the composition condition management unit 204 assigns, to the composition condition, an environmental information tag indicating environmental information that was assigned to the part composed by the image composition unit 202 using the composition condition, either as is or after editing. In this way, the composition condition management unit 204 associates and classifies the environmental information and the composition condition, and stores them in the database 300 as composition condition data 307.

[0069] 9 is a diagram showing an example of synthesis condition data 307 classified by hierarchically structured environmental information. The hierarchical structure can store environmental information in a hierarchical structure, such as major classifications (e.g., "factory" and "road"), medium classifications (e.g., "front-end processing room" and "intersection"), and minor classifications (e.g., "etching" and "Nihonbashi intersection").

[0070] According to such a hierarchical structure, even if there is no synthesis condition with the same environmental information, it is possible to extract a synthesis condition with similar environmental information. A similar classification based on a hierarchical structure can also be applied to parts and learning data.

[0071] Such a hierarchical structure can be defined by the user as desired, or can be defined by the user according to the environment in which the video camera that captures the video data that is the source of the composite parts is placed, or the purpose of the image analysis.

[0072] In one example, a video camera that captures video data is assigned an ID linked to environmental information, and the video data captured by the video camera, various data generated based on the video data, and the synthesis conditions under which the data are synthesized are assigned the same ID as the video camera. In this way, a series of data is grouped by environmental information.

[0073] 9, when creating other learning data later, the user specifies environmental information via the input / output interface 102, and the synthesis condition management unit 204 identifies parts having environmental information identical or similar to the specified environmental conditions (for example, only major classifications match). Then, as synthesis conditions for synthesizing the parts, synthesis conditions having environmental information identical or similar to the specified environmental conditions are specified. The image synthesis unit 202 then uses these to generate learning data.

[0074] As a specific example, the synthesis conditions generated as synthesis conditions for a part tagged with the environmental information tag "First factory, front-end process room" are stored in database 300 as synthesis condition data 307 tagged with the environmental information tag "First factory, front-end process room", and this synthesis condition can be later reused for image synthesis of a part obtained based on video data tagged with the environmental information tag "First factory, front-end process room".

[0075] According to this embodiment, when learning data suitable for image analysis performed by a machine learning model is automatically obtained, it is possible to select appropriate synthesis conditions and perform image synthesis. EXAMPLES

[0076] Fig. 10 is an overall block diagram of a machine learning system according to another embodiment. The parts that differ from the embodiment in Fig. 8 will be described.

[0077] In the above embodiment, the environmental information can be specified by an environmental information tag defined by the user or the origin of the image data. As another example, the environmental information can be specified by extracting image features using an autoencoder or a convolution neural network (CNN).

[0078] The environmental information management unit 205 includes a feature extractor that extracts image features using an autoencoder or a CNN. The environmental information management unit 205 records the extracted features in the database 300 as environmental information data 308.

[0079] 11 is a conceptual diagram for explaining the process of automatically estimating environmental information. After the image synthesis unit 202 creates the synthesis conditions 501, the synthesis condition management unit 204 inputs at least one of the video data 301, the still image 2000, the background image 302, the object image 303, the correct answer information 304, and the occlusion image 305 (sometimes referred to as "autoencoder input" for convenience) to the autoencoder or CNN included in the environmental information management unit 205, and extracts features (S11001).

[0080] The obtained feature amount can be linked to parts, synthesis conditions, and learning data as environmental information data 308, similarly to the first and second embodiments. In synthesis condition data 307, the synthesis conditions are linked to the feature amount and stored.

[0081] As a specific example, a feature amount may be added instead of or in addition to the environmental information tag in Fig. 7. In this case, since the feature amount differs depending on the image data, it is possible to group the image data by a known method such as clustering in order to facilitate reuse.

[0082] The more information in the autoencoder input, the more detailed the features can be obtained, but this makes it difficult to understand intuitively, so for visualization purposes, it is possible to use known techniques such as principal component analysis to represent them in low dimensions. In this case, a certain range of features (feature vectors) can be linked to a group of parts, synthesis conditions, and learning data.

[0083] As a typical example, a still image 2000 is used as an autoencoder input. When the environment information management unit 205 obtains environment information (feature amount) as the autoencoder output, it links the feature amount to video data 301 that is the source of the still image. It also links the feature amount to a background image 302, an object image 303, correct answer information 304, and an occlusion image 305, which are generated from the still image 2000. It also links the feature amount to the synthesis conditions under which these parts are synthesized.

[0084] In the example of Figure 7, data originating from a single video data 301 has the same environmental information tag, so the environmental information (features) extracted collectively can be linked to data having the same environmental information tag as the still image from which the features were extracted.

[0085] The environmental information management unit 205 records the feature amount thus linked to other data in the database 300 as environmental information data 308.

[0086] 12 shows an example of a GUI that is displayed on the display via the input / output interface 102 when the user extracts features. The extracted feature components are reduced in dimension and visualized in a graph, making it easier to understand the correspondence with environmental information.

[0087] According to the present embodiment, when linking environmental information with a synthesis condition, it is possible to automatically identify the environmental information. In addition, by expressing the environmental information as a feature vector, it is possible to automatically select the synthesis condition linked with the environmental information.

[0088] 13 is a diagram showing the concept of a process for automatically selecting appropriate synthesis conditions for an environment by automatically extracting environmental conditions (feature amounts). A feature amount extractor 1301 using an autoencoder or CNN extracts feature amounts from autoencoder input and records them as environmental information data 308. At this time, the data may be grouped by clustering using a known method.

[0089] As described above, the extracted feature amount is linked to, for example, the still image 2000 that is the autoencoder input. In addition, the extracted feature amount is linked to the video data 301, background image 302, object image 303, correct answer information 304, and occlusion image 305 related to the still image 2000 that is the autoencoder input. In the case where learning data 306 has already been generated from these autoencoder inputs, the extracted feature amount is also linked to the learning data. In addition, the feature amount is also linked to the synthesis condition data 307 used when the learning data was generated. This information is stored in advance in a database.

[0090] In the example of FIG. 13, the database contains environment information data indicating environment A and environment information data indicating environment B, and each part, learning data 306, and synthesis condition data 307 are linked to each other.

[0091] When attempting to generate learning data from parts obtained from an unknown environment, the synthesis condition management unit 204 sends, for example, a still image that is the source of the parts to the environment information management unit 205 .

[0092] Since image data such as still images reflect the environment in which the parts exist, it is preferable to generate a composite image using synthesis conditions linked to environmental conditions that match the still images.Feature values ​​are used to quantify and extract the environmental conditions from the still images.

[0093] The environmental information management unit 205 extracts features from a still image using an autoencoder input in a feature extractor 1301. Then, the existing environmental information data 308 is searched to extract environmental information data having similar features and notifies the synthesis condition management unit 204.

[0094] The synthesis condition management unit 204 uses synthesis condition data 307 linked to the extracted environmental information data to perform image synthesis of parts obtained from an unknown environment, thereby enabling image synthesis that is suited to the environmental information.

[0095] In the example of Figure 13, the environmental information management unit 205 extracts features from a still image 2000 obtained from an unknown environment C using a feature extractor 1301, and compares the features with the features of environments A and B, which are existing environmental information data 308, to extract the environmental information data 308 having the closest features.

[0096] The environmental information management unit 205 notifies the extracted environmental information data 308 to the synthesis condition management unit 204. The synthesis condition management unit 204 calls up synthesis condition data linked to the notified environmental information data from the database 300. The synthesis condition management unit 204 sends the called synthesis condition data 307 to the image synthesis unit 202, and uses the template to image-synthesize the parts obtained from the unknown environment C.

[0097] According to this embodiment, the synthesis conditions that have been created once are linked to the environmental information and registered in the template, and the appropriate synthesis conditions can be selected from multiple candidates based on the difference in the environmental information. Therefore, it is possible to select appropriate synthesis conditions even for images with unknown environmental information.

[0098] The above method of estimating environmental information using features can be used in conjunction with the method of Fig. 9 in which environmental information is organized into a hierarchical structure. For example, after extracting synthesis conditions by major classification, the synthesis conditions can be further narrowed down by environmental information using features. This method is expected to further improve the accuracy of selecting synthesis conditions. EXAMPLES

[0099] In the following, an example of a GUI for adding correct answer information to a target image in the preprocessing unit 201 and a special effect that is an example of a template for synthesizing images in the image synthesis unit 202 will be described.

[0100] FIG. 14 is an explanatory diagram of a GUI when attaching correct answer information 304 to an object image 303 in the case of skeleton detection in the processing in the pre-processing unit 201. The user attaches correct answer information 304, which is a correct answer value, to one cut-out object image 303 in relative coordinates specific to the object image, to generate an object image with correct answer information. In the example of FIG. 14, the left shoulder of the person, which is the object, is located at (x 6 ,y 6 )

[0101] FIG. 15 is an explanatory diagram when the object image with correct answer information (object image 303 and correct answer information 304) generated in FIG.

[0102] Basically, when pasting the object image 303 onto the background image, the relative coordinates of the correct answer information 304 are converted into a coordinate system specific to the background image 302. In the example of FIG. 15, based on the template of the synthesis condition data 307, the pasting position of the object image 303 is expressed as (x min ,y min ) for each template. Depending on the template, a certain position variation is given randomly. In the example of Figure 14, the person's left shoulder is specified by the relative coordinate (x 6 ,y 6 ), so in the composite image composed of the background image and the person's left shoulder, the coordinates are (x min +x 6 ,y min +y 6 ) (if there is no variation).

[0103] Here, when a magnifying glass or the like is used in the occlusion image 305, it is desirable to perform coordinate shifting to make the representation more consistent with the actual situation. In the example of Fig. 15, the "magnifying glass" which is the occlusion image 305 overlaps with the "person" which is the object image 303. In the template for the synthesis process, if the object image 303 overlaps with a predetermined range based on the occlusion image 305, it is specified that a correction is performed when changing the relative coordinates of the correct answer information 304 to a coordinate system specific to the background image 302.

[0104] Specifically, in the occlusion image 305 of the "loupe," the transparent part of the lens is cut out. Here, the composition condition template is as follows: "If the occlusion image is a magnifying glass, the coordinates of the part of the object image within the frame of the magnifying glass are relative coordinates (x 1 ,y 1 ) to (x 1 +x 2 ,y 1 +y 2 )" In the example of FIG. 15, the coordinates of the left eye position are corrected to reproduce the refraction of the lens of the magnifying glass.

[0105] As described above, it is possible to add special effects such as transforming the images of designated portions of the background image 302, the object image 303, and the occlusion image 305, changing the brightness, or blurring them.

[0106] Special effects defined by such templates are useful for expressing, for example, transparent or semi-transparent occlusion, light reflection, and the like. EXAMPLES

[0107] In the following, a case will be described where segmentation is performed by image analysis, a preprocessing unit 201 assigns category classification to the segmented region to obtain correct answer information, and an image is synthesized in an image synthesis unit 202. Basically, the configuration may be the same as that of the embodiment in Fig. 8, but the configuration specific to this embodiment will be described below.

[0108] The pre-processing unit 201 displays a still image 2000 on a display via the input / output interface 102, and has the user specify a segmentation region from the still image 2000 using an appropriate GUI, and assign a category classification to the segmentation region.

[0109] 16 is a diagram for explaining the processing of this embodiment. In the preprocessing of this embodiment, the segmented regions are stored as parts in the object image 303. In addition, the category classification (table, pliers, cutter, magnifying glass, person, . . . ) associated with each part is stored in the correct answer information 304.

[0110] In the segmentation process, there is no need to distinguish between object images 303 and occlusion images 305 as in the embodiment of FIG. 1, since the segmented region can be both an object and an occlusion at the same time.

[0111] When compositing segmented regions, the image composition unit 202 specifies the pasting positions of the segmentation regions as described in the first to third embodiments, and also specifies layers for each region as composition conditions, creates templates, and stores them as part of composition condition data 307. For example, a table (layer 0), a wrench (layer 1), and a cutter (layer 2) are defined in the template. The defined templates are associated with parts in the same manner as in the embodiments already described.

[0112] When generating a composite image from each part, the image synthesis unit 202 creates the composite image based on synthesis conditions associated with the parts so that the category and drawing order defined in advance in the template match.

[0113] For example, if you specify layers as described above, a composite image will be generated in which a wrench and a cutter are placed on a table. EXAMPLES

[0114] In the following embodiment, an example of adjusting the environmental information according to a more specific work environment will be described. For example, consider the case of tracking behavior in receiving and sorting cardboard boxes in the transportation industry, or detecting objects in deviation work. In this case, the objects of image analysis are cardboard boxes and workers (people).

[0115] The problem in this case is that since the work is non-routine, there are few similar situations, making it difficult to reuse the correct answer information 304. As a result, it is possible that the following environmental changes may cause false detection (mistaking an object that is not the detection target for the target object).

[0116] Examples of environmental changes include changes in the size of the work area, the placement of cardboard boxes, the number of workers, the clothes of the workers, the position of the lighting, changes in the condition of the workers and the cardboard boxes, etc. In addition to the above, the presence of extraneous objects (e.g., stepladders, dollies, traffic cones (registered trademark), etc.) that appear in the image may also be considered.

[0117] FIG. 17 is an explanatory diagram showing the concept of image synthesis of learning data that can respond to the above-mentioned environmental changes. In the present embodiment, the object image 303 and the occlusion image 305 are all or partly given the same correct answer information (label) to images of different forms. For example, in image recognition, an object image of a stationary person and an object image of a running person are both given the same label of "worker" and treated as the same in the system. Also, an object image of a closed cardboard box and an object image of an open cardboard box are both given the same label of "cardboard box" and treated as the same in the system. In this way, object images that are the same but have different forms depending on the situation can be obtained by, for example, cutting out still images from different timings of video image data and extracting the object images.

[0118] When synthesizing images, object images are randomly selected from multiple types of object images with the same label according to the synthesis condition template. This makes it possible to generate a synthetic image that reflects environmental changes. For example, if the synthesis condition template dictates that one object image of a "worker" be selected, an object image of a stationary person and an object image of a running person are randomly selected and the images are synthesized.

[0119] In this embodiment, the correct answer information 304 is also given to the occlusion image 305. That is, correct answer values ​​are also given to objects other than the target object associated with the environment, and image synthesis is performed to generate learning data. In this way, it is possible to reduce false positives.

[0120] 18 is a conceptual diagram for explaining the effect of attaching the correct answer information 304 to the occlusion image 305 of this embodiment. The method of attaching the correct answer information 304 to the occlusion image 305 is similar to the method of attaching the correct answer information 304 to the object image 303.

[0121] A case will be described where a still image 1801 to be subjected to image analysis is analyzed by the machine learning model 1802 of this embodiment. The machine learning model 1802 of this embodiment is capable of object detection of a "color cone" which is an occlusion.

[0122] A still image 1801 includes objects "worker" and "cardboard box," and also includes a "traffic cone" that is not an object. If the machine learning model 1802 could only recognize the objects "cardboard box" and "worker," it might erroneously recognize the "traffic cone" as a "cardboard box." However, the machine learning model of this embodiment, which can also recognize the "traffic cone," can reduce erroneous recognition.

[0123] In the upper application, processing is performed on the detection targets "worker" and "cardboard box," so information on "traffic cone" is filtered from the output 1803 of the machine learning model (S1804). The filtered data is processed by the upper application to perform processing such as behavior tracking (S1805). In the case of a data set in which only "worker" and "cardboard box" are assigned correct values ​​without assigning a correct value for "traffic cone," the similarity of "traffic cone" to the feature values ​​of "worker" or "cardboard box" stored during learning is evaluated at the time of inference, and it is recognized as one of them. However, in the case of this embodiment using a data set in which correct values ​​for "traffic cone," "worker," and "cardboard box" are assigned, at the time of inference, "traffic cone" is closest to "traffic cone" among the three feature values ​​stored during learning, so that the probability of misrecognition can be expected to be reduced. Therefore, the filtering described above is effective.

[0124] According to the above-described embodiment, it is possible to reflect environmental conditions such as illuminance, type of light, size, and distance according to the environment, which may cause objects to look different even though they are the same. For example, in a factory where similar tasks are performed, workers wear similar work clothes, but the equipment, processing materials, and layout are different. In addition, it is possible to reflect differences in environmental conditions in the learning data, such as some workers standing directly in front of the camera while others stand at an angle or at the edge.

[0125] According to the above embodiment, it is possible to generate learning data with a small amount of processing, which reduces energy consumption, reduces carbon emissions, prevents global warming, and contributes to the realization of a sustainable society. [Explanation of symbols]

[0126] Machine learning system 1, central processing unit 101, input / output interface 102, memory 200, database 300

Claims

1. An image synthesis unit is provided, The image synthesis unit includes: A background image, a first image and answer information associated with the first image, a second image, and a synthesis condition associated with environmental information are input; generating a composite image from the first image, the second image, and the background image using the composite condition; generating learning data from the synthetic image and the correct answer information; The environmental information is defined in a hierarchical classification. A system for generating training data.

2. the background image, the first image, the correct answer information, and the second image are each associated with environmental information; The learning data generating system according to claim 1.

3. environmental information associated with a synthesis condition used when generating a synthetic image from the first image, the second image, and the background image, The environmental information is the same as or similar to the environmental information associated with the first image, the second image, and the background image. The learning data generating system according to claim 2.

4. A synthesis condition management unit is provided, The synthesis condition management unit The synthesis conditions are compiled into a database based on a predetermined classification. The learning data generating system according to claim 1.

5. An image synthesis unit, The image synthesis unit includes: A background image, a first image and answer information associated with the first image, a second image, and a synthesis condition associated with environmental information are input; generating a composite image from the first image, the second image, and the background image using the composite condition; generating learning data from the synthetic image and the correct answer information; The environmental information is composed of features obtained from an image. A system for generating training data.

6. An image synthesis unit is provided, The image synthesis unit includes: A background image, a first image and answer information associated with the first image, a second image, and a synthesis condition associated with environmental information are input; generating a composite image from the first image, the second image, and the background image using the composite condition; generating learning data from the synthetic image and the correct answer information; Equipped with an environmental information management department, The environmental information management unit extracting features from the images, and storing the features in a database in association with a synthesis condition under which the first image, the second image, and the background image were synthesized; A system for generating training data.

7. The environmental information management unit extracting features from the first image, the second image, and an image related to the background image obtained from an unknown environment, and comparing the extracted features with features in the database; The image synthesis unit includes: synthesizing the first image, the second image, and the background image obtained from the unknown environment using a synthesis condition associated with the feature amount extracted based on the comparison result; The learning data generating system according to claim 6.

8. The synthesis conditions are Specifying at least one of the type and number of at least one of the first image and the second image to be combined Specifying the position of at least one of the first image and the second image relative to the background image - Variation of at least one of the first image and the second image relative to the background image from a specified position. - the order in which the first image and the second image overlap The size of at least one of the first image and the second image. The orientation of at least one of the first image and the second image. -Brightness of the composite image ・Resolution of the composite image - Special effects on composite images A template including at least one of the following information: The learning data generating system according to claim 1.

9. The special effect on the composite image is an effect of transforming an image of a specified portion of at least one of the background image, the first image, and the second image. The learning data generating system according to claim 8.

10. An image synthesis unit is provided, The image synthesis unit includes: A background image, a first image and answer information associated with the first image, a second image, and a synthesis condition associated with environmental information are input; generating a composite image from the first image, the second image, and the background image using the composite condition; generating learning data from the synthetic image and the correct answer information; A pre-processing unit is provided. The pre-treatment unit includes: extracting the background image, the first image, and the second image from a still image; adding the correct answer information to the first image to generate correct answer information associated with the first image; A system for generating training data.

11. The pre-treatment unit includes: the background image, the first image, the correct answer information, and the second image are linked to environmental information linked to the still image, respectively; The learning data generating system according to claim 10.

12. A method for learning a machine learning model, comprising: performing machine learning on a machine learning model using learning data generated by the learning data generation system according to claim 1.

13. A first step of preparing synthesis conditions linked to environmental information; a second step of retrieving the background image, the object image and the non-object image extracted from the image data; a third step of generating a composite image by combining the background image, the object image, and the non-object image using a combination condition associated with an environmental condition that matches the image data; A fourth step of generating learning data using the synthetic image and answer information added to the object image; Run using features of the image data to select environmental conditions that match the image data; How to generate training data.

14. A first step of preparing synthesis conditions linked to environmental information; a second step of retrieving the background image, the object image and the non-object image extracted from the image data; a third step of generating a composite image by combining the background image, the object image, and the non-object image using a combination condition associated with an environmental condition that matches the image data; A fourth step of generating learning data using the synthetic image and answer information added to the object image; Run The environmental conditions adapted to the image data are defined in a hierarchical classification. How to generate training data.

Citation Information

Patent Citations

  • Training data generation apparatus, training data generation method, and training data generation program

    JP2020149086A

  • Information processing apparatus

    JP2021033707A

  • Generating synthetic images and / or training machine learning model(s) based on the synthetic images

    WO2020102767A1