Learning data generation device and learning data generation method
The learning data generation device synthesizes object images in train front images to create diverse and realistic dangerous scenarios, addressing imbalance and domain shift, enhancing detection accuracy and safety in train front monitoring.
Patent Information
- Application Number
- JP2024001507
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-22
AI Technical Summary
The challenge in train front monitoring using AI models is the imbalance and domain shift issues due to the scarcity of dangerous images, leading to decreased detection accuracy when applying learned models to commercial routes.
A learning data generation device and method that synthesizes object images at specific positions in train front images based on track portions to create realistic dangerous scenarios, using object detection AI to determine appropriate image processing areas and types, thereby generating diverse and relevant learning data.
This approach addresses the imbalance and domain shift issues by creating a large dataset of realistic dangerous situations, improving detection accuracy and safety in train front monitoring.
Smart Images

Figure 2025107938000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a learning data generation device and the like.
Background Art
[0002] In the field of railways, the development of a train front monitoring system for detecting obstacles on the track using an AI (Artificial Intelligence) model that determines whether a detection target object appears in a captured image in front of the train is underway. Such an AI model is generated by performing a learning process using a group of captured learning images including a captured image in which no detection target object, which is an obstacle such as a person or a car, appears, that is, a safety image of a situation to be determined as safe, and a captured image in which a detection target object appears, that is, a dangerous image of a situation to be determined as dangerous.
[0003] However, a dangerous situation is a situation that should not occur, and the possibility of occurring in the real world is extremely low. Therefore, it is extremely difficult to prepare a dangerous image. When performing a learning process with unbalanced learning data in which the proportion of dangerous images included in the images used for the learning process is extremely small, an AI model that may not be able to make a correct determination for a real problem may result. This is a problem called so-called imbalance. This problem is a general problem typified by defective product detection using an AI model, not limited to train front monitoring. As a technique for obtaining an appropriate recognition rate when the amount of teacher data is insufficient, for example, the technique of Patent Document 1 is known.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] As a method of preparing a dangerous image for train front monitoring, a method can be considered in which a situation determined to be dangerous is deliberately created using a test track and the situation is photographed. However, since the background that occupies most of the photographed image is the background of the test track, it is significantly different from the commercial route to which actual train front monitoring is applied. For this reason, there is a problem that the detection accuracy may decrease when an AI model learned using the photographed image is applied to a commercial route. This is a problem called so-called domain shift.
[0006] The problem to be solved by the present invention is to realize a technique for preparing an appropriate train front image of learning data for train front monitoring by AI.
Means for Solving the Problem
[0007] A first invention for solving the above problems is A learning data generation device for generating learning data for train front monitoring by AI (Artificial Intelligence), A perspective degree determination means (for example, the perspective degree determination unit 204 in FIG. 5) that determines the perspective degree of a given synthesis position in the train front image based on the track portion shown in the given train front image, An image synthesis means (for example, the image synthesis unit 212 in FIG. 5) that generates a train front image of the learning data to be determined as a dangerous situation in the train front monitoring by synthesizing a given object image of a size corresponding to the perspective degree at the synthesis position, A learning data generation device comprising.
[0008] As another invention, A learning data generation method for a computer to generate learning data for train front monitoring by AI (Artificial Intelligence), Determining the perspective degree of a given synthesis position in the train front image based on the track portion shown in the given train front image (for example, step S7 in FIG. 4), By synthesizing a given object image of a size corresponding to the degree of proximity at the synthesis position, a forward train image of the learning data to be determined as a dangerous situation in the forward train monitoring is generated (for example, step S19 in FIG. 4). A learning data generation method including this may be configured.
[0009] According to the first invention or the like, it becomes possible to prepare an appropriate forward train image of the learning data for forward train monitoring by AI. That is, a dangerous image, which is a forward train image of the learning data to be determined as a dangerous situation where there is some object in front of the train, can be generated by a simple method of synthesizing a given object image at a given synthesis position in the forward train image.
[0010] Since the forward train image is an image of the three-dimensional space in front of the train as a two-dimensional space, by determining the degree of proximity of the synthesis position and synthesizing an object image of a size corresponding to that degree of proximity, a forward train image of the learning data with less discomfort that conforms to the depth in front of the train can be generated. The track laid so as to extend in front of the train always appears in the forward train image. Therefore, it is possible to appropriately determine the degree of proximity of the synthesis position based on the track portion appearing in the forward train image.
[0011] Also, it is relatively easy to photograph and acquire a large number of forward train images at various locations on the operating line. By using these forward train images to generate a forward train image of the learning data, a large number of forward train images, which are the learning data to be determined as dangerous situations at various locations on the operating line, can be easily generated.
[0012] Since a dangerous image that seems to be a dangerous situation on the operating line can be easily generated, it becomes possible to solve the problems of imbalance and domain shift related to the learning data for forward train monitoring.
[0013] The second invention is the above invention, in which Of the images of the front of the train, an image portion corresponding to a range of obstacles or a range near obstacles in train operation set based on the track is determined based on the track portion, and an image processing area is set that includes the composite position and at least a part of which overlaps with the image portion. The image processing area setting means (for example, the image processing area setting unit 206 in FIG. 5), further comprises, The image synthesizing means synthesizes the object image in the image processing area. It is a learning data generation device.
[0014] According to the second invention, a train image of appropriate learning data for determining a dangerous situation, such as the presence of an object in a range of obstacles or a range near obstacles in train operation, can be generated.
[0015] The third invention is the above-mentioned invention, Based on the detection result obtained by inputting a composite partial image in which the object image is synthesized into a partial image of the image processing area to an object detection AI learned based on a captured image of the object of the object image, an object image determination means for determining whether to adopt the object image (for example, the object image determination unit 210 in FIG. 5), further comprises, The image synthesizing means generates, as the train front image of the learning data, an image in which the object image determined to be adopted by the object image determination means is synthesized in the image processing area. It is a learning data generation device.
[0016] According to the third invention, for a composite partial image in which an object image is synthesized into a partial image of an image processing area of a train front image, it is possible to determine whether to adopt the object image as an image to be synthesized into the train front image based on whether the object detection AI detects the object. Thereby, an image suitable as learning data for the AI used for train front monitoring can be generated.
[0017] The fourth invention is the above-mentioned invention, Object type specifying means for specifying the type of object (for example, the object type specifying unit 202 in FIG. 5), further comprising, the image processing area setting means sets the image processing area based on the size of the object of the type specified by the object type specifying means and the degree of perspective, It is a learning data generation device.
[0018] According to the fourth invention, an image with a rich variety of objects in front of the train and little sense of incongruity can be generated. Objects that may exist in front of the train as situations to be determined as dangerous include various types of objects such as people, automobiles, and bicycles.
[0019] The fifth invention is the above-mentioned invention, synthetic partial image acquisition means (for example, the synthetic partial image acquisition unit 208 in FIG. 5) that acquires a synthetic partial image in which the object image is synthesized with a partial image of the image processing area by specifying the size and position of the object to be synthesized in the image processing area for the AI for image generation, It is a learning data generation device further comprising.
[0020] According to the fifth invention, by using the AI for image generation, it becomes possible to mechanically acquire a synthetic partial image in which an object image is synthesized with a partial image of the image processing area of the image in front of the train.
[0021] The sixth invention is the above-mentioned invention, object type specifying means for specifying the type of object (for example, the object type specifying unit 202 in FIG. 5), further comprising, the object detection AI is for each type of object to be detected, the synthetic partial image acquisition means specifies an object of the type specified by the object type specifying means for the AI for image generation, the object image determination means determines whether to adopt the object image using the object detection AI corresponding to the type of object specified by the object type specifying means, It is a learning data generation device.
[0022] According to the sixth invention, it is determined whether to adopt, as an image, an object image of a specified object type synthesized into the front image of the train using an AI for object detection corresponding to the specified object type. Thereby, the detection accuracy of the object can be improved. In addition, an image more suitable as learning data for the AI can be generated.
[0023] The seventh invention is the above-described invention, in which the object includes at least a person, the types of the object include differences in the postures of people, and it is a learning data generation device.
[0024] According to the seventh invention, an image in which people in various postures exist in front of the train can be generated as the front image of the train of the learning data to be determined as a dangerous situation. Then, by using the AI learned using this learning data, the detection accuracy of people in front of the train monitoring can be improved, and further improvement in safety can be achieved.
[0025] The eighth invention is the above-described invention, in which image selection means (for example, the image selection unit 214 in FIG. 5) that selects M train front images (N > M) from N train front images based on a given dangerous situation occurrence ratio, is further provided, and the M train front images selected by the image selection means are made into processing targets of the perspective determination means and the image synthesis means, thereby generating a data set of the learning data composed of N train front images including M train front images to be determined as a dangerous situation. and it is a learning data generation device.
[0026] According to the eighth invention, it is possible to generate a dataset of learning data consisting of N front train images including M dangerous situations to be determined according to the occurrence rate of dangerous situations. That is, it is possible to generate a realistic learning dataset that reproduces the occurrence rate of actual or assumed dangerous situations. By using an AI trained with this dataset of learning data, it becomes possible to appropriately determine situations that should be determined as dangerous for front train monitoring without causing the problem of imbalance, and it becomes possible to further improve safety.
Brief Description of the Drawings
[0027]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Mode for Carrying Out the Invention
[0028] Hereinafter, preferred embodiments of the present invention will be described with reference to the drawings. Note that the form to which the present invention is applicable is not limited to the following embodiments. Also, in the description of the drawings, the same reference numerals are assigned to the same elements.
[0029] The learning data generation device of this embodiment is a device that generates learning data for train front monitoring by AI (Artificial Intelligence). Train front monitoring by AI refers to detecting the presence or absence of obstacles in front of a train using an AI model that determines whether a detection target object is captured in the input captured image (the presence or absence of the detection target object) and outputs the result. The captured image is an image captured by a camera mounted on a railway vehicle of the track in front of the railway vehicle. The detection target object is an obstacle that affects the running of the railway vehicle, for example, an object such as a person, an animal, a car, or a bicycle on or near the track.
[0030] Such an AI model is a learned AI model generated by performing a learning process on an unlearned AI model using a group of learning images, which is a set of learning images associating (corresponding) the presence or absence of a detection target object in the captured image. The learning data generation device generates a learning image associated with "detection target object present" among the learning images used for learning such an AI model, that is, a learning image that should be determined as a "dangerous" situation from the perspective of train front monitoring (hereinafter, appropriately referred to as a "dangerous image").
[0031] FIGS. 1 to 3 are diagrams for explaining the outline of the generation of learning images (dangerous images) that should be determined as dangerous situations by the learning data generation device 1. Further, FIG. 4 is a flowchart for explaining the process executed by the learning data generation device 1, which is a process of generating a data set of learning data consisting of N (N>M) learning images including M learning images (dangerous images) that should be determined as dangerous situations.
[0032] The learning data generation device 1 first selects M images (M = N × risk situation occurrence ratio) from the given N train front images 10 based on the given risk situation occurrence ratio (step S1). Here, the train front image 10 is, for example, as shown in the upper part of FIG. 1, an image of the front of the train taken by a camera mounted on an operating vehicle on the operating route to be monitored for the front of the train, and is a captured image in which the detection target object is not shown, that is, an image (safe image) to be determined as a "safe" situation. The risk situation occurrence ratio can be, for example, a ratio based on the past occurrence ratio of dangerous situations on the operating route or a ratio based on locations on the operating route where dangerous situations are likely to occur.
[0033] Then, for each of the selected M train front images (safe images) 10, an image (object image) 20 of an object that is a given detection target object (in the example of FIG. 1, a "maintenance worker") is synthesized to generate a learning image (dangerous image) 30 shown in FIG. 2. By repeating the process (steps S3 to S19), M learning images (dangerous images) 30 to be determined as dangerous situations are generated.
[0034] Specifically, the learning data generation device 1 first sets the type of object to be the detection target object (step S3). This specification may be selected according to the user's specification from among the previously prepared candidates, or may be randomly selected.
[0035] Next, an image processing area 16 in the train front image 10 is set (steps S5 to S9). The image processing area 16 is an area corresponding to the position where the detection target object is to be present. Details will be described later, but it is set based on the synthesis position 12 of the detection target object, its size, etc. Subsequently, an object image 20 corresponding to the size of the image processing area 16 in the train front image 10 is generated (step S11). The object image 20 can be generated by specifying the type, posture, size, etc. of the object using an AI for image generation. Next, a synthesized partial image 22 is generated by synthesizing the object image 20 with a partial image 18 corresponding to the image processing area 16 in the train front image 10 (step S13).
[0036] Then, according to the detection result obtained by inputting the composite partial image 22 into the object detection AI 3 corresponding to the set object type, it is determined whether to adopt the object image 20 (adoption or not) (step S15). The object detection AI 3 is an AI model generated by training an untrained AI model with an image actually capturing an object of the corresponding type as training data for each type of object. The object detection AI 3 outputs, as a detection result, the probability indicating whether an object of the corresponding type is shown in the input image. The training data generation device 1 determines whether to adopt the object image 20, for example, by comparing the probability output as the detection result from the object detection AI 3 with a predetermined threshold value.
[0037] If the object image 20 is adopted (step S17: YES), the object image 20 is synthesized into the image processing area 16 of the train front image 10 to generate a training image 30 (step S19). If the object image 20 is not adopted (step S17: NO), the process returns to step S11, and again, similarly, a new object image 20 is generated, and so on. Through the above processing, a training image (hazard image) 30, which is a train front image of the training data to be determined as a dangerous situation, is generated.
[0038] In this way, when M training images (hazard images) 30 to be determined as dangerous situations are generated, together with the remaining (N - M) train front images (safe images) 30 that were not selected, a data set of training data consisting of N training images including the M training images (hazard images) to be determined as dangerous situations is formed (step S21).
[0039] FIG. 3 is a diagram for explaining in detail the setting of the image processing area 16 (corresponding to steps S5 to S9 in FIG. 4). As shown in FIG. 3, when setting the image processing area 16, first, the composite position 12 of the detection target object is set (step S5 in FIG. 4). The composite position 12 is a position where the detection target object is to be present in the forward image 10 of the train, and it may be set according to an external input by the user, or it may be set randomly from the range of obstacles or near the range of obstacles during train operation, such as on the track or in its vicinity. Also, since it is considered inappropriate to detect an object only when it is at a position too close to the train for forward train monitoring, a position at a certain distance from the train, for example, the lower edge portion of the image, may be set as the exclusion range, and the composite position 12 may be set from outside this exclusion range.
[0040] Next, the track portion 14 in the forward image 10 of the train is extracted. The track portion 14 refers to the area between the two rails on the left and right in the image. The extraction of this track portion 14 can be performed, for example, by semantic segmentation on the forward image 10 of the train. Next, based on the track portion 14, the distance of the composite position 12 is determined (step S7 in FIG. 4). Specifically, a horizontal line 13 passing through the composite position 12 in the forward image 10 of the train is defined, and the width of the track portion 14 along this horizontal line 13, that is, the length of the track gauge in the forward image 10 of the train, is calculated as the number of pixels. Since the track gauge length is known, the length in the real space per pixel in the forward image 10 of the train can be obtained from the calculated number of pixels. This is taken as the distance of the composite position 12 in the forward image 10 of the train.
[0041] Subsequently, a rectangular area including the composite position 12 as the lower side is set as the image processing area 16 (step S9 in FIG. 4). This is because an object in contact with the ground surface is assumed to be the detection target object. Note that the shape of the image processing area 16 is not limited to a rectangular shape and may be other shapes. The size of the image processing area 16 (for example, the vertical (height) and horizontal (width) lengths) is determined according to the type of object and the distance of the composite position 12.
[0042] That is, for an object with a size (e.g., height) h in real space, the number of pixels n (= h / x) corresponding to the size in the forward train image 10 when the object is assumed to be at the composite position 12 is calculated from the length x in real space per pixel, which is the perspective degree of the composite position 12. Then, the number of pixels obtained by adding a predetermined margin to the number of pixels n is set as the length of one side (e.g., the vertical length) of the image processing area 16. The same applies to the length of the other side of the image processing area 16. Note that the aspect ratio of the image processing area 16 may be determined according to the type of object, and the length of the other side of the image processing area 16 may be determined according to this aspect ratio. For example, if the type of object is "a maintenance worker (person) standing upright", the vertical size of the image processing area 16 is determined based on the height of the person, and the horizontal size of the image processing area 16 is determined based on the aspect ratio corresponding to the upright person.
[0043] Also, it is desirable to set the image processing area 16 so that it at least partially overlaps with the range of obstacles or the vicinity of obstacles in train operation, such as on or near the track. For example, if the composite position 12 is within the range of obstacles or the vicinity of obstacles in train operation, the image processing area 16 may be set such that the approximate center of the lower side is the composite position 12. If the composite position 12 is outside the range of obstacles or the vicinity of obstacles in train operation, the image processing area 16 may be set by adjusting the composite position 12 on the lower side so as to overlap with the range of obstacles or the vicinity of obstacles in train operation.
[0044] Since the forward train image 10 is a two-dimensional image of a three-dimensional space with depth, even for the same object, the size in the image may vary depending on the position where it exists. Therefore, the size of the image processing area 16 is determined according to the perspective degree of the composite position 12 in the forward train image 10, and a natural learning image (hazard image) 30 that reproduces the sense of depth of the object is generated.
[0045] [Functional Configuration] FIG. 5 shows an example of the functional configuration of the learning data generation device 1. According to FIG. 5, the learning data generation device 1 includes an operation unit 102, a display unit 104, a communication unit 106, a processing unit 200, and a storage unit 300, and is realized as a kind of computer system. Note that the learning data generation device 1 may be realized by a single computer, or may be configured by connecting a plurality of computers.
[0046] The operation unit 102 is realized by an input device such as a keyboard, a mouse, a touch panel, various switches, etc., and outputs an operation signal corresponding to the performed operation to the processing unit 200. The display unit 104 is realized by a display device such as a liquid crystal display or a touch panel, and performs various displays based on the display signal from the processing unit 200. The communication unit 106 is a communication device realized by, for example, a wireless communication module, a router, a modem, a jack of a wired communication cable, a control circuit, etc., and connects to a given communication network to perform data communication with an external device.
[0047] The processing unit 200 is a processor realized by an arithmetic device or an arithmetic circuit such as a CPU (Central Processing Unit) or an FPGA (Field Programmable Gate Array), and performs overall control of the learning data generation device 1 based on the programs and data stored in the storage unit 300, the input data from the operation unit 102 and the communication unit 106, etc.
[0048] Further, the processing unit 200 has, as functional processing blocks, an object type specifying unit 202, a distance determination unit 204, an image processing area setting unit 206, a composite partial image acquisition unit 208, an object image determination unit 210, an image composite unit 212, and an image selection unit 214. Each of these functional units of the processing unit 200 can be realized either software-wise by the processing unit 200 executing a program or by a dedicated arithmetic circuit. In this embodiment, the former software realization will be described.
[0049] The object type specifying unit 202 specifies the type of the object. The object includes at least a person, and the type of the object includes differences in the postures of the person.
[0050] Specifically, a plurality of types of objects to be detected as obstacles in the front monitoring of the train are determined in advance, and the specification can be made from among these. Also, the specification of the type of the object can be performed according to an operation input by the user via the operation unit 102, for example.
[0051] The depth determination unit 204 determines the depth of a given composite position 12 in the train front image 10 based on the track portion shown in the train front image 10 selected by the image selection unit 214.
[0052] Specifically, a track portion 14 which is the area between two left and right rails in the train front image 10 is extracted. Then, the width of the track portion 14 along the horizontal line 13 passing through the composite position 12 in the train front image 10, that is, the length of the track gauge in the train front image 10, is calculated as the number of pixels, and the length in the real space per pixel in the train front image 10 is obtained. This is taken as the depth of the composite position 12 in the train front image 10 (see FIG. 3). The composite position 12 is a position where it is desired to have a detection target object in the train front image 10, and may be set according to an external input by the user via the operation unit 102, for example, or may be set randomly from among the obstacle range or the vicinity of the obstacle range in train operation such as on or near the track.
[0053] The image processing area setting unit 206 determines, based on the track portion 14, an image portion corresponding to the obstacle range or the vicinity of the obstacle range in train operation set based on the track in the train front image 10, and sets an image processing area 16 that includes the composite position 12 and at least a part of which overlaps with the image portion. Also, the image processing area 16 is set based on the size of the object of the type specified by the object type specifying unit 202 and the depth.
[0054] Specifically, an area with a predetermined shape (e.g., rectangular shape) including the synthesis position 12 on the lower side is set as the image processing area 16. The size of the image processing area 16 (e.g., the length in the vertical (height) and horizontal (width) directions) is determined according to the type of the object and the degree of proximity of the synthesis position 12. That is, for an object with a size (e.g., height) h in the real space, from the length x in the real space per pixel, which is the degree of proximity of the synthesis position 12, the number of pixels n (= h / x) corresponding to the size in the front image 10 of the train when the object is present at the synthesis position 12 is calculated. Then, the number of pixels obtained by adding a predetermined margin to the number of pixels n is set as the length of one side (e.g., the vertical length) of the image processing area 16. The same applies to the length of the other side of the image processing area 16 (see FIG. 3).
[0055] Also, the image processing area 16 is set so as to at least partially overlap with the range of obstacles or the vicinity of obstacles during train operation, such as on or near the track. That is, if the synthesis position 12 is within the range of obstacles or the vicinity of obstacles during train operation, the image processing area 16 is set such that the approximate center of the lower side is the synthesis position 12. If the synthesis position 12 is outside the range of obstacles or the vicinity of obstacles during train operation, the synthesis position 12 on the lower side is adjusted so as to overlap with the range of obstacles or the vicinity of obstacles during train operation, and the image processing area 16 is set.
[0056] The composite partial image acquisition unit 208 obtains a composite partial image 22 in which an object image 20 is synthesized into the partial image 18 of the image processing area 16 by specifying an object of the type specified by the object type specifying unit 202 to a predetermined AI for image generation and specifying the size and position of the object to be synthesized into the image processing area 16.
[0057] Based on the captured image of the object in the object image 20, the object image determination unit 210 determines whether to adopt the object image 20 based on the detection result obtained by inputting the composite partial image 22 in which the object image 20 is combined with the partial image 18 of the image processing area 16 into the object detection AI 3 learned based on the captured image of the object. The object detection AI 3 corresponds to each type of object to be detected, and uses the object detection AI 3 corresponding to the type of object specified by the object type specifying unit 202 to determine whether to adopt the object image 20.
[0058] Specifically, the object detection AI 3 is an AI model generated by training an untrained AI model using, as training data, an image obtained by actually capturing an object of the corresponding type for each type of object. The object detection AI 3 outputs, as a detection result, the accuracy (probability) indicating whether or not an object of the corresponding type appears in the input image. The object image determination unit 210 determines whether to adopt the object image 20 by comparing, for example, the accuracy output as a detection result from the object detection AI 3 with a predetermined threshold value. The object detection AI 3 is prepared and stored in advance for each type of object as object detection AI model data 310.
[0059] The image synthesis unit 212 generates a training image (hazard image) 30, which is a train front image of the training data to be determined as a hazardous situation in train front monitoring, by synthesizing a given object image 20 of a size corresponding to the perspective at the synthesis position 12. Specifically, the object image 20 determined to be adopted by the object image determination unit 210 is synthesized into the image processing area 16.
[0060] The image selection unit 214 selects M (N > M) front - train images from N front - train images based on a given risk - situation occurrence ratio. The risk - situation occurrence ratio can be, for example, a ratio based on the occurrence ratio of past dangerous situations on the operating route, or a ratio based on locations on the operating route where dangerous situations are likely to occur. By using the M front - train images selected by the image selection unit 214 as the processing targets for each of the depth determination unit 204, the image processing area setting unit 206, and the image synthesis unit 212, a data set of learning data consisting of N front - train images including the M front - train images to be determined as dangerous situations can be generated.
[0061] The storage unit 300 is realized by an external storage device constructed in a cloud environment in addition to an IC (Integrated Circuit) memory such as a ROM (Read Only Memory) or a RAM (Random Access Memory), and a storage device such as a hard disk. It stores programs, data, etc. for the processing unit 200 to integrally control the learning - data generation device 1, is used as a working area for the processing unit 200, and temporarily stores the calculation results executed by the processing unit 200, input data from the operation unit 102 and the communication unit 106, etc.
[0062] In this embodiment, the storage unit 300 stores a learning - data generation program 302, object - detection AI model data 310, front - train - image data 320 which is data of the front - train image 10, and learning - image data 322 which is data of the generated learning image 30.
[0063] [Operation and Effect] As described above, according to this embodiment, it is possible to prepare appropriate front - train images of learning data for train - front monitoring by AI. That is, a learning image 30 which is a front - train image of learning data to be determined as a dangerous situation where there is some object in front of the train can be generated by a simple method of synthesizing a given object image 20 at a given synthesis position in the front - train image 10.
[0064] Since the forward train image 10 is an image of the three-dimensional space in front of the train as a two-dimensional space, by determining the perspective of the synthesis position 12 and synthesizing the object image 20 with a size corresponding to that perspective, it is possible to generate a learning image 30 with less discomfort that fits the depth in front of the train. The track laid so as to extend in front of the train always appears in the forward train image 10. Therefore, it is possible to appropriately determine the perspective of the synthesis position 12 based on the track portion appearing in the forward train image.
[0065] Also, it is relatively easy to capture and acquire a large number of forward train images 10 at various locations on the operating route. By using these forward train images 10 to generate the learning image 30, it is possible to easily generate a large number of learning images 30, which are learning data to be determined as dangerous situations at various locations on the operating route. Since it is possible to easily generate a dangerous image as if it were a dangerous situation on the operating route, it is possible to solve the problems of imbalance and domain shift related to the learning data for forward train monitoring.
[0066] Note that the applicable embodiments of the present invention are not limited to the above-described embodiments, and it goes without saying that they can be appropriately changed without departing from the spirit of the present invention.
Explanation of Reference Numerals
[0067] 1... Learning data generation device 200... Processing unit 202... Object type designation unit 204... Perspective determination unit 206... Image processing area setting unit 208... Composite partial image acquisition unit 210... Object image determination unit 212... Image synthesis unit 214... Image selection unit 300... Storage unit 302... Learning data generation program 310... AI model data for object detection 320... Forward train image data 322... Learning image data 10… Front image of the train (safety image) 30… Training image (hazard image) 12… Composite position 13… Horizontal line 14… Track section 16… Image processing area 18… Partial image 20… Object image 22… Composite partial image 3… AI for object detection
Claims
1. A learning data generation device for generating learning data for train front monitoring by AI (Artificial Intelligence), a perspective degree determination means for determining the perspective degree of a given composite position in the train front image based on the track portion shown in the given train front image; an image synthesis means for generating a train front image of the learning data to be determined as a dangerous situation in the train front monitoring by synthesizing a given object image of a size corresponding to the perspective degree at the composite position; A learning data generation device comprising:
2. An image processing area setting means for determining an image portion corresponding to a train operation obstacle range or an obstacle vicinity range set based on the track in the train front image based on the track portion, and setting an image processing area including the composite position and at least a part of which overlaps with the image portion; further comprising: The image synthesis means synthesizes the object image in the image processing area. The learning data generation device according to Claim 1.
3. An object image determination means for determining whether to adopt the object image based on a detection result obtained by inputting a composite partial image obtained by synthesizing the object image into a partial image of the image processing area to an object detection AI learned based on a captured image of the object of the object image; further comprising: The image synthesis means generates, as the train front image of the learning data, an image obtained by synthesizing the object image determined to be adopted by the object image determination means in the image processing area. The learning data generation device according to Claim 2.
4. An object type designating means for designating the type of the object; further comprising: The image processing area setting means sets the image processing area based on the size of the object of the type designated by the object type designating means and the perspective degree. The learning data generation device according to Claim 2.
5. A composite partial image acquisition means for acquiring a composite partial image obtained by synthesizing the object image into a partial image of the image processing area by designating the size and position of the object to be synthesized in the image processing area for an AI for image generation; The learning data generation device according to Claim 3, further comprising:
6. An object type designating means for designating the type of the object; further comprising: The object detection AI exists for each type of object to be detected. The synthetic partial image acquisition means designates an object of the type designated by the object type designation means to the AI for image generation. The object image determination means determines whether the object image is acceptable or not using the object detection AI corresponding to the type of the object designated by the object type designation means. The learning data generation device according to claim 5.
7. The object includes at least a person. The types of the object include differences in the postures of people. The learning data generation device according to claim 4.
8. Image selection means for selecting M (N > M) front train images from N front train images based on a given risk situation occurrence rate, and further comprising: generating a data set of the learning data consisting of N front train images including M front train images to be determined as dangerous situations by using the M front train images selected by the image selection means as processing targets of the depth determination means and the image synthesis means. The learning data generation device according to any one of claims 1 to 7.
9. A learning data generation method for a computer to generate learning data for train front monitoring by AI (Artificial Intelligence), determining the depth of a given synthesis position in the front train image based on the track portion shown in the given front train image; generating a front train image of the learning data to be determined as a dangerous situation in the front train monitoring by synthesizing a given object image of a size corresponding to the depth at the synthesis position; A learning data generation method including the above.
Citation Information
Patent Citations
Image recognition learning device, image recognition learning method, image recognition learning program and terminal device
JP2021002270A