Mobile body training data generation method and mobile body training data generation system
The method and system generate mobile object learning data through virtual simulation, addressing inefficiencies in existing methods by reducing labor and improving efficiency in creating annotated data for machine learning models.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
- Filing Date
- 2025-09-19
- Publication Date
- 2026-04-23
AI Technical Summary
Existing methods for generating learning data for machine learning models that detect mobile objects require significant labor and are inefficient, particularly in adding annotation information.
A method and system that utilize a simulation space to virtually place mobile objects and imaging devices, generating virtual images with annotation information to create learning data, reducing the need for manual labor and improving efficiency.
Reduces the effort required to prepare learning data and enhances the efficiency of machine learning by generating large amounts of annotated learning data without direct site visits, enabling accurate annotation even in complex scenarios.
Smart Images

Figure JP2025033059_23042026_PF_FP_ABST
Abstract
Description
Method for Generating Mobile Object Learning Data and Mobile Object Learning Data Generation System
[0001] The present disclosure relates to a method for generating mobile object learning data and a mobile object learning data generation system.
[0002] Patent Document 1 discloses an image data collection device including a distance detection unit that detects the distance to an object, an imaging unit that captures an image of the object to acquire image data, an object detection unit that detects the presence or absence of the object based on the detection value of the distance detection unit, and an image discrimination unit that discriminates the image data of the imaging unit according to the detection result of the object detection unit. This image data collection device detects the presence or absence and position of an object based on the detection value of the distance detection unit, creates label data representing the area where the object appears in the image data of the imaging unit based on the position when the object is detected, and stores the image data and the label data.
[0003] Japanese Patent Application Laid-Open No. 2021-33572
[0004] In Patent Document 1, it is assumed that the process of detecting the distance to an object and the process of imaging the object are executed at least in a real environment where the object exists. Therefore, there is a problem that a great deal of labor is required for generating and preparing the learning data necessary for performing machine learning using deep learning. In addition, in order to perform efficient machine learning, additional information called annotation is often added to the learning data, but this also requires a corresponding amount of labor, and it is considered that there is room for improvement in preparing efficient learning data.
[0005] The present disclosure has been devised in view of the above-described conventional situation, and aims to reduce the labor required for preparing the learning data necessary for machine learning of a model for detecting a mobile object and to assist in improving the efficiency of machine learning.
[0006] This disclosure provides a method for generating mobile object learning data, which involves placing a mobile object and an imaging device that virtually images the mobile object on a simulation space that reproduces the field space to be learned, moving the mobile object as seen from the imaging device on the simulation space while the imaging device virtually images the mobile object to generate a virtual image, generating annotation information that identifies the mobile object in the virtual image based on the virtual image of the mobile object, adding the annotation information to the virtual image, and generating learning data necessary for training a detection model that detects any mobile object present in the field space.
[0007] Furthermore, this disclosure provides a mobile object learning data generation system comprising a processor and memory, wherein the processor, in cooperation with the memory, places a mobile object and an imaging device that virtually images the mobile object on a simulation space that reproduces the field space to be learned, moves the mobile object as seen from the imaging device on the simulation space and virtually images the mobile object with the imaging device to generate a virtual image, generates annotation information that identifies the mobile object in the virtual image based on the virtual image of the mobile object, adds the annotation information to the virtual image, and generates learning data necessary for learning a detection model that detects any mobile object present in the field space.
[0008] These comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or recording media, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0009] According to this disclosure, it is possible to reduce the effort required to prepare training data necessary for machine learning of models that detect moving objects, and to support the efficiency of machine learning.
[0010] Figure 8 shows an example of the system configuration of the mobile object learning data generation system according to this embodiment. Figure 9 shows an example of the data structure of on-site information for on-site objects. Figure 10 shows an example of the data structure of on-site information for virtual cameras. Figure 11 shows an example of the data structure of mobile object position information. Figure 12 shows an example of the data structure of annotation information. Figure 13 shows an example of the overview of the learning image generation process. Figure 23 shows another example of a learning image. Figure 34 shows a flowchart showing an example of the operation procedure of the mobile object simulation device in chronological order. Figure 45 shows an example of the operation procedure of the annotation information generation process in chronological order. Figure 56 shows a flowchart showing an example of the operation procedure of the mobile object learning image generation device in chronological order.
[0011] Hereinafter, with appropriate reference to the drawings, embodiments specifically disclosing the mobile learning data generation method and mobile learning data generation system according to this disclosure will be described in detail. However, unnecessarily detailed explanations may be omitted. For example, detailed explanations of already well-known matters and redundant explanations of substantially identical configurations may be omitted. This is to avoid the following explanation becoming unnecessarily verbose and to facilitate understanding by those skilled in the art. The attached drawings and the following explanation are provided to enable those skilled in the art to fully understand this disclosure and are not intended to limit the subject matter of the claims. Furthermore, in the following explanation, the same elements may be assigned the same reference numerals to simplify or omit explanations.
[0012] 1. Configuration of the Mobile Object Learning Data Generation System First, the system configuration of the mobile object learning data generation system 100 according to this embodiment will be described with reference to Figure 1. Figure 1 is a diagram showing an example of the system configuration of the mobile object learning data generation system according to this embodiment. The mobile object learning data generation system 100 includes at least a mobile object image recognition learning data generation system 10, a mobile object image recognition device 40, and a field work device 50. The learning data generation system 10 and the mobile object image recognition device 40, and the mobile object image recognition device 40 and the field work device 50 are connected via a wired or wireless network so that data communication is possible between them. Wireless networks may include, for example, Wide Area Network (WAN), Local Area Network (LAN), Long Term Evolution (LTE), 4G, 5G, and other mobile communications, power line communications, short-range wireless communications (e.g., Bluetooth® communications), or networks for mobile phones. Wired networks may include, for example, wired LANs or wired WANs.
[0013] The mobile object image recognition learning data generation system 10 comprises a mobile object simulation device 20 and a mobile object learning image generation device 30. The mobile object simulation device 20 and the mobile object learning image generation device 30 are connected to each other via the wired or wireless network described above, enabling data communication between them. In the mobile object image recognition learning data generation system 10, the mobile object simulation device 20 and the mobile object learning image generation device 30 are configured as separate units, but they may also be configured as an integrated unit.
[0014] The mobile object simulation device 20 is configured using, for example, a personal computer (PC) or a server computer. The mobile object simulation device 20 virtually places a virtual camera 70 and one or more mobile objects TG1 (e.g., a vehicle, drone, person, etc.) to be detected on a computer simulation space 80 (see Figure 6) that virtually reproduces the field space to be learned, and executes a process to move the mobile objects on the computer simulation space 80. While moving the mobile objects, the mobile object simulation device 20 executes a process to generate annotation information for identifying the position of the mobile objects in the virtual captured images based on the virtual captured images virtually captured by the virtual camera 70. The mobile object simulation device 20 includes a processor 21 and a memory 22. Although not shown in Figure 1, the mobile object simulation device 20 may further include an input device such as a mouse that can accept user input.
[0015] The processor 21 is composed of at least one of the following: a Central Processing Unit (CPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), or a Graphical Processing Unit (GPU). The processor 21 functions as a controller that manages the overall operation of the mobile simulation device 20. The processor 21 performs control processing to coordinate the operation of each part of the mobile simulation device 20, data input / output processing between each part of the mobile simulation device 20, data calculation processing, and data storage processing. The processor 21 operates according to the program stored in the memory 22. When operating, the processor 21 uses the memory 22 to cooperate in executing various processes and temporarily stores data generated or acquired by the processor 22 in the memory 22. The processor 21 works in cooperation with the memory 22 to realize the functions of the mobile body simulation unit 23 and the annotation information generation unit 24.
[0016] The mobile object simulation unit 23 virtually places a virtual camera 70 and one or more mobile objects TG1 (e.g., a vehicle, drone, person, etc.) to be detected on a computer simulation space 80 (see Figure 6) that virtually reproduces the field space to be learned, and executes a process to move the mobile objects on the computer simulation space 80. The field objects virtually placed on the computer simulation space 80 are not limited to the virtual camera 70 and the mobile objects TG1, but also include various obstacles placed in the real field space that forms the basis of the computer simulation space 80. Details of such obstacles, virtual camera 70, and mobile object TG1 information (specifically, field information 25 and mobile object position information 26) will be described later with reference to Figures 2, 3, and 4, respectively. The mobile object simulation unit 23 also executes a process to generate a virtual image of the mobile object TG1 as seen from the virtual camera 70 on the computer simulation space 80 (i.e., a virtual image of the mobile object TG1 virtually captured by the virtual camera 70).
[0017] The annotation information generation unit 24 generates annotation information to identify the position of the mobile object TG1 in the virtual image, based on the virtual image of the mobile object TG1 as seen from the virtual camera 70 generated by the mobile object simulation unit 23. The annotation information is additional information for identifying the position of the mobile object TG1 in the virtual image of the virtual camera 70, and includes at least one of the following: a point indicating the position of the mobile object TG1 in the virtual image, a rectangle, an arbitrary shape outline, or text indicating the type of mobile object TG1. Details of the annotation information 27 will be described later with reference to Figure 5.
[0018] Memory 22 is configured using, for example, Random Access Memory (RAM) and Read Only Memory (ROM), and temporarily stores programs necessary for the operation of the mobile simulation device 20, as well as data acquired or generated during operation. RAM is, for example, work memory used during the operation of the mobile simulation device 20. ROM stores, for example, programs for controlling the mobile simulation device 20 in advance. Memory 22 temporarily stores field information 25, mobile location information 26, and annotation information 27 generated by the processor 21.
[0019] Here, we will explain the field information 25, mobile object position information 26, and annotation information 27 with reference to Figures 2 to 5. Figure 2 is a diagram showing an example of the data structure of field information 25a for a field object. Figure 3 is a diagram showing an example of the data structure of field information 25b for a virtual camera 70. Figure 4 is a diagram showing an example of the data structure of mobile object position information 26. Figure 5 is a diagram showing an example of the data structure of annotation information 27.
[0020] As shown in Figure 2, the field information 25a for field objects consists of a record for each field object that includes a field object ID that identifies the field object, a time, coordinates, and a CG type. To distinguish between field objects and moving objects, moving objects are treated differently from field objects, while field objects are treated as objects that do not move. The time indicates the time when the field object in question was placed in the computer simulation space 80. The coordinates indicate the three-dimensional coordinates that specify the position in which the field object in question was placed in the computer simulation space 80. The CG type indicates what kind of object the field object in question is. This field information 25a is an example of field information 25 generated by the mobile body simulation unit 23 of the processor 21 of the mobile body simulation device 20 at the timing when the object in question is placed in the computer simulation space 80.
[0021] As shown in Figure 3, the field information 25b for the virtual camera 70 consists of a record for each virtual camera 70 that includes a virtual camera ID that identifies the virtual camera 70, a time, coordinates, CG type, direction, field of view, focal length, and shooting time interval. The time indicates the time when the corresponding virtual camera 70 was placed in the computer simulation space 80. The coordinates indicate the three-dimensional coordinates that specify the position where the corresponding virtual camera 70 was placed in the computer simulation space 80. The CG type indicates what type of camera device the corresponding virtual camera is. The direction, field of view, focal length, and shooting time interval indicate the camera parameters set for the corresponding virtual camera, respectively. This field information 25b is an example of field information 25 generated by the mobile simulation unit 23 of the processor 21 of the mobile simulation device 20 at the timing when the corresponding virtual camera 70 is placed in the computer simulation space 80.
[0022] As shown in Figure 4, the mobile object position information 26 consists of a record for each mobile object that includes a mobile object ID to identify the mobile object, a time, coordinates, and a CG type. The time indicates the time when the mobile object was placed in the computer simulation space 80 or the time at a certain moment after it moved. The coordinates indicate the three-dimensional coordinates that specify the position where the mobile object was placed in the computer simulation space 80 or the position at a certain moment after it moved. The CG type indicates what kind of object the mobile object is. This mobile object position information 26 is generated by the mobile object simulation unit 23 of the processor 21 of the mobile object simulation device 20 at the time when the mobile object was placed in the computer simulation space 80 or at a certain moment after it moved.
[0023] As shown in Figure 5, the annotation information 27 consists of a record for each pair of a moving object and a virtual camera that images the moving object, which includes a virtual camera ID that identifies the virtual camera, a moving object ID that identifies the moving object, a time, an annotation type, and annotation data. Although not shown in Figure 5, the annotation information 27 may also include actual data of the virtual image of the moving object TG1 as seen by the virtual camera 70 generated by the moving object simulation unit 23, or information about the storage location of the actual data. The time indicates the type of annotation generated based on the virtual image captured by the virtual camera of the corresponding moving object. The annotation data indicates the pixel position (2D coordinate) in the virtual image that identifies the range of existence of the annotation for identifying the moving object present in the virtual image of the virtual camera 70. This annotation information 27 is generated by the annotation generation unit 24 of the processor 21 of the moving object simulation device 20 at the timing when the annotation is generated corresponding to the moving object.
[0024] The mobile object learning image generation device 30 is configured using, for example, a personal computer (PC) or a server computer. The mobile object learning image generation device 30 uses on-site information 25, mobile object position information 26, and annotation information 27 sent from the mobile object simulation device 20 to generate training images 35 necessary for machine learning to generate a machine learning model 45 for detecting mobile objects. The mobile object learning image generation device 30 includes a processor 31 and memory 32. Although not shown in Figure 1, the mobile object learning image generation device 30 may further include an input device such as a mouse that can accept user input.
[0025] The processor 31 is composed of at least one of the following: a CPU, DSP, FPGA, or GPU. The processor 31 functions as a controller that oversees the overall operation of the mobile learning image generation device 30. The processor 31 performs control processing to coordinate the operation of each part of the mobile learning image generation device 30, data input / output processing between the mobile learning image generation device 30 and each part, data calculation processing, and data storage processing. The processor 31 operates according to the program stored in the memory 32. During operation, the processor 31 uses the memory 32 to cooperate in executing various processes and temporarily stores the data generated or acquired by the processor 31 in the memory 32. By cooperating with the memory 32, the processor 31 realizes the functions of the CG image generation unit 33 and the annotation unit 34.
[0026] The CG image generation unit 33 generates a Computer Graphic (CG) image in which various objects indicated by the site information 25 and the mobile object indicated by the mobile object position information 26 are placed in the computer simulation space 80, based on the site information 25 and mobile object position information 26 sent from the mobile object simulation device 20. Since this method of generating CG images can be realized using already known publicly available techniques, a detailed explanation is omitted here.
[0027] The annotation unit 34 uses the annotation information 27 sent from the mobile simulation device 20 to generate training images 35 by adding at least one annotation indicated by the annotation information to the CG images generated by the CG image generation unit 33. The annotation unit 34 may generate training images 35 in a format that conforms to the specifications of the machine learning model 45.
[0028] Here, the training images 35 will be explained with reference to Figures 6 and 7. Figure 6 is a diagram showing an example of the overview of the training image generation process. Figure 7 is a diagram showing another example of a training image. The training images 35 are input images necessary for the machine learning process in which the mobile image recognition device 40 constructs (for example, generates or updates) a machine learning model 45. When a large number of very diverse training images 35 are prepared, the mobile image recognition device 40 performs machine learning on the machine learning model.
[0029] As shown in Figure 6, a virtual camera 70 is positioned on the upper part of a pole in the z-direction in a computer simulation space 80 illustrating a parking lot PK1, and a mobile object TG1 to be detected is positioned there. The CG image generation unit 33 of the processor 31 of the mobile object learning image generation device 30 generates a CG image IMG1 with the mobile object TG1 as the subject as seen from the virtual camera 70. In this CG image IMG1, the mobile object TG1 is shown almost in the center, and other objects around the parking lot PK1 may also be shown. The annotation unit 34 of the processor 31 of the mobile object learning image generation device 30 generates a learning image IMG2 by adding annotation information LB1 to the CG image IMG1. The annotation information LB1 of the learning image IMG2 consists of, for example, the outline of the vehicle which is the mobile object TG1 and the area enclosed by that outline filled with a predetermined color.
[0030] As shown in Figure 7, the training image IMG3 may contain multiple moving objects (e.g., vehicles) in a single image. The training image IMG3 is assigned annotation information LB11, LB12, LB13, LB14, LB15, LB16, LB17, LB18, LB19, LB20, LB21, LB22, LB23, and LB24 for each vehicle traveling in each of the multiple lanes.
[0031] Memory 32 is configured, for example, using RAM and ROM, and temporarily stores programs necessary for the operation of the mobile object learning image generation device 30, as well as data acquired or generated during operation. RAM is, for example, work memory used during the operation of the mobile object learning image generation device 30. ROM stores, for example, programs for controlling the mobile object learning image generation device 30 in advance. Memory 32 temporarily stores field information 25, mobile object position information 26, and annotation information 27 sent from the mobile object simulation device 20.
[0032] The mobile object image recognition device 40 is configured using, for example, a personal computer (PC) or a server computer. The mobile object image recognition device 40 receives and acquires a plurality of training images 35 generated by the mobile object learning image generation device 30. Using the plurality of training images 35, the mobile object image recognition device 40 performs machine learning processing to construct (e.g., generate or update) a machine learning model capable of detecting moving objects in the field image 56, which is a virtual image captured by the field work device 50. The mobile object image recognition device 40 includes a processor 41 and memory 42. Although not shown in Figure 1, the mobile object image recognition device 40 may further include an input device such as a mouse that can accept user input.
[0033] The processor 41 is composed of at least one of the following: a CPU, DSP, FPGA, or GPU. The processor 41 functions as a controller that oversees the overall operation of the mobile object learning image recognition device 40. The processor 41 performs control processing to coordinate the operation of each part of the mobile object learning image recognition device 40, data input / output processing between each part of the mobile object learning image recognition device 40, data calculation processing, and data storage processing. The processor 41 operates according to the program stored in the memory 42. When operating, the processor 41 uses the memory 42 to cooperate in executing various processes and temporarily stores data generated or acquired by the processor 41 in the memory 42. By cooperating with the memory 42, the processor 41 realizes the functions of the learning unit 43 and the inference unit 44.
[0034] The learning unit 43 uses multiple training images 35 stored in the memory 42 to perform machine learning processing to construct (e.g., generate or update) a machine learning model capable of detecting moving objects in virtual images captured by the field work device 50. The learning unit 43 stores the machine learning model 45 obtained through the machine learning processing in the memory 42.
[0035] The inference unit 44 refers to a machine learning model stored in memory 42, receives a field image 56, which is a virtual image captured at various work sites and sent from the field work device 50, detects moving objects in the field image 56, and executes a process to generate information that identifies the moving object. The inference unit 44 generates moving object position information 46 that indicates the position of the moving object identified in the field image 56 and feeds it back (transmits) it to the field work device 50. This moving object position information 46 may be configured to have the same items as the moving object position information 26 shown in Figure 4, or it may be configured to further include information indicating the position of the moving object in the virtual image captured by the moving object, the type of the moving object, or both.
[0036] Memory 42 is configured, for example, using RAM and ROM, and temporarily stores programs necessary for the operation of the mobile learning image recognition device 40, as well as data acquired or generated during operation. RAM is, for example, work memory used during the operation of the mobile learning image recognition device 40. ROM stores, for example, programs for controlling the mobile learning image recognition device 40 in advance. Memory 42 temporarily stores multiple training images 35 sent from the mobile learning image generation device 30. Memory 42 also stores machine learning models 45 constructed (for example, generated or updated) by machine learning performed by the processor 41.
[0037] The machine learning model 45 detects moving objects in the input virtual image and outputs information to identify those moving objects. The machine learning model 45 is constructed (e.g., generated or updated) by machine learning performed by the learning unit 43 of the processor 41 and stored in the memory 42.
[0038] The field work device 50 is configured using, for example, a personal computer (PC). The field work device 50 is used for each type of task performed in the field space, and generates a field image 56 by capturing images of the field space, and requests the mobile object image recognition device 40 to identify the position of a subject (for example, a moving object) that appears in the field image 56. The field work device 50 includes a processor 51 and a memory 52. Although not shown in Figure 1, the field work device 50 may further include an input device such as a mouse that can accept user input.
[0039] The processor 51 is composed of at least one of the following: a CPU, DSP, FPGA, or GPU. The processor 51 functions as a controller that oversees the overall operation of the field work device 50. The processor 51 performs control processing to coordinate the operation of each part of the field work device 50, data input / output processing between each part of the field work device 50, data calculation processing, and data storage processing. The processor 51 operates according to the program stored in the memory 52. When operating, the processor 51 uses the memory 52 to cooperate in executing various processes and temporarily stores data generated or acquired by the processor 51 in the memory 52. By cooperating with the memory 52, the processor 51 realizes the functions of the mobile information registration unit 54.
[0040] The mobile object information registration unit 54 executes the process of storing (registering) the mobile object position information 46, which has been fed back (returned) from the mobile object image recognition device 40, in the memory 52.
[0041] Memory 52 is configured using, for example, RAM and ROM, and temporarily stores programs necessary for the operation of the field work device 50, as well as data acquired or generated during operation. RAM is, for example, work memory used during the operation of the field work device 50. ROM stores, for example, programs for controlling the field work device 50 in advance. Memory 52 temporarily stores field images 56 captured by the image acquisition device 53 and mobile object position information 46 fed back from the mobile object image recognition device 40. Memory 52 also stores field master data 55, which is various data necessary for performing tasks in the field space.
[0042] The site master 55 is a database that stores various data necessary for the execution of operations in the site space handled by the site operation device 50.
[0043] The image capturing device 53 captures and generates a virtual captured image (i.e., the site image 56) showing the state of the site space handled by the site operation device 50 and sends it to the processor 51. The processor 51 generates a request for recognition processing including this site image 56 and sends it to the moving object image recognition device 40. The response (feedback) to this recognition processing request becomes the moving object position information 46.
[0044] 2. Operation of the Moving Object Learning Data Generation System Next, referring to FIGS. 8 to 10, the operation procedure of the moving object learning data generation system 100 according to the present embodiment will be described. FIG. 8 is a flowchart showing an example of the operation procedure of the moving object simulation device in time series. FIG. 9 is a flowchart showing an example of the operation procedure of the annotation information generation process of FIG. 8 in time series. FIG. 10 is a flowchart showing an example of the operation procedure of the moving object learning image generation device in time series. A series of processes in FIGS. 8 and 9 are mainly executed by the processor 21 of the moving object simulation device 20. A series of processes in FIG. 10 are mainly executed by the processor 31 of the moving object learning image generation device 30.
[0045] In FIG. 8, the processor 21 initializes the simulation process of the moving object, prepares a computer simulation space 80 that reproduces the site space to be learned, and arranges various objects and moving objects in the computer simulation space 80 (St1). In this initialization, the processor 21 sets the information necessary for the simulation process based on user operations or inputs from other systems. This necessary information includes, for example, the initial position, moving speed, moving destination of the moving object on the computer simulation space 80, the three-dimensional shape of the computer simulation space 80, the arrangement position of the passage through which the moving object can move in the computer simulation space 80, the arrangement position of obstacles, the installation position, direction, angle of view, focal length, shooting time interval, etc. of the virtual camera 70 arranged in the computer simulation space 80.
[0046] The processor 21 sets the number of annotation information to be generated for a single moving object (i.e., the number of training images to be generated) based on user input (St2). The moving object simulation device 20 repeats the generation of annotation information for a single moving object, matching the number of annotation information to be generated set in step St2 (i.e., the series of processes from steps St3 to St6). The processor 21 performs a process to move (specifically update) the moving object in the computer simulation space 80 based on time-series data 26a, among the various objects and moving objects placed in the computer simulation space 80 during initialization in step St1 (St3). This process in step St3 may be performed using any simulation method, for example, based on the occurrence of discrete movements or based on continuous time steps. The time-series data 26a is data indicating the position of the moving object at each time (specifically, three-dimensional coordinates indicating the position in the computer simulation space 80).
[0047] In step St3, the processor 21 executes a process to generate annotation information to identify the moving object while it is moving (St4). Details of the process in step St4 will be described later with reference to Figure 9. In step St3, the processor 21 outputs the current time on the computer simulation space 80, the field information 25, the moving object position information 26, and the annotation information 27 obtained while updating the position of the moving object to the moving object learning image generation device 30 (St5). If the processor 21 determines that it has generated a number of annotation information items that match the number of annotation information items set in step St2 (St6, YES), it terminates the series of processes shown in Figure 8. On the other hand, if it determines that it has not yet generated a number of annotation information items that match the number of annotation information items set in step St2 (St6, NO), the processor 21 returns to step St3.
[0048] In FIG. 9, the processor 21 uses the on-site information 25 and the moving object position information 26 to virtually capture the moving object TG1, which is the subject seen from the virtual camera 70 arranged on the computer simulation space 80, and generates a virtual captured image (see the CG image IMG1 in FIG. 6) (St11). The processor 21 uses the moving object position information 26 to associate and organize the virtual captured image generated in step St21 with the position of the moving object, and projects the portion of the virtual captured image to which the annotation of the moving object should be added (St12). The processor 21 generates annotation information 27 (see FIG. 5) according to which pixel of the virtual captured image on which the portion to be annotated in step St12 is projected and what kind of annotation of the moving object should be given (St13). Note that the processor 21 may determine what kind of annotation information to generate according to the specifications of the machine learning model used in the moving object image recognition device 40, and generate annotation information according to the determination result.
[0049] In FIG. 10, the processor 31 reads and acquires various information (specifically, the on-site information 25, the moving object position information 26, and the annotation information 27) sent from the moving object simulation device 20 from the memory 32 (St21). The processor 31 uses the on-site information 25 and the moving object position information 26 acquired in step St21 to generate a CG image corresponding to the virtual captured image when the moving object TG1, which is the subject seen from the virtual camera 70 arranged on the computer simulation space 80, is virtually captured (St22).
[0050] The processor 31 generates a learning image 35 in which an annotation using the annotation information 27 acquired in step St21 is added (for example, superimposed) to the CG image generated in step St22 (St23). The annotation using the annotation information 27 corresponds to, for example, a point indicating the position of the moving object, a rectangle, an outline of an arbitrary shape, a text for identifying the type of the moving object, or a combination thereof, in order to facilitate the identification of the moving object reflected in the CG image. The processor 31 outputs the learning image 35 generated in step St23 to the moving object image recognition device 40 (St24).
[0051] As described above, the mobile object learning data system 100 according to this embodiment allows for the generation of a large amount of learning data (learning images) in a computer simulation space 80, where various objects and mobile objects are virtually arranged to resemble the actual site space, without having to go directly to the site space to be learned. These generated learning images can then be used and applied to machine learning. This eliminates the need to take photographs in the actual site space, as well as the need for manual annotation work to identify the mobile object. It also makes it possible to generate learning images under conditions that occur infrequently in the actual site space, such as extreme bad weather. Furthermore, even if a large portion of the virtual image of the mobile object to be annotated is obscured by obstacles (e.g., walls, other people, vehicles), it becomes possible to accurately add annotation information to images that are difficult for humans to judge.
[0052] (Summary of this disclosure) Based on the above description of embodiments, the following technical concepts corresponding to the items below are disclosed.
[0053] (Item 1) The method for generating mobile object learning data according to this disclosure involves placing a mobile object (vehicle TG1) and an imaging device (virtual camera 70) that virtually images the mobile object on a simulation space (80) that reproduces the field space to be learned, moving the mobile object as seen from the imaging device on the simulation space, and generating a virtual image by virtually imaging the mobile object with the imaging device, generating annotation information that identifies the mobile object in the virtual image based on the virtual image of the mobile object, and adding the annotation information to the virtual image to generate learning data necessary for training a detection model that detects any mobile object present in the field space. As a result, the method for generating mobile object learning data allows for the generation of multiple learning images after virtually placing various objects and mobile objects on a computer simulation space, which not only reduces the effort required to prepare the learning data (learning images) necessary for machine learning of a model that detects mobile objects (machine learning model), but also supports the efficiency of machine learning.
[0054] (Item 2) The mobile object learning data generation method described in Item 1 further comprises generating multiple annotation pieces for each mobile object, and generating the same number of learning data for each mobile object as the number of annotation pieces generated. As a result, the mobile object learning data generation method can generate multiple learning images for the same mobile object at each time point while moving the mobile object along the time axis in a computer simulation space, thereby supporting the improvement of machine learning accuracy.
[0055] (Item 3) In the mobile object learning data generation method described in Item 1 or 2, multiple mobile objects are captured in a single virtual image captured virtually by the imaging device, and annotation information for each of the multiple mobile objects is generated from the single virtual image. As a result, the mobile object learning data generation method makes it possible to efficiently generate training images useful for machine learning of a machine learning model that can simultaneously detect multiple mobile objects even in situations where multiple mobile objects are included in the field of view.
[0056] (Item 4) In the mobile object learning data generation method described in any one of Items 1 to 3, the annotation information includes at least one of a point, a rectangle, an arbitrary-shaped outline indicating the position of the mobile object in the virtual image, and text indicating the type of mobile object. As a result, the mobile object learning data generation method can efficiently generate annotation information that makes it possible to visually determine the presence of a mobile object in the image.
[0057] (Item 5) In the method for generating mobile object learning data described in any one of Items 1 to 4, the imaging device is movable in the simulation space. As a result, according to the method for generating mobile object learning data, even if the virtually placed virtual camera moves along the time axis in the real-world space corresponding to the computer simulation space, the learning data (training images) necessary for machine learning of a model (machine learning model) that detects mobile objects can be efficiently generated.
[0058] (Item 6) The mobile object learning data generation system according to the present disclosure comprises a processor (21, 31) and a memory (22, 32), wherein the processor cooperates with the memory to place a mobile object (vehicle TG1) and an imaging device (image capturing device 53) that virtually images the mobile object on a simulation space (80) that reproduces the field space to be learned, move the mobile object as seen from the imaging device on the simulation space and virtually image the mobile object with the imaging device, generate annotation information that identifies the mobile object in the virtual image based on the virtual image of the mobile object, and attach the annotation information to the virtual image to generate learning data necessary for learning a detection model that detects any mobile object present in the field space. As a result, the mobile object learning data generation system can generate multiple learning images after virtually placing various objects and mobile objects on a computer simulation space, which not only reduces the effort required to prepare learning data (learning images) necessary for machine learning of a model that detects mobile objects (machine learning model), but also supports the efficiency of machine learning.
[0059] While embodiments have been described above with reference to the attached drawings, this disclosure is not limited to such examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents can be conceived within the scope of the claims, and these are also understood to fall within the technical scope of this disclosure. Furthermore, the components of the embodiments described above can be combined in any way without departing from the spirit of the invention.
[0060] Although various embodiments have been described above with reference to the drawings, it goes without saying that the present invention is not limited to these examples. It is clear to those skilled in the art that various modifications or alterations can be conceived within the scope of the claims, and these will naturally also fall within the technical scope of the present invention. Furthermore, the components of the above embodiments may be combined in any way without departing from the spirit of the invention.
[0061] This application is based on Japanese Patent Application No. 2024-184120 filed on October 18, 2024, and its contents are incorporated herein by reference.
[0062] The technology disclosed herein is useful as a method and system for generating mobile object learning data, which reduces the effort required to prepare the learning data necessary for machine learning of a model that detects moving objects, and supports the efficiency of machine learning.
[0063] 10 Mobile object image recognition learning data generation system 20 Mobile object simulation device 21 Processor 22 Memory 23 Mobile object simulation unit 24 Annotation information generation unit 25 Field information 26 Mobile object position information 27 Annotation information 30 Mobile object learning image generation device 31 Processor 32 Memory 33 CG image generation unit 34 Annotation application unit 35 Learning image 40 Mobile object image recognition device 41 Processor 42 Memory 43 Learning unit 44 Inference unit 45 Machine learning model 46 Mobile object position information 50 Field work device 51 Processor 52 Memory 53 Image acquisition device 54 Mobile object information registration unit 55 Field master 56 Field image 100 Mobile object learning data generation system
Claims
1. A method for generating mobile object learning data, comprising:
1. Placing a mobile object and an imaging device that virtually images the mobile object in a simulation space that reproduces the field space to be learned; 2. Generating a virtual image by virtually imaging the mobile object with the imaging device while moving the mobile object as seen from the imaging device in the simulation space; 3. Generating annotation information that identifies the mobile object in the virtual image based on the virtual image of the mobile object; 4. Adding the annotation information to the virtual image to generate learning data necessary for training a detection model that detects any mobile object present in the field space.
2. A method for generating mobile body learning data according to claim 1, further comprising generating multiple annotation pieces for each mobile body, and generating the same number of learning data for each mobile body as the number of annotation pieces generated.
3. A method for generating mobile object learning data according to claim 1, wherein a plurality of mobile objects are captured in a single virtual image captured virtually by the imaging device, and annotation information for each of the plurality of mobile objects is generated from the single virtual image.
4. The method for generating mobile object learning data according to claim 1, wherein the annotation information includes at least one of a point indicating the position of the mobile object in the virtual image, a rectangle, an outline of an arbitrary shape, and text indicating the type of mobile object.
5. The method for generating mobile learning data according to claim 1, wherein the imaging device is movable in the simulation space.
6. A mobile object learning data generation system comprising a processor and memory, wherein the processor, in cooperation with the memory, places a mobile object and an imaging device that virtually images the mobile object on a simulation space that reproduces the field space to be learned, generates a virtual image by virtually imaging the mobile object with the imaging device while moving the mobile object as seen from the imaging device on the simulation space, generates annotation information that identifies the mobile object in the virtual image based on the virtual image of the mobile object, and adds the annotation information to the virtual image to generate learning data necessary for learning a detection model that detects any mobile object present in the field space.
Citation Information
Patent Citations
Learning data generation device, learning data generation method, machine learning method, and program
JP2019023858A
Machine learning device, machine learning method, and program
JP2019207662A
Image data generation device, image data generation method, and image data generation program
JP2024064413A