Method and system for generating mobile learning data.

By using a simulation space to generate virtual images with annotation information, the method addresses the inefficiencies of manual labor in preparing training data for machine learning models, improving the efficiency and reducing the effort needed for detecting moving objects.

JP2026073685APending Publication Date: 2026-05-01PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

The existing methods for generating learning data for machine learning models that detect moving objects require significant labor for data preparation and annotation, which is inefficient and labor-intensive.

Method used

A method and system that utilize a simulation space to virtually place mobile objects and an imaging device, generating virtual images with annotation information to create training data for detection models, reducing the need for manual labor and real-world data collection.

Benefits of technology

This approach reduces the effort required to prepare training data for machine learning models, enhancing efficiency by allowing for the generation of large amounts of annotated data without direct real-world data collection, including scenarios that are difficult or infrequent in real environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073685000001_ABST
    Figure 2026073685000001_ABST
Patent Text Reader

Abstract

This reduces the effort required to prepare the training data necessary for machine learning models that detect moving objects, thereby supporting the efficiency of machine learning. [Solution] The method for generating mobile object learning data includes: placing a mobile object and an imaging device that virtually images the mobile object on a simulation space that reproduces the field space to be learned; moving the mobile object as seen from the imaging device on the simulation space while the imaging device virtually images the mobile object to generate a virtual image; generating annotation information that identifies the mobile object in the virtual image based on the virtual image of the mobile object; attaching the annotation information to the virtual image; and generating learning data necessary for training a detection model that detects any mobile object present in the field space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method for generating mobile learning data and a mobile learning data generation system.

Background Art

[0002] Patent Document 1 discloses an image data collection device including a distance detection unit that detects the distance to an object, an imaging unit that captures an object to acquire image data, an object detection unit that detects the presence or absence of an object based on the detection value of the distance detection unit, and an image discrimination unit that discriminates the image data of the imaging unit according to the detection result of the object detection unit. This image data collection device detects the presence or absence and position of an object based on the detection value of the distance detection unit, creates label data representing the area where the object appears in the image data of the imaging unit based on the position when the object is detected, and stores the image data and the label data.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In Patent Document 1, it is assumed that the process of detecting the distance to an object and the process of imaging the object are at least executed in a real environment where the object exists. For this reason, there is a problem that a great deal of labor is required for the generation and preparation of learning data necessary for performing machine learning using deep learning. In addition, in order to perform efficient machine learning, additional information called annotation is often added to the learning data, but the addition of this annotation also requires a corresponding amount of labor, and it is considered that there is room for improvement in the preparation of efficient learning data.

[0005] This disclosure was devised in light of the conventional circumstances described above, and aims to reduce the effort required to prepare training data necessary for machine learning of models that detect moving objects, thereby supporting the efficiency of machine learning. [Means for solving the problem]

[0006] This disclosure provides a method for generating mobile object learning data, which involves placing a mobile object and an imaging device that virtually images the mobile object on a simulation space that reproduces the field space to be learned, moving the mobile object as seen from the imaging device on the simulation space while the imaging device virtually images the mobile object to generate a virtual image, generating annotation information that identifies the mobile object in the virtual image based on the virtual image of the mobile object, adding the annotation information to the virtual image, and generating learning data necessary for training a detection model that detects any mobile object present in the field space.

[0007] Furthermore, this disclosure provides a mobile object learning data generation system comprising a processor and memory, wherein the processor, in cooperation with the memory, places a mobile object and an imaging device that virtually images the mobile object on a simulation space that reproduces the field space to be learned, moves the mobile object as seen from the imaging device on the simulation space and virtually images the mobile object with the imaging device to generate a virtual image, generates annotation information that identifies the mobile object in the virtual image based on the virtual image of the mobile object, adds the annotation information to the virtual image, and generates learning data necessary for learning a detection model that detects any mobile object present in the field space.

[0008] These comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or recording media, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media. [Effects of the Invention]

[0009] According to this disclosure, it is possible to reduce the effort required to prepare training data necessary for machine learning of models that detect moving objects, and to support the efficiency of machine learning. [Brief explanation of the drawing]

[0010] [Figure 1] This figure shows an example of the system configuration of the mobile learning data generation system according to this embodiment. [Figure 2] A diagram showing an example of a data structure for field information used in field objects. [Figure 3] A diagram showing an example of a data structure for on-site information used by virtual cameras. [Figure 4] A diagram showing an example of a data structure for mobile object position information. [Figure 5] This diagram shows an example of the data structure of annotation information. [Figure 6] A diagram showing an example of the process for generating training images. [Figure 7] A diagram showing another example of training images. [Figure 8] A flowchart showing an example of the operation procedure for a mobile simulation device in chronological order. [Figure 9] A flowchart showing a time-series example of the operation procedure for the annotation information generation process in Figure 8. [Figure 10] A flowchart showing a time-series example of the operation procedure for a mobile learning image generation device. [Modes for carrying out the invention]

[0011] Hereinafter, with appropriate reference to the drawings, embodiments specifically disclosing the mobile learning data generation method and mobile learning data generation system according to this disclosure will be described in detail. However, unnecessarily detailed explanations may be omitted. For example, detailed explanations of already well-known matters and redundant explanations of substantially identical configurations may be omitted. This is to avoid the following explanation becoming unnecessarily verbose and to facilitate understanding by those skilled in the art. The attached drawings and the following explanation are provided to enable those skilled in the art to fully understand this disclosure and are not intended to limit the subject matter of the claims. Furthermore, in the following explanation, the same elements may be assigned the same reference numerals to simplify or omit explanations.

[0012] 1. Configuration of the mobile learning data generation system First, the system configuration of the mobile object learning data generation system 100 according to this embodiment will be described with reference to Figure 1. Figure 1 is a diagram showing an example of the system configuration of the mobile object learning data generation system according to this embodiment. The mobile object learning data generation system 100 includes at least a mobile object image recognition learning data generation system 10, a mobile object image recognition device 40, and a field work device 50. The learning data generation system 10 and the mobile object image recognition device 40, and the mobile object image recognition device 40 and the field work device 50 are connected to each other via a wired or wireless network so that data communication is possible between them. The wireless network may be, for example, a Wide Area Network (WAN), Local Area Network (LAN), Long Term Evolution (LTE), 4G, 5G or other mobile communication, power line communication, short-range wireless communication (e.g., Bluetooth® communication), or a network for mobile phones. The wired network may be, for example, a wired LAN or a wired WAN.

[0013] The mobile object image recognition learning data generation system 10 comprises a mobile object simulation device 20 and a mobile object learning image generation device 30. The mobile object simulation device 20 and the mobile object learning image generation device 30 are connected to each other via the wired or wireless network described above, enabling data communication between them. In the mobile object image recognition learning data generation system 10, the mobile object simulation device 20 and the mobile object learning image generation device 30 are configured as separate units, but they may also be configured as an integrated unit.

[0014] The mobile object simulation device 20 is configured using, for example, a personal computer (PC) or a server computer. The mobile object simulation device 20 virtually places a virtual camera 70 and one or more mobile objects TG1 (e.g., a vehicle, drone, person, etc.) to be detected on a computer simulation space 80 (see Figure 6) that virtually reproduces the field space to be learned, and executes the process of moving the mobile objects on the computer simulation space 80. While moving the mobile objects, the mobile object simulation device 20 executes the process of generating annotation information to identify the position of the mobile objects in the virtual captured images based on the virtual captured images virtually captured by the virtual camera 70. The mobile object simulation device 20 includes a processor 21 and memory 22. Although not shown in Figure 1, the mobile object simulation device 20 may further include an input device such as a mouse that can accept user operation.

[0015] The processor 21 is composed of at least one of, for example, a Central Processing Unit (CPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), or a Graphical Processing Unit (GPU). The processor 21 functions as a controller that manages the overall operation of the mobile body simulation device 20. The processor 21 performs control processing for coordinating the operations of each part of the mobile body simulation device 20, input / output processing of data between each part of the mobile body simulation device 20, arithmetic processing of data, and storage processing of data. The processor 21 operates according to a program stored in the memory 22. When operating, the processor 21 uses the memory 22 to cooperate to execute various processes, and temporarily stores the data generated or acquired by the processor 22 in the memory 22. By cooperating with the memory 22, the processor 21 realizes the functions of the mobile body simulation unit 23 and the annotation information generation unit 24.

[0016] The mobile body simulation unit 23 virtually arranges at least a virtual camera 70 and one or more mobile bodies TG1 (for example, vehicles, drones, people, etc.) to be detected on a computer simulation space 80 (see FIG. 6) that virtually reproduces the on-site space to be learned, and executes a process of moving the mobile body on the computer simulation space 80. The on-site objects virtually arranged on the computer simulation space 80 are not limited to the virtual camera 70 and the mobile body TG1, and also include various obstacles arranged in the real on-site space that is the basis of the computer simulation space 80. Details of such obstacles, the virtual camera 70, and the information of the mobile body TG1 (specifically, the on-site information 25 and the mobile body position information 26) will be described later with reference to FIGS. 2, 3, and 4 respectively. Further, the mobile body simulation unit 23 executes a process of generating a virtual captured image of the mobile body TG1 as seen from the virtual camera 70 on the computer simulation space 80 (that is, a virtual captured image of the mobile body TG1 virtually captured by the virtual camera 70).

[0017] The annotation information generation unit 24 executes a process of generating annotation information for specifying the position of the moving object TG1 in the virtual captured image based on the virtual captured image of the moving object TG1 viewed from the virtual camera 70 generated by the moving object simulation unit 23. The annotation information is additional information for specifying the position of the moving object TG1 in the virtual captured image of the virtual camera 70, and includes, for example, at least one of a point indicating the position of the moving object TG1 in the virtual captured image, a rectangle, an arbitrarily shaped contour line, and text indicating the type of the moving object TG; The details of the annotation information 27 will be described later with reference to FIG. 5.

[0018] The memory 22 is configured using, for example, a Random Access Memory (RAM) and a Read Only Memory (ROM), and temporarily holds programs necessary for the operation of the moving object simulation device 20 and data acquired or generated during the operation. The RAM is, for example, a work memory used during the operation of the moving object simulation device 20. The ROM stores and holds, for example, a program for controlling the moving object simulation device 20 in advance. The memory 22 temporarily stores the site information 25, the moving object position information 26, and the annotation information 27 generated by the processor 21.

[0019] Here, the site information 25, the moving object position information 26, and the annotation information 27 will be described with reference to FIGS. 2 to 5 respectively. FIG. 2 is a diagram showing an example of the data structure of the site information for the site object. FIG. 3 is a diagram showing an example of the data structure of the site information for the virtual camera

[0018] . FIG. 4 is a diagram showing an example of the data structure of the moving object position information . FIG. 5 is a diagram showing an example of the data structure of the annotation information .

[0020] As shown in Figure 2, the field information 25a for a field object consists of a record for each field object that includes a field object ID that identifies the field object, a time, coordinates, and a CG type. To distinguish between field objects and moving objects, moving objects are treated differently from field objects, while field objects are treated as objects that do not move. The time indicates the time when the field object in question was placed in the computer simulation space 80. The coordinates indicate the three-dimensional coordinates that identify the position in which the field object in question was placed in the computer simulation space 80. The CG type indicates what kind of object the field object in question is. This field information 25a is an example of field information 25 generated by the mobile body simulation unit 23 of the processor 21 of the mobile body simulation device 20 at the timing when the object in question is placed in the computer simulation space 80.

[0021] As shown in Figure 3, the field information 25b for the virtual camera 70 consists of a record for each virtual camera 70 that includes a virtual camera ID that identifies the virtual camera 70, a time, coordinates, CG type, direction, field of view, focal length, and shooting time interval. The time indicates the time when the virtual camera 70 was placed in the computer simulation space 80. The coordinates indicate the three-dimensional coordinates that identify the position where the virtual camera 70 was placed in the computer simulation space 80. The CG type indicates what type of camera device the virtual camera is. The direction, field of view, focal length, and shooting time interval indicate the camera parameters set for the virtual camera, respectively. This field information 25b is an example of field information 25 generated by the mobile simulation unit 23 of the processor 21 of the mobile simulation device 20 at the timing when the virtual camera 70 is placed in the computer simulation space 80.

[0022] As shown in Figure 4, the mobile object position information 26 consists of a record for each mobile object that includes a mobile object ID to identify the mobile object, a time, coordinates, and a CG type. The time indicates the time when the mobile object was placed in the computer simulation space 80 or the time at a certain moment after it moved. The coordinates indicate the three-dimensional coordinates that specify the position where the mobile object was placed in the computer simulation space 80 or the position at a certain moment after it moved. The CG type indicates what kind of object the mobile object is. This mobile object position information 26 is generated by the mobile object simulation unit 23 of the processor 21 of the mobile object simulation device 20 at the time the mobile object was placed in the computer simulation space 80 or at a certain moment after it moved.

[0023] As shown in Figure 5, the annotation information 27 consists of a record for each pair of a moving object and a virtual camera that images the moving object, which includes a virtual camera ID that identifies the virtual camera, a moving object ID that identifies the moving object, a time, an annotation type, and annotation data. Although not shown in Figure 5, the annotation information 27 may also include actual data of the virtual image of the moving object TG1 as seen by the virtual camera 70 generated by the moving object simulation unit 23, or information about the storage location of the actual data. The time indicates the type of annotation generated based on the virtual image captured by the virtual camera of the relevant moving object. The annotation data indicates the pixel position (2D coordinate) in the virtual image that identifies the range of existence of the annotation for identifying the moving object present in the virtual image of the virtual camera 70. This annotation information 27 is generated by the annotation generation unit 24 of the processor 21 of the moving object simulation device 20 at the timing when the annotation is generated corresponding to the relevant moving object.

[0024] The mobile object learning image generation device 30 is configured using, for example, a personal computer (PC) or a server computer. The mobile object learning image generation device 30 uses on-site information 25, mobile object position information 26, and annotation information 27 sent from the mobile object simulation device 20 to generate training images 35 necessary for machine learning to generate a machine learning model 45 for detecting mobile objects. The mobile object learning image generation device 30 includes a processor 31 and memory 32. Although not shown in Figure 1, the mobile object learning image generation device 30 may further include an input device such as a mouse that can accept user input.

[0025] The processor 31 is composed of at least one of the following: a CPU, DSP, FPGA, or GPU. The processor 31 functions as a controller that oversees the overall operation of the mobile learning image generation device 30. The processor 31 performs control processing to coordinate the operation of each part of the mobile learning image generation device 30, data input / output processing between each part of the mobile learning image generation device 30, data calculation processing, and data storage processing. The processor 31 operates according to the program stored in the memory 32. During operation, the processor 31 uses the memory 32 to perform various processes in cooperation with it and temporarily stores the data generated or acquired by the processor 31 in the memory 32. By cooperating with the memory 32, the processor 31 realizes the functions of the CG image generation unit 33 and the annotation unit 34.

[0026] The CG image generation unit 33 generates a Computer Graphic (CG) image in which various objects indicated by the site information 25 and the mobile object indicated by the mobile object position information 26 are placed in the computer simulation space 80, based on the site information 25 and mobile object position information 26 sent from the mobile object simulation device 20. Since this method of generating CG images can be realized using already known publicly available techniques, a detailed explanation is omitted here.

[0027] The annotation unit 34 uses the annotation information 27 sent from the mobile simulation device 20 to generate training images 35 by adding at least one annotation indicated by the annotation information to the CG images generated by the CG image generation unit 33. The annotation unit 34 may generate training images 35 in a format that conforms to the specifications of the machine learning model 45.

[0028] Here, the training images 35 will be explained with reference to Figures 6 and 7. Figure 6 is a diagram showing an example of the overview of the training image generation process. Figure 7 is a diagram showing another example of a training image. The training images 35 are input images necessary for the machine learning process in which the mobile image recognition device 40 constructs (e.g., generates or updates) a machine learning model 45. When a large number of very diverse training images 35 are prepared, the mobile image recognition device 40 performs machine learning on the machine learning model.

[0029] As shown in Figure 6, a virtual camera 70 is positioned at the top of a pole in the z-direction on a computer simulation space 80 illustrating a parking lot PK1, and the mobile object TG1 to be detected is positioned there. The CG image generation unit 33 of the processor 31 of the mobile object learning image generation device 30 generates a CG image IMG1 with the mobile object TG1 as the subject as seen from the virtual camera 70. In this CG image IMG1, the mobile object TG1 is shown almost in the center, and other objects around the parking lot PK1 may also be shown. The annotation unit 34 of the processor 31 of the mobile object learning image generation device 30 generates a learning image IMG2 by adding annotation information LB1 to the CG image IMG1. The annotation information LB1 of the learning image IMG2 consists of, for example, the outline of the vehicle which is the mobile object TG1 and the area enclosed by that outline filled with a predetermined color.

[0030] As shown in Figure 7, the training image IMG3 may also contain multiple moving objects (e.g., vehicles) in a single image. The training image IMG3 is assigned annotation information LB11, LB12, LB13, LB14, LB15, LB16, LB17, LB18, LB19, LB20, LB21, LB22, LB23, and LB24 for each vehicle traveling in each of the multiple lanes.

[0031] Memory 32 is configured, for example, using RAM and ROM, and temporarily stores programs necessary for the operation of the mobile object learning image generation device 30, as well as data acquired or generated during operation. RAM is, for example, work memory used during the operation of the mobile object learning image generation device 30. ROM stores, for example, programs for controlling the mobile object learning image generation device 30 in advance. Memory 32 temporarily stores field information 25, mobile object position information 26, and annotation information 27 sent from the mobile object simulation device 20.

[0032] The mobile object image recognition device 40 is configured using, for example, a personal computer (PC) or a server computer. The mobile object image recognition device 40 receives and acquires a plurality of training images 35 generated by the mobile object learning image generation device 30. Using the plurality of training images 35, the mobile object image recognition device 40 performs machine learning processing to construct (e.g., generate or update) a machine learning model capable of detecting moving objects in the field image 56, which is a virtual image captured by the field work device 50. The mobile object image recognition device 40 includes a processor 41 and memory 42. Although not shown in Figure 1, the mobile object image recognition device 40 may further include an input device such as a mouse that can accept user input.

[0033] The processor 41 is composed of at least one of the following: a CPU, DSP, FPGA, or GPU. The processor 41 functions as a controller that oversees the overall operation of the mobile learning image recognition device 40. The processor 41 performs control processing to coordinate the operation of each part of the mobile learning image recognition device 40, data input / output processing between each part of the mobile learning image recognition device 40, data calculation processing, and data storage processing. The processor 41 operates according to the program stored in the memory 42. When operating, the processor 41 uses the memory 42 to cooperate in executing various processes and temporarily stores data generated or acquired by the processor 41 in the memory 42. By cooperating with the memory 42, the processor 41 realizes the functions of the learning unit 43 and the inference unit 44.

[0034] The learning unit 43 uses multiple training images 35 stored in memory 42 to perform machine learning processing to construct (e.g., generate or update) a machine learning model capable of detecting moving objects in virtual images captured by the field work device 50. The learning unit 43 stores the machine learning model 45 obtained through the machine learning processing in memory 42.

[0035] The inference unit 44 refers to a machine learning model stored in memory 42, receives a field image 56, which is a virtual image captured at various work sites and sent from the field work device 50, detects moving objects in the field image 56, and executes a process to generate information that identifies the moving object. The inference unit 44 generates moving object position information 46 that indicates the position of the moving object identified in the field image 56 and feeds it back (transmits) it to the field work device 50. This moving object position information 46 may have the same items as the moving object position information 26 shown in Figure 4, or it may be further configured to include information indicating the position of the moving object in the virtual image captured by the moving object, the type of the moving object, or both.

[0036] Memory 42 is configured, for example, using RAM and ROM, and temporarily stores programs necessary for the operation of the mobile learning image recognition device 40, as well as data acquired or generated during operation. RAM is, for example, work memory used during the operation of the mobile learning image recognition device 40. ROM stores, for example, programs for controlling the mobile learning image recognition device 40 in advance. Memory 42 temporarily stores multiple training images 35 sent from the mobile learning image generation device 30. Memory 42 also stores machine learning models 45 constructed (for example, generated or updated) by machine learning performed by the processor 41.

[0037] The machine learning model 45 detects moving objects in the input virtual image and outputs information to identify those moving objects. The machine learning model 45 is constructed (e.g., generated or updated) by machine learning performed by the learning unit 43 of the processor 41 and stored in the memory 42.

[0038] The field work device 50 is configured using, for example, a personal computer (PC). The field work device 50 is used for each type of task performed in the field space, and generates a field image 56 by capturing images of the field space, and requests the mobile object image recognition device 40 to identify the position of a subject (for example, a moving object) that appears in the field image 56. The field work device 50 includes a processor 51 and a memory 52. ​​Although not shown in Figure 1, the field work device 50 may further include an input device such as a mouse that can accept user input.

[0039] The processor 51 is composed of at least one of the following: a CPU, DSP, FPGA, or GPU. The processor 51 functions as a controller that oversees the overall operation of the field work device 50. The processor 51 performs control processing to coordinate the operation of each part of the field work device 50, data input / output processing between each part of the field work device 50, data calculation processing, and data storage processing. The processor 51 operates according to the program stored in the memory 52. ​​When operating, the processor 51 uses the memory 52 to perform various processes in cooperation with it and temporarily stores data generated or acquired by the processor 51 in the memory 52. ​​By cooperating with the memory 52, the processor 51 realizes the functions of the mobile information registration unit 54.

[0040] The mobile object information registration unit 54 performs the process of storing (registering) the mobile object position information 46, which is fed back (responded) from the mobile object image recognition device 40, in the memory 52.

[0041] Memory 52 is configured, for example, using RAM and ROM, and temporarily stores programs necessary for the operation of the field work device 50, as well as data acquired or generated during operation. RAM is, for example, work memory used during the operation of the field work device 50. ROM stores, for example, programs for controlling the field work device 50 in advance. Memory 52 temporarily stores field images 56 captured by the image acquisition device 53 and mobile object position information 46 fed back from the mobile object image recognition device 40. Memory 52 also stores field master data 55, which is various data necessary for performing tasks in the field space.

[0042] The field master 55 is a database that stores various types of data necessary for executing tasks in the field space handled by the field work device 50.

[0043] The image acquisition device 53 captures and generates a virtual image (i.e., a field image 56) that shows the state of the field space handled by the field work equipment 50, and sends it to the processor 51. The processor 51 generates a recognition processing request that includes this field image 56 and sends it to the mobile object image recognition device 40. The response (feedback) to this recognition processing request becomes the mobile object position information 46.

[0044] 2. Operation of the Mobile Learning Data Generation System Next, the operation procedure of the mobile object learning data generation system 100 according to this embodiment will be described with reference to Figures 8 to 10. Figure 8 is a flowchart showing an example of the operation procedure of the mobile object simulation device in chronological order. Figure 9 is a flowchart showing an example of the operation procedure of the annotation information generation process in Figure 8 in chronological order. Figure 10 is a flowchart showing an example of the operation procedure of the mobile object learning image generation device in chronological order. The series of processes in Figures 8 and 9 are mainly executed by the processor 21 of the mobile object simulation device 20. The series of processes in Figure 10 are mainly executed by the processor 31 of the mobile object learning image generation device 30.

[0045] In Figure 8, the processor 21 initializes the simulation process for the moving object, prepares a computer simulation space 80 that reproduces the target field space, and places various objects and the moving object in the computer simulation space 80 (St1). In this initialization, the processor 21 sets the information necessary for the simulation process based on user operation or input from other systems. This necessary information includes, for example, the initial position of the moving object in the computer simulation space 80, its movement speed, destination, the three-dimensional shape of the computer simulation space 80, the placement of pathways that the moving object can move through in the computer simulation space 80, the placement of obstacles, the installation position, direction, field of view, focal length, and shooting time interval of the virtual camera 70 placed in the computer simulation space 80.

[0046] The processor 21 sets the number of annotation information to be generated for a single moving object (i.e., the number of training images to be generated) based on user input (St2). The moving object simulation device 20 repeats the generation of annotation information for a single moving object, matching the number of annotation information to be generated set in step St2 (i.e., the series of processes from steps St3 to St6). The processor 21 performs a process to move (specifically update) the moving object in the computer simulation space 80 based on time-series data 26a, among the various objects and moving objects placed in the computer simulation space 80 during initialization in step St1 (St3). This process in step St3 may be performed using any simulation method, for example, based on the occurrence of discrete movements or based on a continuous time step. The time-series data 26a is data indicating the position of the moving object at each time (specifically, 3D coordinates indicating the position in the computer simulation space 80).

[0047] In step St3, the processor 21 executes a process to generate annotation information to identify the moving object while it is moving (St4). Details of the process in step St4 will be described later with reference to Figure 9. In step St3, the processor 21 outputs the current time on the computer simulation space 80, the field information 25, the moving object position information 26, and the annotation information 27 obtained while updating the position of the moving object to the moving object learning image generation device 30 (St5). If the processor 21 determines that it has generated a number of annotation information items that match the number of annotation information items set in step St2 (St6, YES), it terminates the series of processes shown in Figure 8. On the other hand, if it determines that it has not yet generated a number of annotation information items that match the number of annotation information items set in step St2 (St6, NO), the processor 21 returns to step St3.

[0048] In Figure 9, the processor 21 uses the field information 25 and the mobile object position information 26 to generate a virtual image (see CG image IMG1 in Figure 6) of the mobile object TG1, which is the subject, as seen from a virtual camera 70 placed in the computer simulation space 80 (St11). The processor 21 uses the mobile object position information 26 to organize the virtual image generated in step St21 in correspondence with the position of the mobile object, and projects the parts of the mobile object to be annotated onto the virtual image (St12). The processor 21 generates annotation information 27 (see Figure 5) according to which pixels of the virtual image projected in step St12 should be annotated with what kind of mobile object annotation (St13). The processor 21 may also determine what kind of annotation information to generate according to the specifications of the machine learning model used in the mobile object image recognition device 40, and generate annotation information according to the determination result.

[0049] In Figure 10, the processor 31 reads and acquires various information (specifically, site information 25, mobile object position information 26, and annotation information 27) sent from the mobile object simulation device 20 from the memory 32 (St21). Using the site information 25 and mobile object position information 26 acquired in step St21, the processor 31 generates a CG image that corresponds to a virtual image taken when the mobile object TG1, which is the subject as seen from a virtual camera 70 placed in the computer simulation space 80, is virtually captured (St22).

[0050] The processor 31 generates a training image 35 by adding annotations (e.g., superimposing) to the CG image generated in step St22 using the annotation information 27 acquired in step St21 (St23). The annotations using the annotation information 27 include, for example, points indicating the position of a moving object, rectangles, contour lines of arbitrary shapes, text identifying the type of moving object, or a combination thereof, in order to make it easier to identify moving objects in the CG image. The processor 31 outputs the training image 35 generated in step St23 to the moving object image recognition device 40 (St24).

[0051] As described above, the mobile object learning data system 100 according to this embodiment allows for the generation of a large amount of learning data (learning images) in a computer simulation space 80, where various objects and mobile objects are virtually placed in a manner similar to the real-world environment, without having to directly visit the actual field space to be learned. These generated learning images can then be used and applied to machine learning. This eliminates the need to take photographs in the real-world environment, as well as the need for manual annotation work to identify the mobile object. Furthermore, it enables the generation of learning images under conditions that occur infrequently in the real-world environment, such as extreme bad weather. In addition, even if a large portion of the virtual image of the mobile object to be annotated is obscured by obstacles (e.g., walls, other people, vehicles), accurate annotation information can be added to images that are difficult for humans to judge.

[0052] (Summary of this disclosure) The above description of embodiments discloses the technical concepts corresponding to the following items.

[0053] (Item 1) The method for generating mobile learning data related to this disclosure is: A mobile object (vehicle TG1) and an imaging device (virtual camera 70) that virtually images the mobile object are placed on a simulation space (80) that reproduces the field space to be studied. The moving object, as viewed from the imaging device, is moved in the simulation space, and the imaging device virtually images the moving object to generate a virtual image. Annotation information is generated to identify the moving object in the virtual image based on the virtual image of the moving object. The annotation information is added to the virtual captured image to generate training data necessary for training a detection model that detects any moving object present in the field space. As a result, this mobile object learning data generation method allows for the virtual placement of various objects and mobile objects in a computer simulation space, thereby generating multiple training images. This not only reduces the effort required to prepare the training data (training images) necessary for machine learning a model that detects mobile objects (machine learning model), but also supports the efficiency of machine learning.

[0054] (Item 2) In the mobile learning data generation method described in item 1, The method further includes generating multiple annotation pieces for each of the moving objects, The same number of training data for each mobile object as the number of annotation information generated are generated. This means that, according to the method for generating mobile object learning data, multiple training images can be generated for the same mobile object at each point in time while moving the mobile object along the time axis in a computer simulation space, thereby supporting the improvement of machine learning accuracy.

[0055] (Item 3) In the method for generating mobile object learning data described in item 1 or 2, multiple mobile objects are captured in a single virtual image captured virtually by the imaging device, and annotation information for each of the multiple mobile objects is generated from the single virtual image. This means that, according to the method for generating mobile object learning data, it is possible to efficiently generate training images useful for machine learning models that can simultaneously detect multiple mobile objects even in situations where multiple mobile objects are included in the field of view.

[0056] (Item 4) In the mobile learning data generation method described in any one of items 1 to 3, The annotation information includes at least one of the following: a point indicating the position of the moving object in the virtual image, a rectangle, an outline of an arbitrary shape, and text indicating the type of the moving object. As a result, the mobile object learning data generation method can efficiently generate annotation information that makes it possible to visually identify the presence of a mobile object in the captured image.

[0057] (Item 5) In the mobile learning data generation method described in any one of items 1 to 4, The imaging device is movable within the simulation space. As a result, according to this method for generating mobile object learning data, even if the virtually positioned virtual cameras move along the time axis in the real-world space corresponding to the computer simulation space, it is possible to efficiently generate the learning data (training images) necessary for machine learning of a model (machine learning model) that detects moving objects in the same way.

[0058] (Item 6) The mobile learning data generation system related to this disclosure is It is equipped with a processor (21, 31) and memory (22, 32), The aforementioned processor, in cooperation with the memory, A mobile object (vehicle TG1) and an imaging device (image capturing device 53) that virtually captures images of the mobile object are placed on a simulation space (80) that reproduces the field space to be studied. The moving object, as viewed from the imaging device, is moved in the simulation space while the imaging device virtually images the moving object. Annotation information is generated to identify the moving object in the virtual image based on the virtual image of the moving object. The annotation information is added to the virtual captured image to generate training data necessary for training a detection model that detects any moving object present in the field space. As a result, the mobile object learning data generation system can generate multiple training images by virtually placing various objects and mobile objects in a computer simulation space. This not only reduces the effort required to prepare the training data (training images) necessary for machine learning a model that detects mobile objects (machine learning model), but also supports the efficiency of machine learning.

[0059] While embodiments have been described above with reference to the attached drawings, this disclosure is not limited to such examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents can be conceived within the scope of the claims, and these are also understood to fall within the technical scope of this disclosure. Furthermore, the components of the embodiments described above can be combined in any way without departing from the spirit of the invention. [Industrial applicability]

[0060] The technology disclosed herein is useful as a method and system for generating mobile object learning data, which reduces the effort required to prepare the learning data necessary for machine learning of a model that detects moving objects, and supports the efficiency of machine learning. [Explanation of Symbols]

[0061] 10. Mobile Object Image Recognition Learning Data Generation System 20 Mobile Simulation Device 21 processors 22 memory 23 Mobile Simulation Department 24 Annotation Information Generation Unit 25. Field Information 26 Mobile location information 27 Annotation Information 30 Mobile object learning image generation device 31 processors 32 memory 33 CG Image Generation Unit 34 Annotation section 35 images for learning 40 Mobile object image recognition device 41 processors 42 memory 43 Learning Department 44 Reasoning part 45 Machine Learning Models 46 Mobile object location information 50 Field work equipment 51 processors 52 memory 53 Image acquisition devices 54 Mobile Information Registration Unit 55 Field Master 56 On-site images 100 Mobile Learning Data Generation System

Claims

1. A mobile object (TG1) and an imaging device (70) that virtually images the mobile object are placed on a simulation space (80) that reproduces the field space to be studied. The moving object, as viewed from the imaging device, is moved in the simulation space, and the imaging device virtually images the moving object to generate a virtual image. Annotation information is generated to identify the moving object in the virtual image based on the virtual image of the moving object. The annotation information is added to the virtual captured image, and training data necessary for training a detection model that detects any moving object present in the field space is generated. A method for generating mobile object learning data.

2. The method further includes generating multiple annotation pieces for each of the moving objects, The same number of the aforementioned annotation information items as the number of the aforementioned learning data items for each moving object are generated. The method for generating mobile learning data according to claim 1.

3. Multiple moving objects are captured in a single virtual image captured virtually by the imaging device. Multiple annotation pieces for each of the moving objects are generated from a single virtual image. The method for generating mobile learning data according to claim 1.

4. The annotation information includes at least one of the following: a point indicating the position of the moving object in the virtual image, a rectangle, an arbitrary-shaped outline, and text indicating the type of the moving object. The method for generating mobile learning data according to claim 1.

5. The imaging device is movable in the simulation space. The method for generating mobile learning data according to claim 1.

6. Equipped with a processor and memory, The aforementioned processor, in cooperation with the memory, A moving object and an imaging device that virtually images the moving object are placed in a simulation space that reproduces the real-world environment to be studied. The moving object, as viewed from the imaging device, is moved in the simulation space, and the imaging device virtually images the moving object to generate a virtual image. Annotation information is generated to identify the moving object in the virtual image based on the virtual image of the moving object. The annotation information is added to the virtual captured image, and training data necessary for training a detection model that detects any moving object present in the field space is generated. Mobile object learning data generation system.

Citation Information

Patent Citations

  • Learning data generation device, learning data generation method, machine learning method, and program

    JP2019023858A

  • Machine learning device, machine learning method, and program

    JP2019207662A

  • Image data collection device and image data sorting method

    JP2021033572A