Training Data Generation Method and Training Data Generation System
By using tracks as pseudo-labels to automatically generate labeled training data, the time and cost of data labeling are reduced, facilitating efficient training and optimization of object recognition models.
Patent Information
- Application Number
- JP2022162348
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-10-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-10-07
AI Technical Summary
Labeling data for object recognition models based on machine learning is time-consuming and costly due to the need for manual annotation.
A method and system for generating labeled training data using tracks as pseudo-labels, automatically obtained by tracking the same moving object across a series of images, reducing the manual effort in data labeling.
Significantly reduces time and cost in generating labeled training data, enabling efficient and effective training of object recognition models, allowing for quick updates and optimization.
Smart Images

Figure 0007697441000001 
Figure 0007697441000002 
Figure 0007697441000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an object identification model based on machine learning.
Background Art
[0002] Patent Document 1 discloses a tracking device that tracks an object (e.g., a person) by using a recognition model. The recognition model extracts an object from an image captured by a surveillance camera. Then, the recognition model extracts feature amounts of the extracted object and tracks the extracted object.
[0003] Non-Patent Document 1 discloses a tracker called "ByteTrack".
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] An object recognition model based on machine learning is used to identify objects in images. To realize an excellent object recognition model, it is necessary to train the object recognition model using a sufficient amount of labeled training data. However, generally, labeling (annotating) data requires a great deal of time and a lot of manpower, and therefore, it is costly.
Means for Solving the Problem
[0007] A first aspect of the present disclosure relates to a training data generation method for generating labeled training data used for training an object recognition model based on machine learning. The training data generation method includes detecting moving objects in a series of images, automatically obtaining a track, which is information representing a time series of the same moving object in a series of images, by tracking the same moving object in a series of images using a tracker, generating labeled training data by assigning the track as a label to the series of images and
[0008] A second aspect of the present disclosure relates to a training data generation system for generating labeled training data used for training an object recognition model based on machine learning. The training data generation system includes one or more processors. The one or more processors are configured to detect moving objects in a series of images, automatically obtain a track, which is information representing a time series of the same moving object in a series of images, by tracking the same moving object in a series of images using a tracker, generate labeled training data by assigning the track as a label to the series of images as described above.
Advantages of the Invention
[0009] According to the present disclosure, a track is used as a label in labeled training data. The track can be automatically acquired by tracking the same moving object in a series of images. Therefore, it is possible to greatly reduce the manual work in labeling data, that is, in generating labeled training data. As a result, time and costs are greatly saved.
[0010] Furthermore, since it is possible to acquire labeled training data while saving time and costs, it is possible to quickly train an object recognition model using a sufficient amount of labeled training data. That is, it is possible to train the object recognition model efficiently and effectively. As a result, the object recognition model is further optimized.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Embodiments for Carrying Out the Invention
[0012] Embodiments of the present disclosure will be described with reference to the accompanying drawings.
[0013] 1. Overview 1-1. Object Identification Model FIG. 1 is a conceptual diagram for explaining an object identification model MDL according to the present embodiment. The object identification model MDL is used to identify an object in an image. Typically, the object identified by the object identification model MDL is a moving object. Examples of moving objects include humans (pedestrians), vehicles, motorcycles, bicycles, robots, and the like.
[0014] The object identification model MDL is based on machine learning. For example, the object identification model MDL is based on a Transformer, which is a type of deep learning model. As another example, the object identification model MDL may be based on a CNN (Convolutional Neural Network).
[0015] Typically, the object identification model MDL performs feature extraction to identify an object. That is, the object identification model MDL extracts feature amounts of an object detected in an image and identifies the object based on the extracted feature amounts.
[0016] The object recognition model MDL may identify (specify) the same object in different images captured by two or more different cameras. In that case, the same moving object can be tracked across two or more different cameras. In the example shown in FIG. 1, the image IMG1 is captured by the camera C1, and the other image IMG2 is captured by the other camera C2. The object recognition model MDL identifies (re-identifies) the same pedestrian in the two different images IMG1 and IMG2. Such an object recognition model MDL is also called a "human re-identification model" or "person re-identification model". The object recognition model MDL may be a person re-identification model based on a Transformer.
[0017] In order to realize an excellent object recognition model MDL, it is necessary to train the object recognition model MDL using a sufficient amount of labeled training data. However, generally, labeling (annotating) data requires a great deal of time and a lot of manpower, and therefore, it is costly.
[0018] Therefore, the present disclosure provides a technique that can reduce the manual work in labeling (annotating) data, that is, generating labeled training data. The present disclosure further provides a technique that can train the object recognition model MDL using a sufficient amount of labeled training data.
[0019] 1-2. System Configuration FIG. 2 is a block diagram showing an example of the system configuration according to the present embodiment. The system according to the present embodiment includes a video collection unit 100, a training data generation system 200, a model training system 300, and an object recognition system 400.
[0020] The video collection unit 100 collects videos. For example, the video collection unit 100 communicates with at least one camera and collects videos captured by the at least one camera. The at least one camera is installed in streets, buildings, etc. As another example, the video collection unit 100 may collect videos from video posting sites. The video collection unit 100 supplies the collected video data to the training data generation system 200.
[0021] The training data generation system 200 receives video data from the video collection unit 100. The training data generation system 200 automatically or almost automatically generates labeled training data LAD based on the video data. The labeled training data LAD is training data (training data) with labels assigned to each object in the image. The labeled training data LAD is also called annotated training data. Details of the generation of the labeled training data LAD will be described later.
[0022] The model training system 300 acquires the labeled training data LAD generated by the training data generation system 200. The model training system 300 trains (trains) the object identification model MDL based on the labeled training data LAD. In other words, the model training system 300 trains the object identification model MDL by using the labeled training data LAD. Here, "supervised learning" or "semi-supervised learning" is used to train the object identification model MDL.
[0023] The object recognition system 400 acquires an object recognition model MDL trained by the model training system 300. The object recognition system 400 performs object recognition processing using the object recognition model MDL. More specifically, the object recognition system 400 acquires video data and inputs the video data into the object recognition model MDL to identify the objects in the video data.
[0024] The training data generation system 200, the model training system 300, and the object recognition system 400 may be a distributed system. That is, the training data generation system 200, the model training system 300, and the object recognition system 400 may be constructed on different nodes (computers) that communicate with each other. As another example, some of the training data generation system 200, the model training system 300, and the object recognition system 400 may be constructed on a single node (computer).
[0025] 1-3. Tracks Used as Labels FIG. 3 is a conceptual diagram for explaining "tracks". FIG. 3 shows a series of images IMG at different time steps (t = t1, t2, t3, ···) included in the video. Each image IMG shows at least one moving object. Examples of moving objects include humans (pedestrians), vehicles, motorcycles, bicycles, robots, and the like.
[0026] The training data generation system 200 detects the moving objects in the series of images IMG included in the video. The bounding box BX represents the position of the detected moving object in the image IMG. The training data generation system 200 acquires the information of the bounding box BX of each moving object in the series of images IMG.
[0027] As a moving object moves, the bounding box BX representing the moving object also moves within a series of images IMG. Multiple bounding boxes BX representing the same moving object in a series of images IMG at different time steps are spatially continuous. Therefore, by focusing on the movement of the bounding box BX, multiple bounding boxes BX representing the same moving object in a series of images IMG can be identified. For example, in FIG. 3, multiple bounding boxes BX1[t] (t = t1, t2, t3, ···) in a series of images IMG at different time steps represent the same pedestrian. Multiple bounding boxes BX2[t] (t = t1, t2, t3, ···) in a series of images IMG at different time steps represent the same vehicle. By identifying multiple bounding boxes BX representing the same moving object in a series of images IMG, it becomes possible to track the same moving object in a series of images IMG.
[0028] A "tracker" is software that automatically tracks the same moving object in a series of images IMG based on a tracking algorithm. For example, "ByteTrack" is known as a powerful tracker (see Non-Patent Document 1 above).
[0029] The tracker (i.e., the tracking algorithm) tracks the same moving object in a series of images IMG based on the movement of the bounding box BX. More specifically, the tracker tracks the same moving object in a series of images IMG by identifying a plurality of bounding boxes BXi[t] (where t = t1, t2, t3, ···) that represent the same moving object in the series of images IMG. Here, i (= 1, 2, 3, ···) is the identifier of the plurality of bounding boxes BX that represent the same moving object. The tracker associates the plurality of bounding boxes BXi[t] that represent the same moving object in a series of images IMG at different time steps with each other. Note that the tracker does not require feature extraction to track the same moving object. The tracker tracks the same moving object based on the movement of the bounding box BX without performing feature extraction.
[0030] "Track TRi" is information representing the time series of the same moving object in a series of images IMG. More specifically, track TRi is identification information indicating a plurality of bounding boxes BXi[t] that represent the same moving object in a series of images IMG at different time steps. Note that track TRi is not the identification information of the moving object itself. For example, track TRi does not indicate who the pedestrian is. At this stage, there is no need to know who the pedestrian is.
[0031] As described above, track TRi can be automatically obtained by a tracker that tracks the same moving object in a series of images IMG. According to the present embodiment, such a track TRi is used as a label in the labeled training data LAD.
[0032] FIG. 4 is a conceptual diagram for explaining the labeled training data LAD according to the present embodiment. A track TRi is assigned as a label to a series of images IMG included in a video. The track TRi can also be called a "pseudo label". A series of images IMG to which the track TRi is assigned as a label is the labeled training data LAD.
[0033] The training data generation system 200 uses a tracker to track the same moving object in a series of images IMG. In other words, the training data generation system 200 tracks the same moving object in a series of images IMG based on a tracking algorithm. As a result, the training data generation system 200 can automatically obtain a track TRi, which is information representing the time series of the same moving object in a series of images IMG. The training data generation system 200 generates the labeled training data LAD by assigning the track TRi as a label to a series of images IMG.
[0034] 1-4. Effects As described above, according to the present embodiment, the track TRi is used as a label in the labeled training data LAD. The track TRi can be automatically obtained by tracking the same moving object in a series of images IMG. Therefore, it is possible to significantly reduce the manual work in labeling (annotating) data, that is, in generating the labeled training data LAD. As a result, time and costs are significantly saved.
[0035] Furthermore, since it is possible to obtain labeled training data LAD while saving time and cost, it becomes possible to quickly train the object identification model MDL using a sufficient amount of labeled training data LAD. That is, it becomes possible to train the object identification model MDL efficiently and effectively. As a result, the object identification model MDL is further optimized. For example, the object identification model MDL can be updated without delay in the environment (e.g., region, season). In other words, it becomes possible to optimize (fine-tune) the object identification model MDL in consideration of the latest environment.
[0036] Hereinafter, specific examples of the training data generation system 200 and the model training system 300 will be described.
[0037] 2. Training Data Generation System FIG. 5 is a block diagram showing a configuration example of the training data generation system 200 according to the present embodiment. The training data generation system 200 includes an I / O (Input / Output) interface 201, an HMI (Human Machine Interface) 202, one or more processors 203 (hereinafter simply referred to as "processor 203"), and one or more storage devices 204 (hereinafter simply referred to as "storage device 204").
[0038] The I / O interface 201 receives various data from the outside and outputs various data to the outside. For example, the I / O interface 201 includes a network interface controller (NIC).
[0039] The HMI 202 is an interface that provides information to the user and receives information from the user. More specifically, the HMI 202 includes an input device and an output device. Examples of the input device include a touch panel, a keyboard, etc. Examples of the output device include a display, etc.
[0040] Processor 203 executes various processes. For example, processor 203 includes a CPU (Central Processing Unit). Storage device 204 stores various information necessary for processing. Examples of storage device 204 include volatile memory, non-volatile memory, HDD (Hard Disk Drive), SSD (Solid State Drive), and the like.
[0041] Processor 203 executes training data generation processing. In the training data generation processing, processor 203 acquires video data VID from video collection unit 100 via I / O interface 201. The video data VID is stored in storage device 204. Processor 203 automatically or almost automatically generates labeled training data LAD based on the video data VID. The labeled training data LAD is stored in storage device 204. Further, processor 203 outputs the labeled training data LAD to model training system 300 (see FIG. 2) via I / O interface 201.
[0042] Training data generation program 205 is a computer program that processor 203 executes to perform training data generation processing. Training data generation program 205 is stored in storage device 204. Training data generation program 205 may be recorded on a computer-readable recording medium. Training data generation program 205 may be provided via a network. The cooperation between processor 203 that executes training data generation program 205 and storage device 204 realizes the training data generation processing.
[0043] Hereinafter, some examples of the training data generation processing will be described.
[0044] 2-1. First example FIG. 6 is a block diagram showing a first example of the functional configuration of the training data generation system 200. The training data generation system 200 includes, as functional blocks, a video input unit 210, an object detection unit 220, a tracker 230, and a training data generation unit 240.
[0045] The video input unit 210 acquires video data VID via the I / O interface 201 or from the storage device 204. The video data VID includes a series of images IMG.
[0046] The object detection unit 220 detects moving objects in a series of images IMG. For example, "YOLOX" is used as the object detection unit 220. The bounding box BX represents the position of the detected moving object in the image IMG. The object detection unit 220 acquires information on the bounding box BX of each moving object in the series of images IMG.
[0047] The tracker 230 automatically tracks the same moving object in a series of images IMG based on a tracking algorithm. For example, ByteTrack (see Non-Patent Document 1 above) is used as the tracker 230. The tracker 230 tracks the same moving object in a series of images IMG based on the movement of the bounding box BX without performing feature extraction. More specifically, the tracker 230 tracks the same moving object in a series of images IMG by identifying a plurality of bounding boxes BXi[t] (t = t1, t2, t3, ···) that represent the same moving object in the series of images IMG. The tracker 230 associates a plurality of bounding boxes BXi[t] that represent the same moving object in a series of images IMG at different time steps with each other.
[0048] Track TRi is identification information indicating a plurality of bounding boxes BXi[t] that represent the same moving object in a series of images IMG at different time steps. In other words, track TRi is information representing the time series of the same moving object in a series of images IMG. Tracking result data TRD indicates track TRi in a series of images IMG. Tracker 230 generates tracking result data TRD by automatically tracking the same moving object.
[0049] Training data generation unit 240 automatically generates labeled training data LAD based on a series of images IMG and tracking result data TRD. More specifically, training data generation unit 240 generates labeled training data LAD by assigning track TRi as a label to a series of images IMG.
[0050] 2-2. Second example There is a possibility that two or more different tracks TRi are assigned to the same moving object. For example, FIG. 7 shows a situation where a certain moving object exits the camera's field of view and then re-enters the same camera's field of view. As a result, two different tracks TRa and TRb may be assigned to the same moving object. Such two or more different tracks TRi assigned to the same moving object are hereinafter referred to as "duplicate tracks".
[0051] The occurrence of duplicate tracks means that two or more different labels are assigned to the same moving object in the labeled training data LAD. If two or more different labels are assigned to the same moving object in the labeled training data LAD, the accuracy of model training may decrease. Therefore, it is desirable to detect duplicate tracks and integrate the duplicate tracks into a single track. For example, the duplicate tracks TRa and TRb shown in FIG. 7 are integrated into a single track TRc as shown in FIG. 8.
[0052] However, manual detection and integration of duplicate tracks require human labor and time. Therefore, the training data generation system 200 may be configured to automatically detect and integrate duplicate tracks. This process is hereinafter referred to as "track integration process".
[0053] FIG. 9 is a block diagram showing a second example of the functional configuration of the training data generation system 200. The training data generation system 200 further includes a track integration unit 250 in addition to the functional configuration described in the first example above. The track integration unit 250 performs a track integration process. That is, the track integration unit 250 automatically detects duplicate tracks based on the tracking result data TRD. When duplicate tracks are detected, the track integration unit 250 automatically integrates the detected duplicate tracks into a single track.
[0054] More specifically, the track integration unit 250 includes a feature extraction model MDL-X. For example, the feature extraction model MDL-X is an existing object recognition model. As another example, the feature extraction model MDL-X may be an object recognition model MDL that has been pre-trained. The track integration unit 250 inputs a series of images IMG to the feature extraction model MDL-X. The feature extraction model MDL-X extracts feature amounts of each moving object detected in the series of images IMG, and calculates the similarity between the detected moving objects based on the extracted feature amounts. The similarity is calculated based on the distance between the feature amounts in the embedding space. The smaller the distance in the embedding space, the higher the similarity.
[0055] The track integration unit 250 acquires the above-described tracking result data TRD. The track integration unit 250 checks whether there are overlapping tracks based on the similarity between the tracking result data TRD and the detected moving objects. If the similarity between the first moving object of the first track and the second moving object of the second track is higher than the threshold value, the track integration unit 250 determines that the first moving object and the second moving object are the same, and the first track and the second track are overlapping tracks. In this case, the track integration unit 250 integrates the first track and the second track into a single track.
[0056] When the track integration process is completed, the track integration unit 250 may present the result of the track integration process to a human checker via the HMI202. For example, the track integration unit 250 presents a series of images IMG and the track TRi modified by the track integration process to the human checker. For example, the track integration unit 250 may display the result of the track integration process on the display of the HMI202.
[0057] The human checker checks the result of the track integration process. For example, the human checker checks whether the automatically detected overlapping tracks are actually the overlapping tracks assigned to the same moving object. As another example, the human checker checks whether the detected overlapping tracks are correctly integrated into a single track. The human checker corrects the result of the track integration process using the HMI202 as needed.
[0058] After checking the result of the track integration process, the human checker approves the result of the track integration process. In response, the result of the track integration process is reflected in the tracking result data TRD. In other words, the result of the track integration process is fed back to the tracking result data TRD. Thereafter, the training data generation unit 240 generates labeled training data LAD based on the series of images IMG and the tracking result data TRD. Therefore, the result of the track integration process is reflected in the labeled training data LAD.
[0059] As described above, according to the second example, duplicate tracks for the same moving object are automatically detected and integrated into a single track. Since there are no more duplicate tracks, a decrease in the accuracy of model training is suppressed. Furthermore, manual work is reduced. Even if a human checker checks the result of the track integration process, the manual work is significantly reduced compared to the case where the human checker performs the track integration process manually.
[0060] 2-3. Third Example FIG. 10 is a block diagram showing a third example of the functional configuration of the training data generation system 200. In the third example, a human checker does not check the result of the track integration process. The result of the track integration process is directly reflected in the tracking result data TRD without being checked by a human. That is, the result of the track integration process is reflected in the labeled training data LAD without being checked by a human.
[0061] According to the third example, compared with the second example described above, manual work is further reduced. Note that errors in the track integration process are tolerated to some extent.
[0062] 3. Model Training System FIG. 11 is a block diagram showing a configuration example of the model training system 300 according to the present embodiment. The model training system 300 includes an I / O (Input / Output) interface 301, an HMI 302, one or more processors 303 (hereinafter simply referred to as "processor 303"), and one or more storage devices 304 (hereinafter simply referred to as "storage device 304").
[0063] The I / O interface 301 receives various data from the outside and outputs various data to the outside. For example, the I / O interface 301 includes a network interface controller (NIC).
[0064] HMI 302 is an interface that provides information to the user and also receives information from the user. More specifically, HMI 302 includes an input device and an output device. Examples of the input device include a touch panel, a keyboard, etc. Examples of the output device include a display, etc.
[0065] The processor 303 executes various processes. For example, the processor 303 includes a CPU. The storage device 304 stores various information necessary for the processes. Examples of the storage device 304 include a volatile memory, a non-volatile memory, an HDD, an SSD, etc.
[0066] The processor 303 executes a model training process. In the model training process, the processor 303 acquires labeled training data LAD via the I / O interface 301. The labeled training data LAD is stored in the storage device 304. The processor 303 trains an object recognition model MDL by using the labeled training data LAD. The trained object recognition model MDL is stored in the storage device 304. Also, the processor 303 outputs the trained object recognition model MDL to an object recognition system 400 (see FIG. 2) via the I / O interface 301.
[0067] The model training program 305 is a computer program that the processor 303 executes to perform the model training process. The model training program 305 is stored in the storage device 304. The model training program 305 may be recorded on a computer-readable recording medium. The model training program 305 may be provided via a network. The cooperation of the processor 303 and the storage device 304 that execute the model training program 305 realizes the model training process.
[0068] Hereinafter, some examples of the model training process will be described.
[0069] 3-1. First Example FIG. 12 is a block diagram showing a first example of the functional configuration of the model training system 300. The model training system 300 includes, as functional blocks, a training data input unit 310, a model input unit 320, and a model training unit 330.
[0070] The training data input unit 310 acquires labeled training data LAD via the I / O interface 301 or from the storage device 304.
[0071] The model input unit 320 acquires an object identification model MDL-O via the I / O interface 301 or from the storage device 304. The object identification model MDL-O is an object identification model before training.
[0072] The model training unit 330 trains the object identification model MDL-O based on the labeled training data LAD. In other words, the model training unit 330 trains the object identification model MDL-O by using the labeled training data LAD. Here, supervised learning or semi-supervised learning is used to train the object identification model MDL-O. As a result, a trained object identification model MDL is obtained.
[0073] 3-2. Second Example FIG. 13 is a block diagram showing a second example of the functional configuration of the model training system 300. The model training system 300 includes, as functional blocks, a training data input unit 310, a model input unit 320, a preliminary training unit 331, and a model training unit 332.
[0074] The preliminary training unit 331 performs pre-training of the object recognition model MDL-O by using an existing dataset. For example, the preliminary training unit 331 performs pre-training of the object recognition model MDL-O based on self-supervised learning. Self-supervised learning does not require labeled training data and only requires a bounding box. As a result of the preliminary training, an object recognition model MDL-P is obtained.
[0075] Note that the object recognition model MDL-P after preliminary training may be used as the feature extraction model MDL-X (see FIGS. 9 and 10) in the above-described track integration process.
[0076] The model training unit 332 further trains the object recognition model MDL-P after preliminary training based on the labeled training data LAD. Here, supervised learning or semi-supervised learning is used to train the object recognition model MDL-P after preliminary training. As a result, a high-precision object recognition model MDL is obtained.
Explanation of Signs
[0077] 100 Video collection unit 200 Training data generation system 201 I / O interface 202 HMI 203 Processor 204 Storage device 205 Training data generation program 210 Video input unit 220 Object detection unit 230 Tracker 240 Training data generation unit 250 Track integration unit 300 Model training system 301 I / O interface 302 HMI 303 Processor 304 Memory device 305 Model training program 310 Training data input section 320 Model input section 330 Model training section 331 Preliminary training section 332 Model training section 400 Object recognition system LAD Labeled training data MDL Object recognition model MDL-X Feature extraction model TRD Tracking result data VID Video data
Claims
1. A training data generation method for generating labeled training data used for supervised learning or semi-supervised learning of an object recognition model based on machine learning, comprising: detecting a moving object in a series of images; automatically obtaining a track, which is information representing a time series of the same moving object in the series of images, by tracking the same moving object in the series of images using a tracker; automatically generating the labeled training data by assigning the track as a label to the series of images; and a training data generation method.
2. The training data generation method according to claim 1, wherein a bounding box represents the position of the detected moving object in the series of images, and the tracker tracks the same moving object based on the movement of the bounding box without performing feature extraction. A training data generation method.
3. The training data generation method according to claim 2, wherein the tracker associates a plurality of bounding boxes representing the same moving object in the series of images with each other, and the track is information indicating the plurality of bounding boxes representing the same moving object in the series of images. A training data generation method.
4. The training data generation method according to any one of claims 1 to 3, further comprising a track integration process, wherein the track integration process includes: detecting two or more different tracks assigned to the same moving object; and integrating the two or more different tracks into a single track. and a training data generation method.
5. The training data generation method according to claim 4, wherein the track integration process includes: extracting feature amounts of each moving object detected in the series of images by inputting the series of images into a feature extraction model, and calculating a similarity between moving objects based on the extracted feature amounts; and when the similarity between a first moving object of a first track and a second moving object of a second track is higher than a threshold value, determining that the first moving object and the second moving object are the same, and integrating the first track and the second track into a single track. and a training data generation method.
6. The training data generation method according to claim 4, further comprising presenting the result of the track integration process to a human checker Training data generation method.
7. The training data generation method according to claim 4, wherein the result of the track integration process is reflected in the labeled training data without being subject to human checking Training data generation method.
8. The training data generation method according to claim 1, wherein the object recognition model is a person re-identification model Training data generation method.
9. A training data generation system for generating labeled training data for supervised learning or semi-supervised learning of an object recognition model based on machine learning, comprising one or more processors, wherein the one or more processors detect moving objects in a series of images, automatically obtain a track, which is information representing the time series of the same moving object in the series of images, by tracking the same moving object in the series of images using a tracker, automatically generate the labeled training data by assigning the track as a label to the series of images configured as Training data generation system.
Citation Information
Patent Citations
Transformer substation safety management and control platform based on deep learning
CN112381778A
Trajectory analysis apparatus and trajectory analysis method
JP2015201005A
Video analyzing device, person searching system and person searching method
JP2020013290A
Concept of generating training data and training machine learning model for use in re-identification
JP2022051683A
Object detection system based on deep learning or the like, and garbage collection vehicle using the same
JP2022066998A