Learning device, learning method, tracking device, tracking method, and recording medium
The learning device addresses the challenge of object association between frames with varying time intervals by extracting frame pairs, detecting objects, associating them, and learning the association method, resulting in improved accuracy and efficiency of object tracking.
Patent Information
- Application Number
- JP2024505688
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-08
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-08
AI Technical Summary
Existing technologies face challenges in accurately associating objects between frames in videos, especially with varying time intervals, which affects the efficiency and accuracy of object tracking and learning.
A learning device and method that acquire a video, extract pairs of frames with different time intervals, detect objects in each frame, associate objects between frames, and learn the object association method based on the association results, using a combination of forward and reverse association results to improve accuracy.
The solution enables accurate and efficient object association and tracking by learning from diverse frame pairs, improving the accuracy of object tracking and reducing computational costs compared to existing methods.
Smart Images

Figure 0007683812000001 
Figure 0007683812000002 
Figure 0007683812000003
Abstract
Description
Technical Field
[0001] This disclosure relates to the technical fields of learning devices, learning methods, tracking devices, tracking methods, and recording media.
Background Art
[0002] Techniques for acquiring environmental information regarding learning images and using the environmental information to perform learning of an object detection model for detecting target objects included in the learning images are described in Patent Document 1. Techniques for performing motion detection by acquiring a target image, deriving a vector related to motion from the acquired target image, and tracking the derived vector, and performing motion detection without increasing the calculation cost are described in Patent Document 2. Techniques for extracting a feature map characterizing the spatial structure of the space imaged in a frame from each of a plurality of frames, capturing a target object imaged in the frame based on each of the plurality of frames, and extracting a region mask indicating the region of the target object, and extracting, for each frame, a region feature representing the feature of the object candidate region based on the feature map, the object candidate region, and the region mask, and performing object association between frames using the plurality of region features extracted for each frame, and accurately associating the same object between frames are described in Patent Document 3. Techniques for acquiring learning data used for learning an image recognizer including a feature extractor and not including a generator, and using a first index used for supervised learning using the labeled images included in the acquired learning data, and regarding the relationship between the feature data output when each of two or more images acquired based on the images included in the learning data is input to the feature extractor, and a second index used for unsupervised learning, and using, without using a third index used for unsupervised learning, the relationship between the output data output when each of two or more images acquired based on the images included in the learning data is input to the image recognizer to learn the image recognizer are described in Patent Document 4.
Prior Art Documents
Patent Documents
[0003] [Patent Document 1] International Publication No. 2021 / 070324 [Patent Document 2] International Publication No. 2020 / 022362 [Patent Document 3] Japanese Patent Application Laid-Open No. 2020-181268 [Patent Document 4] Japanese Patent Application Laid-Open No. 2019-207561 [Summary of the Invention] [Problems to be Solved by the Invention]
[0004] This disclosure aims to provide a learning device, a learning method, a tracking device, a tracking method, and a recording medium for improving the technology described in prior art documents. [Means for Solving the Problems]
[0005] One aspect of the learning device includes an acquisition means for acquiring one video, an extraction means for extracting a plurality of pairs of a first frame and a second frame different from the first frame from a plurality of frames included in the video, a detection means for detecting each of the objects included in the first frame and the objects included in the second frame, a matching means for associating the objects included in the first frame with the objects included in the second frame, and a learning means for causing the matching means to learn the object matching method based on the matching results of the plurality of pairs by the matching means. The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval.
[0006] One aspect of the learning method is to obtain one video, extract a plurality of pairs of a first frame and a second frame different from the first frame from the plurality of frames included in the video, detect each of the objects included in the first frame and the objects included in the second frame, associate the objects included in the first frame with the objects included in the second frame using a corrector, and based on the association results by the corrector for the plurality of pairs, cause the corrector to learn the object association method. The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval.
[0007] The first aspect of the recording medium has a computer program recorded thereon for causing the computer to execute a learning method of obtaining one video, extracting a plurality of pairs of a first frame and a second frame different from the first frame from the plurality of frames included in the video, detecting each of the objects included in the first frame and the objects included in the second frame, associating the objects included in the first frame with the objects included in the second frame using a corrector, and based on the association results by the corrector for the plurality of pairs, causing the corrector to learn the object association method. The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval.
[0008] One aspect of the tracking device includes an acquisition unit that acquires a video, extracts a plurality of pairs of a first frame and a second frame different from the first frame from a plurality of frames included in one video, detects each of the objects included in the first frame and the objects included in the second frame, and has an association unit generated by causing learning of the object association method based on the association results of the plurality of pairs in which the objects included in the first frame and the objects included in the second frame are associated with each other, and a tracking unit that tracks the objects included in the video based on the association of the objects by the association unit. The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval.
[0009] One aspect of the tracking method is a tracking method that acquires a video, extracts a plurality of pairs of a first frame and a second frame different from the first frame from a plurality of frames included in one video, detects each of the objects included in the first frame and the objects included in the second frame, and has an association unit generated by causing learning of the object association method based on the association results of the plurality of pairs in which the objects included in the first frame and the objects included in the second frame are associated with each other, and tracks the objects included in the video based on the association of the objects by the association unit. The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval.
[0010] A second aspect of the recording medium causes a computer to acquire a video, extract a plurality of pairs of a first frame and a second frame different from the first frame from a plurality of frames included in one video, detect each of an object included in the first frame and an object included in the second frame, and perform learning of the object association method based on the association results of the plurality of pairs in which the object included in the first frame and the object included in the second frame are associated with each other, and has an association means generated by causing learning to be performed. A tracking method for tracking an object included in the video based on the association of the object by the association means, wherein the plurality of pairs include a first pair in which a time interval between the first and second frames is a first interval, and a second pair in which a time interval between the first and second frames is a second time interval different from the first interval. A computer program for executing the tracking method is recorded.
Brief Description of Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Embodiments for Carrying Out the Invention
[0012] Hereinafter, embodiments of a learning device, a learning method, a tracking device, a tracking method, and a recording medium will be described with reference to the drawings. [1: First Embodiment]
[0013] The first embodiment of the learning device, the learning method, and the recording medium will be described. Hereinafter, the first embodiment of the learning device, the learning method, and the recording medium will be described using the learning device 1 to which the first embodiment of the learning device, the learning method, and the recording medium is applied. [1-1: Configuration of Learning Device 1]
[0014] The configuration of the learning device 1 according to the first embodiment will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of the learning device 1 according to the first embodiment.
[0015] As shown in FIG. 1, the learning device 1 according to the first embodiment includes an acquisition unit 11, an extraction unit 12, a detection unit 13, an association unit 14, and a learning unit 15.
[0016] The acquisition unit 11 acquires one video MV. The extraction unit 12 extracts a plurality of pairs of a first frame and a second frame different from the first frame from a plurality of frames included in the video MV. The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval.
[0017] The detection unit 13 detects each of the objects included in the first frame and the objects included in the second frame. The association unit 14 associates the objects included in the first frame with the objects included in the second frame. The learning unit 15 causes the association unit 14 to learn the object association method based on the association results by a plurality of sets of the association unit 14. [1-2: Technical effect of the learning device 1]
[0018] The learning device 1 in the first embodiment extracts a plurality of pairs of first and second frames with various time intervals between the frames. Each of the plurality of pairs may be learning data used for the learning of the association. That is, the learning device 1 can prepare many pairs that can be used for the learning of the association. Since the learning device 1 learns the object association method using many pairs of the first and second frames, the association unit 14 that can accurately associate the objects can be obtained. [2: Second embodiment]
[0019] The second embodiment of the learning device, the learning method, and the recording medium will be described. Hereinafter, the second embodiment of the learning device, the learning method, and the recording medium will be described using the learning device 2 to which the second embodiment of the learning device, the learning method, and the recording medium is applied. [2-1: Configuration of the learning device 2]
[0020] With reference to FIG. 2, the configuration of the learning device 2 in the second embodiment will be described. FIG. 2 is a block diagram showing the configuration of the learning device 2 in the second embodiment.
[0021] As shown in FIG. 2, the learning device 2 includes an arithmetic device 21 and a storage device 22. Further, the learning device 2 may include a communication device 23, an input device 24, and an output device 25. However, the learning device 2 may not include at least one of the communication device 23, the input device 24, and the output device 25. The arithmetic device 21, the storage device 22, the communication device 23, the input device 24, and the output device 25 may be connected via a data bus 26.
[0022] The arithmetic device 21 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array). The arithmetic device 21 reads a computer program. For example, the arithmetic device 21 may read the computer program stored in the storage device 22. For example, the arithmetic device 21 may read the computer program stored in a computer-readable and non-transitory recording medium using a recording medium reading device (for example, the input device 24 described later) provided in the learning device 2. The arithmetic device 21 may acquire a computer program from a device (not shown) disposed outside the learning device 2 via the communication device 23 (or other communication device) (that is, it may be downloaded or read). The arithmetic device 21 executes the read computer program. As a result, logical functional blocks for executing the operations to be performed by the learning device 2 are realized within the arithmetic device 21. That is, the arithmetic device 21 can function as a controller for realizing logical functional blocks for executing the operations (in other words, processes) to be performed by the learning device 2.
[0023] FIG. 2 shows an example of the logical functional blocks realized within the arithmetic device 21 for executing the learning operation. As shown in FIG. 2, within the arithmetic device 21, there are realized an acquisition unit 211 which is a specific example of the "acquisition means" described in the appended note below, an extraction unit 212 which is a specific example of the "extraction means" described in the appended note below, a detection unit 213 which is a specific example of the "detection means" described in the appended note below, a correlation unit 214 which is a specific example of the "correlation means" described in the appended note below, and a learning unit 215 which is a specific example of the "learning means" described in the appended note below.
[0024] The storage device 22 can store desired data. For example, the storage device 22 may temporarily store a computer program executed by the arithmetic unit 21. The storage device 22 may temporarily store data that the arithmetic unit 21 temporarily uses when the arithmetic unit 21 is executing a computer program. The storage device 22 may store data that the learning device 2 stores long-term. Note that the storage device 22 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device. That is, the storage device 22 may include a non-temporary recording medium.
[0025] The storage device 22 may store a plurality of moving images MV. The moving image MV may be image data including a plurality of frames. The moving image MV may be used for a learning operation by the learning device 2. However, the storage device 22 does not necessarily have to store the moving image MV.
[0026] The communication device 23 can communicate with a device external to the learning device 2 via a communication network (not shown).
[0027] The input device 24 is a device that receives input of information to the learning device 2 from the outside of the learning device 2. For example, the input device 24 may include an operating device (for example, at least one of a keyboard, a mouse, and a touch panel) that can be operated by an operator of the learning device 2. For example, the input device 24 may include a reading device that can read information recorded as data on an externally attachable recording medium with respect to the learning device 2.
[0028] The output device 25 is a device that outputs information to the outside of the learning device 2. For example, the output device 25 may output information as an image. That is, the output device 25 may include a display device (so-called, display) capable of displaying an image indicating the information to be output. For example, the output device 25 may output information as voice. That is, the output device 25 may include a voice device (so-called, speaker) capable of outputting voice. For example, the output device 25 may output information on paper. That is, the output device 25 may include a printing device (so-called, printer) capable of printing desired information on paper. [2-2: Learning operation performed by learning device 2]
[0029] Referring to FIG. 3, the flow of the learning operation performed by the learning device 2 in the second embodiment will be described. FIG. 3 is a flowchart showing the flow of the learning operation performed by the learning device 2 in the second embodiment. The learning operation performed by the learning device 2 may be an operation performed offline.
[0030] As shown in FIG. 3, the acquisition unit 211 acquires one video MV (step S20). The extraction unit 212 extracts a pair group from a plurality of frames included in the video MV (step S21). One pair may include a first frame and a second frame different from the first frame. The extraction unit 212 extracts a plurality of pairs of a first frame and a second frame different from the first frame from the plurality of frames included in the video MV. The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval. Here, the first frame included in the first pair and the first frame included in the second pair may be different frames. Also, the second frame included in the first pair and the second frame included in the second pair may be different frames. The extraction unit 212 may select any two frames from all the frames included in one video MV to generate a pair. The extraction unit 212 may extract a batch number of pairs using all the frames included in one video MV. The batch number may be, for example, 1024 or more. The batch number is not particularly limited, and any value can be used.
[0031] The extraction unit 212 selects one pair from the extracted pair group (step S22). The extraction unit 212 selects a pair of the first and second frames from among the plurality of extracted pairs.
[0032] The detection unit 213 detects each of the objects included in the first frame and the objects included in the second frame (step S23). The association unit 214 associates the objects included in the first frame and the objects included in the second frame using a learnable association model MM (step S24).
[0033] The association model MM may be a model capable of outputting information regarding the association result between the object included in the first frame and the object included in the second frame when, for example, information regarding the object included in the first frame and information regarding the object included in the second frame are input. The association model MM is typically a model using a neural network, but may also be a model different from the model using a neural network.
[0034] Alternatively, the association model MM may be a model capable of outputting information regarding the association result between the object included in the first frame and the object included in the second frame when, for example, the first frame and the second frame are input. That is, the association model MM may be a model that detects each of the object included in the first frame and the object included in the second frame, and associates the object included in the first frame with the object included in the second frame. In this case, the detection unit 213 may detect each of the object included in the first frame and the object included in the second frame using the association model MM, and the association unit 214 may associate the object included in the first frame with the object included in the second frame using the learnable association model MM. Alternatively, the arithmetic device 21 may include a logical processing block in which the detection unit 213 and the association unit 214 are integrated.
[0035] The extraction unit 212 determines whether there is a pair among the extracted pairs of the plurality of first and second frames for which the processing from step S22 to step S24 is unprocessed (step S25). If there is an unprocessed pair (step S25: Yes), the process proceeds to step S22.
[0036] When the processing from step S22 to step S24 is performed for all pairs (step S25: No), the learning unit 215 causes the association unit 214 to learn the object association method based on the association results by the association unit 214 for the plurality of pairs (step S26).
[0037] Specifically, the learning unit 215 may cause the association model MM used by the association unit 214 to learn the object association method and construct the association model MM. More specifically, the learning unit 215 may adjust the parameters that define the operation of the association model MM. When the association model MM is a neural network, the parameters that define the operation of the association model MM may include at least one of the weights and biases of the neural network. The learning unit 215 may obtain one video MV and update the parameters that define the operation of the association model MM based on the association results of a batch number of pairs. The parameters that define the operation of the association model MM may be, for example, the weights and biases of the neural network. The parameters that define the operation of the association model MM may be stored in the storage device 22.
[0038] The learning unit 215 may calculate a learning loss based on the association result and cause the association model MM to learn the association method from the learning loss. The learning unit 215 may cause the association model MM to perform contrastive learning. The learning unit 215 may calculate a loss function such as the cross-entropy loss of the object between frames and cause the association model MM to learn the association method so that the contrastive loss becomes small (typically, minimum).
[0039] The learning device 2 may construct an association model MM that can be used in online multi-object tracking through a learning operation. [2-3: Extraction example of the first pair]
[0040] FIG. 4 illustrates an extraction example of the first pair by the extraction unit 212. The extraction unit 212 may randomly select a first frame and a second frame from a plurality of frames to extract a pair.
[0041] Figure 4 shows frames [1] to
[10] included in a single video MV. Frames [1] to
[10] may be consecutive frames. For example, the frame following frame [1] may be frame [2], and the frame following frame [2] may be frame [3].
[0042] For example, as shown in FIG. 4, the extraction unit 212 may randomly select frame [1] as the first frame, randomly select frame [3] as the second frame, and randomly extract the pair P1. Also, the extraction unit 212 may randomly select frame [2] as the first frame, randomly select frame [6] as the second frame, and randomly extract the pair P2. Also, the extraction unit 212 may randomly select frame [7] as the first frame, randomly select frame
[10] as the second frame, and randomly extract the pair P3. Also, the extraction unit 212 may randomly select frame [4] as the first frame, randomly select frame
[10] as the second frame, and randomly extract the pair P4.
[0043] Also, for example, a pair of frame [1] and frame [3] may be referred to as a "forward pair", and a pair of frame [3] and frame [1] may be referred to as a "reverse pair", and each may be distinguished as a different pair. [2-4: Extraction example of the second pair]
[0044] FIG. 5 illustrates an extraction example of the second pair by the extraction unit 212. The extraction unit 212 may select a frame before or after a predetermined number of frames from the first frame as the second frame and extract a set.
[0045] FIG. 5 also shows frames [1] to
[10] included in a single video MV. Frames [1] to
[10] may be consecutive frames.
[0046] As shown on the left side of frames [1] to
[10] shown in FIG. 5, a frame two frames before or after the first frame may be selected as the second frame, and a pair may be extracted. In this way, selecting a frame two frames before or after as the second frame may be referred to as "selecting frames with a skip width of 1". Also, as shown on the right side of frames [1] to
[10] shown in FIG. 5, a frame three frames before or after the first frame may be selected as the second frame, and a pair may be extracted. In this way, selecting a frame three frames before or after as the second frame may be referred to as "selecting frames with a skip width of 2". In FIG. 5, the skip widths of 1 and 2 are taken as examples for explanation, but it may be any number of skip widths. The skip width of the frame may be determined automatically or specified manually.
[0047] Note that, as shown in FIG. 6, the object detection operation by the detection unit 213 in step S23 may be performed before the extraction of the pair by the extraction unit 212. The extraction unit 212 may select the same frame and extract different pairs. For this reason, for the frames included in one video MV acquired by the acquisition unit 211, the detection unit 213 may perform object detection before the extraction of the pair by the extraction unit 212. [2-5: Technical effects of the learning device 2]
[0048] In the learning device 2 according to the second embodiment, since the first frame and the second frame are randomly selected from a plurality of frames to extract a pair, and / or a frame a predetermined number of frames before or after one frame is selected as the second frame to extract a pair, various combinations of frames can be generated more easily. Since the learning device 2 uses various combinations of frames, it can learn the association method accurately and efficiently, and can construct an association model MM that can perform accurate association. Also, since the learning device 2 calculates a learning loss based on the association result and causes the association unit 214 to learn the association method from the learning loss, the accuracy of the association can be improved.
[0049] For example, since the learning device of Comparative Example 1 that learns online using a small number of frames, such as about 1 to 10 frames, performs learning using a small number of pairs, the accuracy of the association between objects is likely to decrease. Since the learning device 2 in the second embodiment learns using all the frames included in the video, the accuracy of the association between objects is higher compared to the learning device of Comparative Example 1.
[0050] Also, the learning device of Comparative Example 2 that converts all the frames included in a plurality of videos into batches and performs learning offline arranges the plurality of videos in an array, so it is necessary to select and align the number of frames of the videos. For this reason, the computational cost is high. Also, the association model learned by the learning device of Comparative Example 2 operates offline using all the frames included in the video. In contrast, since the learning device 2 in the second embodiment only needs to convert the frames included in one video into batches, the computational cost can be reduced. In contrast, the association model MM in the second embodiment does not have processing that depends on the number of frames included in one video.
[0051] Also, the association model MM constructed by the learning device 2 in the second embodiment can improve the accuracy of the online object tracking model. [3: Third Embodiment]
[0052] The third embodiment of the learning device, the learning method, and the recording medium will be described. Hereinafter, the third embodiment of the learning device, the learning method, and the recording medium will be described using the learning device 3 to which the third embodiment of the learning device, the learning method, and the recording medium is applied.
[0053] The learning device 3 in the third embodiment includes an arithmetic device 21 and a storage device 22, similar to the learning device 2 in the second embodiment. Further, the learning device 3 may also include a communication device 23, an input device 24, and an output device 25, similar to the learning device 2. However, the learning device 3 may not include at least one of the communication device 23, the input device 24, and the output device 25. The learning operation performed by the learning unit 215 of the learning device 3 in the third embodiment is different from that of the learning device 2 in the second embodiment. Other features of the learning device 3 may be the same as those of the learning device 2. [3-1: Learning operation performed by learning device 3]
[0054] FIG. 7 is a flowchart showing the flow of the learning operation performed by the learning device 3 in the third embodiment. As shown in FIG. 7, the acquisition unit 211 acquires one video MV (step S20).
[0055] The detection unit 213 detects each object included in the frames included in the one video MV acquired by the acquisition unit 211. The detection unit 213 may detect objects in the forward direction, for example (step S30). The detection unit 213 may detect objects in the order of frame capture of the plurality of frames included in the video MV. For example, when the video MV includes frames from frame [1] to frame
[10] , the detection unit 213 may first detect the objects included in frame [1], then detect the objects included in frame [2], then detect the objects included in frame [3],... and finally detect the objects included in frame
[10] .
[0056] Note that the detection unit 213 may detect objects in the reverse direction instead of detecting objects in the forward direction. In this case, the detection unit 213 may detect objects in the reverse order of frame capture of the plurality of frames included in the video MV. For example, when the video MV includes frames from frame [1] to frame
[10] , the detection unit 213 may first detect the objects included in frame
[10] , then detect the objects included in frame [9], then detect the objects included in frame [8],... and finally detect the objects included in frame [1].
[0057] The extraction unit 212 extracts a plurality of pairs of a first frame and a second frame different from the first frame from a plurality of frames included in the moving image MV (step S21). [3-2: Example of pair creation]
[0058] Referring to FIG. 8, the flow of the pair extraction operation performed by the learning device 3 in the third embodiment will be described. As shown in FIG. 8, in the third embodiment, for example, a pair of frame [1] and frame [3] may be extracted as the first forward pair P1F, and a pair of frame [3] and frame [1] may be extracted as the first backward pair P1B. The first forward pair P1F and the first backward pair P1B may be distinguished as different pairs. Similarly, for example, a pair of frame [4] and frame [8] may be extracted as the second forward pair P2F, and a pair of frame [8] and frame [4] may be extracted as the second backward pair P2B. The second forward pair P2F and the second backward pair P2B may be distinguished as different pairs. Further, for example, a pair of frame [7] and frame
[10] may be extracted as the third forward pair P3F, and a pair of frame
[10] and frame [7] may be extracted as the third backward pair P3B. The third forward pair P3F and the third backward pair P3B may be distinguished as different pairs.
[0059] The extraction unit 212 selects one pair from the plurality of extracted pairs of the first and second frames (step S22). The extraction unit 212 selects one pair of the first and second frames from the extracted pair group. The association unit 214 performs forward association (step S24F). The association unit 214 associates the object included in the second frame with the object included in the first frame. The association unit 214 performs backward association (step S24B). The association unit 214 associates the object included in the first frame with the object included in the second frame.
[0060] The extraction unit 212 determines whether there is a pair in the extracted pair group for which the processing in step S22 and steps S24F and S24B has not been performed (step S25). If there is an unprocessed pair (step S25: Yes), the process proceeds to step S22.
[0061] When the processing in step S22 and step S24 has been performed for all pairs (step S25: No), the learning unit 215 causes the association unit 214 to learn the object association method based on the association results by the plurality of sets of association units 14 (step S26).
[0062] The learning unit 215 may cause the association unit 214 to learn the object association method based on the forward association result in which an object included in the first frame is associated with an object included in the second frame, which is a frame after the first frame, and the backward association result in which an object included in the second frame is associated with an object included in the first frame, which is a frame before the second frame.
[0063] FIG. 9 is a conceptual diagram of the forward association result and the backward association result. FIG. 9(a) shows the forward association result in which an object included in frame [3] as the second frame, which is a frame after frame [1] as the first frame, is associated with an object included in frame [1]. FIG. 9(b) shows the backward association result in which an object included in frame [1] as the first frame, which is a frame before frame [3] as the second frame, is associated with an object included in frame [3].
[0064] In the case illustrated in FIG. 9, each frame includes two types of objects. FIG. 9(a) exemplifying the forward association result shows that the association unit 214 has associated object A in frame [3] with object A in frame [1]. Further, FIG. 9(a) shows that the association unit 214 has associated object B in frame [3] with object B in frame [1].
[0065] On the other hand, FIG. 9(b) exemplifying the reverse association result shows that the association unit 214 associates object A in frame [1] with object A in frame [3]. Further, FIG. 9(b) shows that the association unit 214 associates object B in frame [1] with object B in frame [3].
[0066] Comparing the forward association result exemplified in FIG. 9(a) with the reverse association result exemplified in FIG. 9(b), the association unit 214 performs the same association in both the forward association and the reverse association result.
[0067] Also, FIG. 9(c) shows the forward association result of associating the objects included in frame [8] as the second frame, which is a frame after frame [4] as the first frame, with the objects included in frame [4]. FIG. 9(d) shows the reverse association result of associating the objects included in frame [4] as the first frame, which is a frame before frame [8] as the second frame, with the objects included in frame [8].
[0068] FIG. 9(c) exemplifying the forward association result shows that the association unit 214 associates object A in frame [8] with object A in frame [4]. Further, FIG. 9(c) shows that the association unit 214 associates object B in frame [8] with object B in frame [4].
[0069] On the other hand, FIG. 9(d) exemplifying the reverse association result shows that the association unit 214 associates object A in frame [4] with object A in frame [8]. Further, FIG. 9(d) shows that the association unit 214 associates object B in frame [4] with object B in frame [8].
[0070] Comparing the forward association result exemplified in FIG. 9(c) with the reverse association result exemplified in FIG. 9(d), the association unit 214 performs different associations in the forward association and the reverse association.
[0071] The learning unit 215 may cause the association unit 214 to learn the object association method based on a loss function in which the loss increases as the forward association result and the reverse association result become less similar. For example, in the case shown in FIG. 9, since the forward association result shown in FIG. 9(a) and the reverse association result shown in FIG. 9(b) are similar, the loss of the loss function may be small. Also, since the forward association result shown in FIG. 9(c) and the reverse association result shown in FIG. 9(d) are not similar, the loss of the loss function may be large.
[0072] The learning device 3 may perform object association in the forward and reverse directions and learn so that the error between the association results in both directions becomes small. That is, the learning device 3 may perform unsupervised learning. [3-3: Technical effects of the learning device 3]
[0073] The learning device 3 in the third embodiment causes the association unit 214 to learn the object association method based on the forward association result in which the object included in the second frame is associated with the object included in the first frame, which is a frame before the second frame, and the reverse association result in which the object included in the first frame is associated with the object included in the second frame, which is a frame after the first frame. Therefore, learning can be performed without preparing correct data. That is, the learning device 3 can use an unsupervised learning algorithm.
[0074] Since the learning device 3 first performs detection processing for each frame, it is possible to reduce the detection processing for the overlapping portion of the frames in the extracted pair group, and the calculation cost can be reduced. Also, since the learning device 3 adds pairs associated in the reverse direction, the number of pairs that can be used for learning can be efficiently increased.
[0075] In addition, the learning device 3 causes the association unit 214 to learn the object association method based on a loss function in which the loss increases as the forward association result and the reverse association result become less similar, so that the accuracy of object association can be improved. [4: Fourth Embodiment]
[0076] The fourth embodiment of the learning device, the learning method, and the recording medium will be described. Hereinafter, the fourth embodiment of the learning device, the learning method, and the recording medium will be described using the learning device 4 to which the fourth embodiment of the learning device, the learning method, and the recording medium is applied.
[0077] The learning device 4 in the fourth embodiment includes an arithmetic device 21 and a storage device 22, similar to the learning device 3 in the third embodiment. Further, the learning device 4 may include a communication device 23, an input device 24, and an output device 25, similar to the learning device 3. However, the learning device 4 may not include at least one of the communication device 23, the input device 24, and the output device 25. The learning device 4 in the fourth embodiment has different detection operations performed by the detection unit 213 and learning operations performed by the learning unit 215 compared to the learning device 3 in the third embodiment. Other features of the learning device 4 may be the same as other features of the learning device 3. [4-1: Learning Operations Performed by the Learning Device 4]
[0078] In the fourth embodiment, in steps S30 and S31 shown in FIG. 7, the detection unit 213 detects each of the object included in the first frame, the position of the object, the object included in the second frame, and the position of the object. The detection unit 213 may detect each of the object included in the frame included in the single moving image MV acquired by the acquisition unit 211 and the position of the object. The detection unit 213 may detect each of the object and the position of the object, for example, in the forward direction (step S30). Note that the detection unit 213 may detect each of the object and the position of the object in the reverse direction.
[0079] Also, in the fourth embodiment, in step S32 shown in FIG. 7, the learning unit 215 causes the association unit 214 to learn the object association method based on at least one of the comparison result between the position where the first object is included in the first frame in the forward association result and the position where the first object is included in the first frame in the reverse association result, and the comparison result between the position where the second object is included in the second frame in the forward association result and the position where the second object is included in the second frame in the reverse association result.
[0080] FIG. 10 is a conceptual diagram of the forward association result and the reverse association result, similar to FIG. 9. FIG. 10(a) shows the forward association result in which the object included in frame [1] is associated with the object included in frame [3], which is a frame after frame [1]. FIG. 10(b) shows the reverse association result in which the object included in frame [3] is associated with the object included in frame [1], which is a frame before frame [3].
[0081] FIG. 10(a) exemplifying the forward association result illustrates that the association unit 214 associates object A in frame [3] with object A in frame [1]. Also, FIG. 10(a) illustrates that the association unit 214 associates object B in frame [3] with object B in frame [1].
[0082] On the other hand, FIG. 10(b) exemplifying the reverse association result illustrates that the association unit 214 associates object A in frame [1] with object A in frame [3]. Also, FIG. 10(b) illustrates that the association unit 214 associates object B in frame [1] with object B in frame [3].
[0083] As illustrated in FIG. 10(c), the learning unit 215 may cause the association unit 214 to learn the object association method based on the comparison result between the positions where the objects A and B as the first objects are included in the frame [1] of the forward pair P1F in the association result of the forward pair P1F and the positions where the objects A and B are included in the frame [1] of the reverse pair P1B in the reverse association result.
[0084] The learning unit 215 determines whether the forward association result by the association unit 214 matches the reverse association result. For example, as illustrated in FIG. 10(c), the learning unit 215 determines whether the forward association result by the association unit 214 matches the reverse association result based on whether the positions where the objects A and B are included in the frame [1] of the forward pair P1F are the same as the positions where the objects A and B are included in the frame [1] of the reverse pair P1B. In the case illustrated in FIG. 10(c), since the positions where the objects A and B are included in the frame [1] of the forward pair P1F are the same as the positions where the objects A and B are included in the frame [1] of the reverse pair P1B, the learning unit 215 may determine that the association by the association unit 214 is successful.
[0085] FIG. 10(d) shows the forward association result in which the objects included in the frame [4] are associated with the objects included in the frame [8] which is a frame after the frame [4]. FIG. 10(e) shows the reverse association result in which the objects included in the frame [8] are associated with the objects included in the frame [4] which is a frame before the frame [8].
[0086] FIG. 10(d) exemplifying the forward association result illustrates that the association unit 214 associates the object A in the frame [8] with the object A in the frame [4]. Further, FIG. 10(d) illustrates that the association unit 214 associates the object B in the frame [8] with the object B in the frame [4].
[0087] On the other hand, FIG. 10(e) exemplifying the reverse correspondence result illustrates that the correspondence unit 214 has associated object A in frame [8] with object A in frame [4]. Further, FIG. 10(e) illustrates that the correspondence unit 214 has associated object B in frame [8] with object B in frame [4].
[0088] As illustrated in FIG. 10(f), the learning unit 215 may cause the correspondence unit 214 to learn the object correspondence method based on the comparison result between the positions where object A and object B are included in frame [4] of the forward pair P2F and the positions where object A and object B are included in frame [4] of the reverse pair P2B in the correspondence result of the forward pair P2F.
[0089] For example, as illustrated in FIG. 10(f), the learning unit 215 may determine whether the result of forward correspondence by the correspondence unit 214 matches the result of reverse correspondence based on whether the positions where object A and object B are included in frame [4] of the forward pair P2F are the same as the positions where object A and object B are included in frame [4] of the reverse pair P2B. In the case illustrated in FIG. 10(f), since the positions where object A and object B are included in frame [4] of the forward pair P2F are different from the positions where object A and object B are included in frame [4] of the reverse pair P2B, the learning unit 215 may determine that the correspondence by the correspondence unit 214 has failed.
[0090] The learning unit 215 may cause the correspondence unit 214 to learn the object correspondence method so that the overlap between the position where the first object is included in the first frame in the forward correspondence result and the position where the first object is included in the first frame in the reverse correspondence result becomes large. [4-3: Technical Effects of Learning Device 4]
[0091] In the learning device 4 according to the fourth embodiment, in the forward association result, the position where the first object is included in the first frame and, in the reverse association result, the comparison result of the position where the first object is included in the first frame, and in the forward association result, the position where the second object is included in the second frame and, in the reverse association result, the comparison result of the position where the second object is included in the second frame are used. Based on at least one of these, the association unit 214 is made to learn the object association method, so that learning can be performed without preparing correct data. That is, the learning device 4 can use an unsupervised learning algorithm. Further, since the learning device 4 performs association using the information on the position of the object, it can perform the association of the object with higher accuracy compared to the case where the information on the position of the object is not used.
[0092] In addition to the effects of the learning device 3, the learning device 4 according to the fourth embodiment can reflect the consistency of the tracking result in learning from the degree of coincidence between the forward and reverse association results, so that the accuracy of the association can be further improved. [5: Fifth Embodiment]
[0093] The fifth embodiment of the learning device, learning method, and recording medium will be described. Hereinafter, the fifth embodiment of the learning device, learning method, and recording medium will be described using the learning device 5 to which the fifth embodiment of the learning device, learning method, and recording medium is applied.
[0094] The learning device 5 in the fifth embodiment includes an arithmetic device 21 and a storage device 22, similar to at least one of the learning devices 2 to 4 in the second to fourth embodiments. Further, the learning device 5 may include a communication device 23, an input device 24, and an output device 25, similar to at least one of the learning devices 2 to 4 in the second to fourth embodiments. However, the learning device 5 may not include at least one of the communication device 23, the input device 24, and the output device 25. The learning device 5 in the fifth embodiment differs from the learning device 4 in at least one of the learning devices 2 to 4 in the second to fourth embodiments in terms of the information included in the moving image MV acquired by the acquisition unit 211 and the learning operation performed by the learning unit 215. Other features of the learning device 5 may be the same as those of at least one of the learning devices 2 to 4. [5-1: Learning operation performed by learning device 5]
[0095] In the fifth embodiment, in step S20 shown in FIG. 6, the acquisition unit 211 acquires, as the moving image MV, learning information including a sample moving image and correct labels indicating which object each of the sample objects included in each of the plurality of sample frames included in the sample moving image is.
[0096] In step S23, the detection unit 213 detects each of the sample objects included in the first sample frame among the plurality of sample frames and the sample objects included in the second sample frame among the plurality of sample frames.
[0097] In step S24, the association unit 214 associates the sample object included in the first sample frame with the sample object included in the second sample frame.
[0098] The learning unit 215 causes the association unit 214 to learn the object association method based on the association result by the association unit 214 with the correct label. That is, the learning device 5 in the fifth embodiment performs supervised learning. The learning unit 215 may cause the association unit 214 to learn the object association method based on a loss function in which the loss increases as the association result based on the correct label and the association result by the association unit 214 become less similar.
[0099] FIG. 11 is a conceptual diagram illustrating the learning operation performed by the learning device 5 in the fifth embodiment. For example, assume that the one pair selected in step S22 is a pair of frame [6] and frame [3].
[0100] FIG. 11(c) illustrates the object included in frame [6] and the correct label of the object. Further, FIG. 11(d) illustrates the object included in frame [3] and the correct label of the object. As shown in FIGS. 11(c) and (d), the correct label of the round object is "a", and the correct label of the square object is "b". That is, the association unit 214 may determine that the association is successful when the round object included in frame [6] and the round object included in frame [3] are associated as illustrated by the dashed arrow between FIGS. 11(a) and 11(b).
[0101] For example, as illustrated by the solid arrow between FIGS. 11(a) and 11(b), assume that the association unit 214 associates the sample object A included in frame [6] as the first sample frame with the sample object A included in frame [3] as the second sample frame. Since the association result illustrated by the dashed arrow and the association result illustrated by the solid arrow are not similar, the loss of the loss function used by the learning unit 215 may increase. [5-3: Technical Effect of Learning Device 5]
[0102] The learning device 5 in the fifth embodiment causes the association unit 214 to learn the object association method by supervised learning, so that the learning accuracy can be improved. [6: Sixth Embodiment]
[0103] The sixth embodiment of the tracking device, the tracking method, and the recording medium will be described. Hereinafter, the sixth embodiment of the tracking device, the tracking method, and the recording medium will be described using the tracking device 6 to which the sixth embodiment of the tracking device, the tracking method, and the recording medium is applied. [6-1: Configuration of Tracking Device 6]
[0104] With reference to FIG. 12, the configuration of the tracking device 6 in the sixth embodiment will be described. FIG. 12 is a block diagram showing the configuration of the tracking device 6 in the sixth embodiment.
[0105] As shown in FIG. 12, the tracking device 6 includes an arithmetic device 61 and a storage device 62. Further, the tracking device 6 may include a communication device 63, an input device 64, and an output device 65. However, the tracking device 6 may not include at least one of the communication device 63, the input device 64, and the output device 65. The arithmetic device 61, the storage device 62, the communication device 63, the input device 64, and the output device 65 may be connected via a data bus 66.
[0106] The arithmetic unit 61 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array). The arithmetic unit 61 reads a computer program. For example, the arithmetic unit 61 may read the computer program stored in the storage device 62. For example, the arithmetic unit 61 may read the computer program stored in a computer-readable and non-transitory recording medium using a recording medium reading device (for example, the input device 64 described later) provided in the tracking device 6. The arithmetic unit 61 may obtain a computer program from a device (not shown) disposed outside the tracking device 6 via the communication device 63 (or other communication device) (that is, it may be downloaded or read). The arithmetic unit 61 executes the read computer program. As a result, logical functional blocks for executing the operations to be performed by the tracking device 6 are realized within the arithmetic unit 61. That is, the arithmetic unit 61 can function as a controller for realizing logical functional blocks for executing the operations (in other words, processes) to be performed by the tracking device 6.
[0107] FIG. 12 shows an example of the logical functional blocks realized in the arithmetic unit 61 for executing the learning operation. As shown in FIG. 11, in the arithmetic unit 61, an acquisition unit 211, which is a specific example of the “acquisition means” described in the appendix below, and an extraction unit 212, which is a specific example of the “extraction means” described in the appendix below, are realized.
[0108] The storage device 62 can store desired data. For example, the storage device 62 may temporarily store a computer program executed by the arithmetic device 61. The storage device 62 may temporarily store data that the arithmetic device 61 temporarily uses when the arithmetic device 61 is executing a computer program. The storage device 62 may store data that the tracking device 6 stores long-term. Note that the storage device 62 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device. That is, the storage device 62 may include a non-temporary recording medium.
[0109] The storage device 62 may store parameters that define the operation of the association model MM. The association model MM may be an association model MM constructed by at least one of the learning device 2 in the second embodiment to the learning device 5 in the fifth embodiment. However, the storage device 62 may not store parameters that define the operation of the association model MM.
[0110] The communication device 63 can communicate with a device outside the tracking device 6 via a communication network (not shown). The communication device 63 may acquire a moving image MV used for the tracking operation from an imaging device via the communication network.
[0111] The input device 64 is a device that receives input of information to the tracking device 6 from outside the tracking device 6. For example, the input device 64 may include an operating device (e.g., at least one of a keyboard, a mouse, and a touch panel) that can be operated by an operator of the tracking device 6. For example, the input device 64 may include a reading device that can read information recorded as data on an externally attachable recording medium with respect to the tracking device 6.
[0112] The output device 65 is a device that outputs information to the outside of the tracking device 6. For example, the output device 65 may output the information as an image. That is, the output device 65 may include a display device (so-called, display) capable of displaying an image indicating the information to be output. For example, the output device 65 may output the information as sound. That is, the output device 65 may include a sound device (so-called, speaker) capable of outputting sound. For example, the output device 65 may output the information on paper. That is, the output device 65 may include a printing device (so-called, printer) capable of printing desired information on paper. [6-2: Tracking operation performed by the tracking device 6]
[0113] Referring to FIG. 13, the flow of the tracking operation performed by the tracking device 6 in the sixth embodiment will be described. FIG. 13 is a flowchart showing the flow of the tracking operation performed by the tracking device 6 in the sixth embodiment. The tracking device 6 in the sixth embodiment may track an object online.
[0114] As shown in FIG. 12, the acquisition unit 611 acquires the moving image MV (step S60). The acquisition unit 611 may acquire the moving image MV frame by frame.
[0115] The tracking unit 616 tracks the object included in the moving image MV (step S61). The tracking unit 616 may track a plurality of objects included in the moving image MV.
[0116] The tracking unit 616 may have a correspondence model MM constructed by causing learning of an object correspondence method. The tracking unit 616 may track an object based on the correspondence of the objects included in each frame included in the moving image MV by the correspondence model MM. The correspondence model MM may be the correspondence model MM constructed by at least one of the learning device 2 in the second embodiment to the learning device 5 in the fifth embodiment as described above.
[0117] The tracking device 6 in the sixth embodiment can be applied to a scene of tracking a person, particularly a scene of biometric authentication of a moving person. [6-3: Technical Effects of Tracking Device 6]
[0118] Since the tracking device 6 in the sixth embodiment performs tracking using the well-trained association model MM, it can accurately track an object. [7: Supplementary Note]
[0119] Regarding the embodiments described above, the following supplementary notes are further disclosed. [Supplementary Note 1] An acquisition means for acquiring one video, An extraction means for extracting a plurality of pairs of a first frame and a second frame different from the first frame from a plurality of frames included in the video, A detection means for detecting each of the object included in the first frame and the object included in the second frame, An association means for associating the object included in the first frame with the object included in the second frame, A learning means for causing the association means to learn the association method of the object based on the association results of the plurality of pairs by the association means and comprising, The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval Learning device. [Supplementary Note 2] The extraction means randomly selects the first frame and the second frame from the plurality of frames to extract the pair The learning device according to Supplementary Note 1. [Supplementary Note 3] The extraction means selects a frame before or after a predetermined number of frames from the first frame as the second frame to extract the pair The learning device according to Supplementary Note 1 or 2. [Supplementary Note 4] The learning means calculates a learning loss based on the association result, and causes the association means to learn the association method from the learning loss The learning device according to any one of Supplementary Notes 1 to 3. [Supplementary Note 5] The learning means causes the association means to learn the method of associating the object included in the first frame with the object included in the second frame which is a frame subsequent to the first frame, based on the forward association result in which the object included in the first frame is associated with the object included in the second frame, and the backward association result in which the object included in the second frame is associated with the object included in the first frame which is a frame prior to the second frame. The learning device according to any one of Supplementary Notes 1 to 4. [Supplementary Note 6] The learning means causes the association means to learn the method of associating the object, based on a loss function in which the loss increases as the forward association result and the backward association result become less similar. The learning device according to Supplementary Note 5. [Supplementary Note 7] The detection means detects each of the object included in the first frame, the position of the object, the object included in the second frame, and the position of the object. The learning means causes the association means to learn the method of associating the object, based on at least one of the comparison result between the position where the first object is included in the first frame in the forward association result and the position where the first object is included in the first frame in the backward association result, and the comparison result between the position where the second object is included in the second frame in the forward association result and the position where the second object is included in the second frame in the backward association result. The learning device according to Supplementary Note 5 or 6. [Supplementary Note 8] The acquisition means acquires learning information including a sample video and a correct label indicating which object each sample object included in each of the plurality of sample frames included in the sample video is. The detection means detects each of the sample object included in the first sample frame among the plurality of sample frames and the sample object included in the second sample frame among the plurality of sample frames. The correspondence means associates the sample object included in the first sample frame with the sample object included in the second sample frame. The learning means causes the correspondence means to learn the method of associating the objects based on the correct label and the association result by the correspondence means. The learning device according to any one of Appendices 1 to 7. [Appendix 9] An acquisition means for acquiring a video, A plurality of pairs of a first frame and a second frame different from the first frame are extracted from a plurality of frames included in one video, and each of the object included in the first frame and the object included in the second frame is detected. A tracking means for tracking the objects included in the video based on the association result of the plurality of pairs in which the object included in the first frame is associated with the object included in the second frame, and having an association means generated by learning the method of associating the objects based on the association result of the plurality of pairs. Comprising The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval. Tracking device. [Appendix 10] Acquire one video, A plurality of pairs of a first frame and a second frame different from the first frame are extracted from a plurality of frames included in the video, Each of the object included in the first frame and the object included in the second frame is detected, The object included in the first frame is associated with the object included in the second frame using an associator, Based on the association result by the associator of the plurality of pairs, the associator is caused to learn the method of associating the objects, The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval. Learning method [Appendix 11] Cause a computer to acquire one video, extract a plurality of pairs of a first frame and a second frame different from the first frame from a plurality of frames included in the video, detect each of the objects included in the first frame and the objects included in the second frame, associate the objects included in the first frame with the objects included in the second frame using an associator, cause the associator to learn the object association method based on the association results by the associator for the plurality of pairs, the plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval A recording medium on which a computer program for executing the learning method is recorded. [Appendix 12] Acquire a video, extract a plurality of pairs of a first frame and a second frame different from the first frame from a plurality of frames included in one video, detect each of the objects included in the first frame and the objects included in the second frame, and based on the association results of the plurality of pairs in which the objects included in the first frame are associated with the objects included in the second frame, have an association means generated by causing learning of the object association method, and track the objects included in the video based on the association of the objects by the association means A tracking method, comprising: the plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval Tracking method [Appendix 13] Cause a computer to acquire a video, From a plurality of frames included in one video, a plurality of pairs of a first frame and a second frame different from the first frame are extracted, and each of the objects included in the first frame and the objects included in the second frame is detected. Based on the association results of the plurality of pairs in which the objects included in the first frame and the objects included in the second frame are associated, learning of the object association method is performed to generate an association means, and based on the association of the objects by the association means, track the objects included in the video A tracking method, comprising: The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second time interval different from the first interval A recording medium on which a computer program for executing the tracking method is recorded
[0120] At least some of the constituent elements of each of the above embodiments can be appropriately combined with at least some of the other constituent elements of each of the above embodiments. Some of the constituent elements of each of the above embodiments may not be used. Also, to the extent permitted by law, the disclosures of all documents (for example, published gazettes) cited in this disclosure are incorporated by reference to form part of the description of this disclosure
[0121] This disclosure can be appropriately changed within a range not contrary to the technical idea that can be read from the claims and the entire specification. A learning device, a learning method, a tracking device, a tracking method, and a recording medium involving such changes are also included in the technical idea of this disclosure
Description of Reference Numerals
[0122] 1, 2, 3, 4, 5 Learning device 11, 211, 611 Acquisition unit 12, 212 Extraction unit 13, 213 Detection unit 14, 214 Association unit 15, 215 Learning unit MM Association model 6 Tracking device 616 Tracking Unit
Claims
1. An acquisition means for acquiring one video; An extraction means for extracting a plurality of pairs of a first frame and a second frame different from the first frame from a plurality of frames included in the video; A detection means for detecting each of an object included in the first frame and an object included in the second frame; An association means for associating the object included in the first frame with the object included in the second frame; A learning means for causing the association means to learn the object association method based on the association results of the plurality of pairs by the association means and comprising: The plurality of pairs include a first pair in which a time interval between the first and second frames is a first interval, and a second pair in which a time interval between the first and second frames is a second interval different from the first interval A learning device.
2. The extraction means randomly selects the first frame and the second frame from the plurality of frames to extract the pair. The learning device according to Claim 1.
3. The extraction means selects, as the second frame, a frame before or after a predetermined number of frames from the first frame to extract the pair. The learning device according to Claim 1 or 2.
4. The learning means calculates a learning loss based on the association result, and causes the association means to learn the association method from the learning loss. The learning device according to Claim 1 or 2.
5. The learning means causes the association means to learn the object association method based on a forward association result in which an object included in the first frame is associated with an object included in the second frame which is a frame after the first frame, and a reverse association result in which an object included in the second frame is associated with an object included in the first frame which is a frame before the second frame. The learning device according to Claim 1 or 2.
6. An acquisition means for acquiring a video; From a plurality of frames included in one video, a plurality of pairs of a first frame and a second frame different from the first frame are extracted, each of the objects included in the first frame and the objects included in the second frame is detected, and based on the association results of the plurality of pairs in which the objects included in the first frame and the objects included in the second frame are associated, there is an association means generated by causing learning of the object association method, and based on the association of the objects by the association means, a tracking means for tracking the objects included in the video Comprising The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second interval different from the first interval Tracking device
7. A learning method executed by a computer, comprising Obtaining one video From a plurality of frames included in the video, a plurality of pairs of a first frame and a second frame different from the first frame are extracted Detecting each of the objects included in the first frame and the objects included in the second frame Associating the objects included in the first frame and the objects included in the second frame using an associator Based on the association results by the associator of the plurality of pairs, causing the associator to learn the object association method The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second interval different from the first interval Learning method
8. Causing a computer to Obtain one video From a plurality of frames included in the video, a plurality of pairs of a first frame and a second frame different from the first frame are extracted Detecting each of the objects included in the first frame and the objects included in the second frame Associating the objects included in the first frame and the objects included in the second frame using an associator Based on the association results by the associator of the plurality of pairs, causing the associator to learn the object association method The plurality of pairs include a first pair in which the time interval between the first and second frames is a first interval, and a second pair in which the time interval between the first and second frames is a second interval different from the first interval A computer program for causing the learning method to be executed
9. A tracking method executed by a computer, comprising: obtaining a video; extracting a plurality of pairs of a first frame and a second frame different from the first frame from a plurality of frames included in one video, detecting each of an object included in the first frame and an object included in the second frame, and based on the association results of the plurality of pairs in which the object included in the first frame and the object included in the second frame are associated with each other, having an association means generated by causing learning of the object association method, and tracking the object included in the video based on the association of the object by the association means A tracking method, comprising: the plurality of pairs include a first pair in which a time interval between the first and second frames is a first interval, and a second pair in which a time interval between the first and second frames is a second interval different from the first interval A tracking method.
10. Causing a computer to obtain a video; extracting a plurality of pairs of a first frame and a second frame different from the first frame from a plurality of frames included in one video, detecting each of an object included in the first frame and an object included in the second frame, and based on the association results of the plurality of pairs in which the object included in the first frame and the object included in the second frame are associated with each other, having an association means generated by causing learning of the object association method, and tracking the object included in the video based on the association of the object by the association means A tracking method, comprising: the plurality of pairs include a first pair in which a time interval between the first and second frames is a first interval, and a second pair in which a time interval between the first and second frames is a second interval different from the first interval A computer program for causing a computer to execute the tracking method.
Citation Information
Patent Citations
Image data extraction device and image data extraction method
JP2018081545A
Information processing device, information processing method and program
JP2019207561A
Body homologizing device, body homologizing system, body homologizing method, and computer program
JP2020181268A
Motion detection device, feature detection device, fluid detection device, motion detection system, motion detection method, program, and recording medium
WO2020022362A1
Learning device, learning method, object detection device, object detection method, and recording medium
WO2021070324A1