Learning device, inference device, learning method, inference method, and program
The learning device addresses the low performance of conventional object tracking technologies by using a neural network model that learns from both binary and real-valued data, maintains spatial resolution, and incorporates temporal information, resulting in accurate tracking of small objects.
Patent Information
- Application Number
- JP2023190049
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2025-05-19
AI Technical Summary
Conventional object tracking technologies have low performance, especially when tracking small objects, due to issues with spatial resolution loss in neural networks, inaccurate object position identification, and insufficient consideration of temporal information.
A learning device that learns a neural network model for object tracking using both binary and real-valued ground truth data, maintaining spatial resolution through modifications to the neck part of HRNet, and incorporating temporal information by concatenating consecutive frames along the channel dimension.
Enables accurate tracking of small objects by maintaining spatial resolution, improving object position identification, and effectively considering temporal information, thereby enhancing overall tracking performance.
Smart Images

Figure 2025077674000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to object tracking technology in video.
Background Art
[0002] Object tracking technologies such as ball tracking in sports have attracted attention. As conventional technologies for object tracking, for example, there are the technologies disclosed in Non-Patent Documents 1 and 2. Non-Patent Document 1 discloses a technology for detecting an object by post-processing a heatmap obtained by independently inputting each frame of a video into a neural network. Here, a heatmap is a two-dimensional map in which the value becomes higher as it is closer to the object position. By post-processing, the peak of the value in the heatmap is detected, and the position of the object, that is, the (x, y) coordinates in the image of the object are output.
[0003] In the technology disclosed in Non-Patent Document 2, by combining a plurality of consecutive frames and inputting them into a neural network, the detection performance is improved in consideration of temporal information. Also, as post-processing, the heatmap is binarized, and the (x, y) coordinates of the object are output as the geometric mean of the extracted maximum connected component. This improves the robustness against noise. In the learning of the model, a binary mask is generated from the correct object coordinate data assigned to each frame, and the model parameters of the neural network are optimized so as to infer it.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
[0005] However, conventional object tracking technologies have a problem in that the object tracking performance is low, especially when the target object is small.
[0006] The present invention has been made in view of the above points, and an object thereof is to provide a technology that enables accurate tracking of small objects. [Means for Solving the Problems]
[0007] According to the disclosed technology, a learning device that learns a neural network model for tracking an object in an image, Ground truth data generated from a plurality of images in which the object exists, the model is learned using first ground truth data consisting of binary data, and the model is further learned using second ground truth data consisting of real-valued data, which is ground truth data generated from a specific plurality of images among the plurality of images. A learning unit A learning apparatus including the same is provided.
Advantages of the Invention
[0008] According to the disclosed technology, it becomes possible to accurately track small objects.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0010] Hereinafter, embodiments (this embodiment) of the present invention will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the following embodiments.
[0011] Hereinafter, the problems related to the technology of this embodiment will be described in detail, and then the technology according to this embodiment will be explained.
[0012] In this embodiment, it is assumed that the object to be tracked is a small object such as the ball in tennis shown in FIG. 1. However, the technology according to the present invention is not limited to small objects such as balls, and can be applied to the tracking of any object.
[0013] In the example shown in FIG. 1, an image or video obtained by shooting a tennis match is input to an inference device 200 described later, and the inference device 200 outputs an image or video that shows the trajectory of the moving ball.
[0014] (Regarding problems) As conventional techniques for object tracking, there are, for example, the techniques disclosed in Non-Patent Documents 1 and 2 described above. However, the conventional techniques disclosed in Non-Patent Documents 1 and 2 have the following problems (1) to (3).
[0015] Problem (1); The neural network used in the techniques of Non-Patent Documents 1 and 2 is an encoder-decoder type neural network as shown in FIG. 2. In an encoder-decoder type neural network, the spatial resolution decreases in the encoder part. Therefore, the detection and tracking performance of particularly small objects deteriorates.
[0016] As a neural network that is not of the encoder-decoder type, for example, it is conceivable to use High-Resolution Modules (HRMs) disclosed in Non-Patent Document 3. However, when the technology disclosed in Non-Patent Document 3 is used as it is, the spatial resolution of the input data decreases in the previous stage (neck part) of the HRMs.
[0017] Problem (2): In the technologies of Non-Patent Documents 1 and 2, the model is trained to infer a binary mask that is the correct data. In such a learning method, the accuracy of object position identification may be impaired.
[0018] Problem (3): In the technologies of Non-Patent Documents 1 and 2, temporal information is not sufficiently considered. That is, in the technology of Non-Patent Document 1, temporal information is not considered at all, and in the technology of Non-Patent Document 2, temporal information is only considered within consecutive frames that are simultaneously input to the neural network.
[0019] For the reasons (1) to (3) above, the conventional technology has a problem that the tracking performance of an object is low, especially when the target object is small.
[0020] (Outline of the Embodiment) Hereinafter, a technology that solves the above problems and enables accurate tracking even when the target object is small will be described. Specifically, a learning device 100 that learns a model of a neural network for performing object tracking and an inference device 200 that performs object tracking using the learned model will be described.
[0021] The learning device 100 learns a model of a neural network that predicts a heatmap representing the position of an object in the input image group (images). The inference device 200 uses the learned model to detect the (x, y) coordinates of the position of the object from the image group in the input video. In the inference by the inference device 200, the object position is detected by post-processing on the heatmap obtained by the model. In this embodiment, the learning device 100 and the inference device 200 are separate devices, but the learning device 100 and the inference device 200 may be the same device.
[0022] Hereinafter, each of the model, learning process, and inference process in this embodiment will be described.
[0023] (Regarding the Model) First, a model of a neural network that is commonly used in both the learning device 100 and the inference device 200 will be described.
[0024] The model according to this embodiment is a model that can generate a heatmap with the same spatial resolution H×W as the input data (input tensor). In this embodiment, as a model in which the spatial resolution of the input data is not impaired, a model based on HRMs in Non-Patent Document 3 is used, while changing the processing in its neck part to processing in which the spatial resolution is not impaired. Specifically, it is as follows.
[0025] HRNet disclosed in Non-Patent Document 3 is composed of a neck part (stem block) and HRMs (multi-stage high-resolution modules). In each new stage in HRMs, one "convolution block from high resolution to low resolution" is added. Fig. 3 shows an example of HRMs used in this embodiment. HRMs themselves are an existing technology, and their details are as described in Non-Patent Document 3, for example.
[0026] When using the HRNet disclosed in Non-Patent Document 3 as it is, the feature amount input to HRMs is downsized to one-fourth by the neck part. This is shown in Fig. 4. Due to this processing, the spatial resolution may be impaired.
[0027] Therefore, in this embodiment, in order to increase the resolution of the intermediate representation, the stride in the neck part is reduced, and a tensor having a higher spatial resolution is input to HRMs.
[0028] Configuration examples having a neck part and HRMs in this embodiment are shown in Figs. 5 and 6. Either the configuration of Fig. 5 or the configuration of Fig. 6 may be used, but in this embodiment, the configuration shown in Fig. 6 is used.
[0029] In addition, in this embodiment, for example, MIMO (Multiple-In Multiple-Out) disclosed in Non-Patent Document 4 is adopted. In this embodiment, using the MIMO, N consecutive frames are concatenated along the channel dimension, and the obtained tensor of H×W×3N is input to the model of this embodiment. The model generates corresponding N heatmaps with the same spatial resolution as the input (i.e., H×W×N).
[0030] The "3" in the above "3N" means that three channels of R, G, and B are used. Note that the number of channels is not limited to "3", and it may be, for example, "1" (corresponding to grayscale video), or a value larger than "3" (corresponding to multispectral video).
[0031] Note that in this embodiment, using a model based on the above-mentioned HRMs is just an example, and the models that can be used in this embodiment are not limited to models based on HRMs, and any neural network model for image processing can be used.
[0032] (Configuration and operation overview of learning device 100) Subsequently, the configuration and operation overview of the learning device 100 in this embodiment will be described. FIG. 7 shows a configuration example of the learning device 100. As shown in FIG. 7, the learning device 100 in this embodiment includes an input unit 110, a learning unit 120, an output unit 130, and a data storage unit 140.
[0033] Learning data is input from the input unit 110 and stored in the data storage unit 140. The learning data may be video of an object being photographed, or correct data (binary GT map, real-valued GT map) obtained by the processing described later.
[0034] The learning unit 120 holds a model to be learned (specifically, the parameters of the model), and performs learning of the model (optimization of parameters) using the learning data read from the data storage unit 140. For example, the error backpropagation method is used for learning the model.
[0035] The output unit 130 outputs the learned model (learned parameters). The output learned model is input to the inference device 200.
[0036] In the prior art, only binary data is used as the correct answer data during model learning used by the learning unit 120. However, in the present embodiment, in addition to binary data, real-valued data is used as the correct answer data. The real-valued data is designed such that the closer the coordinate value is to the correct (x, y) coordinates, the higher it is. Further, in the present embodiment, an error function that enables model learning with real-valued correct answer data is used.
[0037] Furthermore, in the present embodiment, the model is learned by combining the learning with binary data and the corresponding error function used in the prior art and the learning with the above-described real-valued data and the corresponding error function. Also, with the model, the object position is output as the centroid position considering the values of the output heatmap instead of the conventional geometric mean. Outputting the object position as the centroid position considering the values of the output heatmap is also a feature during inference. Note that the calculation of the centroid position based on the heatmap may be performed by the model or outside the model in the learning unit 120 (or the inference unit 220).
[0038] Hereinafter, the learning process of the above-described model executed by the learning unit 120 will be described in more detail.
[0039] (Details of the learning process) As the overall procedure, first, correct answer data (referred to as a GT (Ground Truth) map) is calculated from the 2D (two-dimensional) object position, and the parameters of the model are optimized so as to minimize the loss (error) between the prediction by the model and the GT map.
[0040] The correct object position p within the image GT ∈R 2 is stored in the data storage unit 140. The learning unit 120 generates a binary GT map y GT from the correct object position p bin based on the following formula (1) and stores it in the data storage unit 140. Note that the binary GT map y bin may also be generated outside the learning device 100.
[0041]
Equation
[0042] Therefore, in this embodiment, in order to capture the object more accurately, a new learning method is used. Specifically, first, the learning unit 120 generates a real-valued GT map y
[0043] based on the following formula (2). The generated real-valued GT map y real is stored in the data storage unit 140. Note that the generation of the real-valued GT map may also be performed outside the learning device 100. real
[0044]
Equation
[0045] When using the real-valued GT map as the ground truth data, the model parameters are optimized by minimizing the following loss function (quality focal loss).
[0046] [Equation] In Equation (3), σ p is the sigmoid output of the model prediction at p, and β is a parameter that controls the downweighting rate. Note that when using the binary GT map, Equation (3) is equivalent to the focal loss.
[0047] (Detailed learning procedure) When applying the real-valued GT map to all the training data, it was found that the detection performance of the object (specifically, the ball) did not improve. Therefore, in this embodiment, in the learning, the real-valued GT map is applied only to the data (samples) that are difficult to detect. The specific procedure is as follows. This learning method may be called HLSM (Hard-to-Localize Sample Mining).
[0048] After performing learning for a predetermined number of epochs using the binary GT map obtained from the video (image group) to be learned, the learning unit 120 makes an inference on the above video (image group) using the model having the parameters at that time, and identifies "a plurality of images" in which the object position as the inference result is far from the correct position (GT position). For example, "a plurality of images" in which the object position as the inference result is separated from the correct position (GT position) by a distance equal to or greater than a threshold value are identified.
[0049] Note that the inference here is made using the centroid position of the values in the heatmap output from the model without performing post - processing using past images described later. For example, when there are multiple candidate regions (blobs) where the "1" of the binarized values in the heatmap are connected, the centroid position of the candidate region with the maximum sum of the values of the candidate region before binarization is taken as the inference result.
[0050] However, it is also possible to perform post - processing using past images described later for this inference here.
[0051] The learning unit 120 generates a real - valued GT map for the above - mentioned "multiple images" using Equation (2), and performs model learning using the real - valued GT map and Equation (3) in the remaining one or more epochs.
[0052] (Configuration and Outline of Operation of Inference Device 200) Next, the configuration and outline of operation of the inference device 200 in the present embodiment will be described. FIG. 9 shows a configuration example of the inference device 200. As shown in FIG. 9, the inference device 200 in the present embodiment includes an input unit 210, an inference unit 220, an output unit 230, and a data storage unit 240.
[0053] The learned parameters are input from the input unit 210 and stored in the data storage unit 240. The parameters are read out by the inference unit 220, and the inference unit 220 holds a model set with the parameters.
[0054] The video (multiple images) to be the object of object tracking is input from the input unit 210. The inference unit 220 detects the object position from each image of the video using the model and generates an image (video) indicating the object position. The output unit 230 outputs the image (video).
[0055] In this embodiment, in the post - processing for the heatmap output from the model, the inference unit 220 performs processing considering the time series. Specifically, the inference unit 220 predicts the object position in the current frame from the detection results in one or more frames immediately before the target current frame, and for example, outputs, as the object position in the current frame, the one among one or more detection candidate positions in the current frame that is closest to the prediction. Also, when obtaining the detection candidate positions in the current frame, the inference unit 220 processes one frame by a plurality of samplings and uses all the results obtained in each of them as candidate positions.
[0056] (Details of Inference Processing) First, the basic flow of the inference process executed by the inference unit 220 will be described. When a video clip consisting of T images is input to the inference unit 220 via the input unit 210, the inference unit 220 samples N consecutive images in order without duplication (that is, the sampling step size here is N). N is an integer of 1 or more.
[0057] Next, the inference unit 220 performs pre - processing on the sampled images to generate a tensor and inputs the tensor to a learned model. The model generates N heatmaps.
[0058] The inference unit 220 binarizes each heatmap with a threshold value (for example, 0.5), detects the connected components (that is, blobs) where 1s are connected, and for each blob, estimates the candidate 2D object position (for example, ball position) together with its confidence. The blobs may be called candidate regions.
[0059] In this processing procedure, the inference unit 220 generally calculates the object position as the geometric center (that is, the centroid) in the blob and calculates the confidence as the size of the blob.
[0060] For each image, the inference unit 220 selects the object position with the highest confidence as the inference result. On the other hand, if no blob is found, the inference unit 220 assumes that no object is detected. In the present embodiment, the following three techniques are introduced for the above basic processing.
[0061] <Regarding the object position> The value of the heatmap within the blob is important for accurately estimating the object position. Therefore, in the present embodiment, the object position is calculated as the center (center of gravity) of the value of the heatmap within the blob (the original value before binarization), and the sum of the values of the heatmap within the blob (the original value before binarization) is used as the confidence.
[0062] <Online tracking> Especially when detecting small objects such as balls, errors are likely to occur if only relying on the detection confidence in the image. Therefore, in the present embodiment, the inference unit 220 performs online tracking considering both the detection confidence and temporal consistency.
[0063] Specifically, for the image at time t+1, the inference unit 220 uses the heatmap generated by the model to detect candidate positions of the object (generally, there are multiple candidate positions (and corresponding confidences)), and predicts the object position at t+1 from the detection results (object positions) in one or more images (frames) before t+1. One or more images before t+1 are, for example, the image at t, the image at t-1, and the image at t-2.
[0064] The inference unit 220 excludes candidate positions that are farther than a threshold from the object position at t+1 predicted from past images among all candidate positions at t+1, and selects the candidate position with the highest confidence among the remaining candidate positions as the inference result (detection result) at t+1.
[0065] The inference unit 220 may select, as the inference result (detection result) at t+1, the candidate position that is closest to the object position at t+1 predicted from past images among all candidate positions at t+1.
[0066] The method of predicting the object position at t+1 from the object position in the past image is not limited to a specific method. For example, as shown in the following formula (4), using the object positions in the images at t, t-1, and t-2, the object position ^ p at t+1 can be predicted.
[0067] [Number] <Oversampling> In this embodiment, the inference unit 220 oversamples the same image in different combinations of MIMO sampling and uses all candidate positions generated as inference results of the same image multiple times. Specifically, it is as follows. Here, as an example, three images are input to the model using MIMO.
[0068] The inference unit 220 takes the image at time 1, the image at time 2, and the image at time 3 as inputs, and uses the method described above to obtain one or more candidate positions for the image at time 1, one or more candidate positions for the image at time 2, and one or more candidate positions for the image at time 3. Next, the inference unit 220 takes the image at time 2, the image at time 3, and the image at time 4 as inputs, and obtains one or more candidate positions for the image at time 2, one or more candidate positions for the image at time 3, and one or more candidate positions for the image at time 4.
[0069] The processing using oversampling as described above is sequentially performed. In the above case, for example, for the image at time 2, two inferences are performed, and one or more candidate positions are obtained in each inference. The inference unit 220 uses all candidate positions obtained by multiple inferences for an image at a certain time as the candidate positions for the image at that time. That is, one candidate position is selected from all candidate positions obtained by multiple inferences based on the predicted position obtained from the past image.
[0070] (Example of the hardware configuration of the device) The learning device 100 and the inference device 200 described in this embodiment can both be realized, for example, by causing a computer to execute a program. This computer may be a physical computer or a virtual machine on the cloud.
[0071] That is, the device can be realized by executing a program corresponding to the processing performed by the device using hardware resources such as a CPU, GPU, and memory built into the computer. The above program can be recorded on a computer-readable recording medium (such as a portable memory), saved, or distributed. It is also possible to provide the above program through a network such as the Internet or email.
[0072] FIG. 10 is a diagram showing an example of the hardware configuration of the computer. The computer in FIG. 10 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, etc., which are mutually connected by a bus BS. Note that the computer may further include a GPU.
[0073] The program for realizing the processing on the computer is provided, for example, by a recording medium 1001 such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the installation of the program does not necessarily have to be performed from the recording medium 1001, and it may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program and also stores necessary files, data, etc.
[0074] When there is an instruction to start a program, the memory device 1003 reads and stores the program from the auxiliary storage device 1002. The CPU 1004 realizes the functions related to the device according to the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network or the like. The display device 1006 displays a GUI (Graphical User Interface) or the like according to the program. The input device 1007 is composed of a keyboard, a mouse, buttons, or a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the calculation result.
[0075] (Effects of the Embodiment) With the technology according to the present embodiment described above, it becomes possible to accurately track small objects.
[0076] Regarding the above embodiment, the following additional remarks are further disclosed.
[0077] <Additional Remark> (Additional Remark Item 1) A learning device that learns a neural network model for tracking an object in an image, a learning unit that learns the model using first correct data that is correct data generated from a plurality of images in which the object exists and that consists of binary data, and further learns the model using second correct data that is correct data generated from a specific plurality of images among the plurality of images and that consists of real-valued data A learning device comprising the same. (Additional Remark Item 2) The specific plurality of images are images in which the position of the object inferred by the model learned using the first correct data is separated from the correct position by a threshold value or more. The learning device according to Additional Remark Item 1. (Additional Remark Item 3) The model outputs a heatmap having the same spatial resolution as the spatial resolution of the input tensor. The learning device according to Additional Remark Item 1. (Additional Remark Item 4) An inference device that tracks an object in an image using a learned neural network model, an input unit that inputs a plurality of images in which the object exists, for an image at a specific time among the plurality of images, based on the detection result of the object for an image at a time earlier than the specific time, from among a plurality of candidate positions of the object obtained using the model, selects one candidate position and outputs the one candidate position as the detection result for the image at the specific time, an inference unit An inference device comprising: (Additional item 5) The inference unit, uses the detection result of the object for an image at a time earlier than the specific time to obtain a predicted position of the object at the specific time, selects the one candidate position from among one or more candidate positions remaining after excluding one or more candidate positions that are more than a threshold value away from the predicted position from among the plurality of candidate positions The inference device according to additional item 4. (Additional item 6) The inference unit performs multiple inferences by oversampling on the image at the specific time, and selects the one candidate position from among the plurality of candidate positions obtained by the multiple inferences The inference device according to additional item 4. (Additional item 7) The model outputs a heatmap with the same spatial resolution as the spatial resolution of the input tensor The inference device according to additional item 4. (Additional item 8) A learning method executed by a learning device that learns a neural network model for tracking an object in an image, a learning step of learning the model using first ground truth data composed of binary data, which is ground truth data generated from a plurality of images in which the object exists, and further learning the model using second ground truth data composed of real-valued data, which is ground truth data generated from a specific plurality of images among the plurality of images A learning method comprising: (Additional item 9) An inference method executed by an inference device that tracks an object in an image using a learned neural network model, an input step of inputting a plurality of images in which the object exists, for an image at a specific time among the plurality of images, based on the detection result of the object for an image at a time earlier than the specific time, from among the plurality of candidate positions of the object obtained using the model, selecting one candidate position and outputting the one candidate position as the detection result for the image at the specific time; and an inference step An inference method comprising: (Appended Claim 10) A non-transitory storage medium storing a program for causing a computer to function as each part in the learning device according to any one of Claims 1 to 3. (Appended Claim 11) A non-transitory storage medium storing a program for causing a computer to function as each part in the inference device according to any one of Claims 4 to 7.
[0078] As described above, the present embodiment has been described, but the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.
Explanation of Signs
[0079] 100 Learning device 110 Input unit 120 Learning unit 130 Output unit 140 Data storage unit 200 Inference device 210 Input unit 220 Inference unit 230 Output unit 240 Data storage unit 1000 Drive device 1001 Recording medium 1002 Auxiliary storage device 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input device 1008 Output device
Claims
1. A learning device that learns a neural network model for tracking an object in an image, comprising: a learning unit that learns the model using first supervised answer data generated from a plurality of images in which the object is present and made up of binary data, and further learns the model using second supervised answer data generated from specific images among the plurality of images and made up of real-valued data. A learning device comprising:
2. The specific plurality of images are images in which the position of the object inferred by the model trained using the first correct answer data is away from the correct answer position by a threshold or more. The learning device according to claim 1 .
3. The model outputs a heatmap with the same spatial resolution as the input tensor. The learning device according to claim 1 .
4. An inference device that tracks objects in an image using a trained neural network model, an input unit for inputting a plurality of images in which the object is present; an inference unit that selects, for an image at a specific time among the plurality of images, one candidate position of the object from among a plurality of candidate positions of the object acquired using the model, based on a detection result of the object for an image at a time prior to the specific time, and outputs the one candidate position as a detection result for the image at the specific time; An inference device comprising:
5. The inference unit is obtaining a predicted position of the object at the specific time using a detection result of the object for an image taken at a time prior to the specific time; Among the plurality of candidate positions, one or more candidate positions that are farther away from the predicted position than a threshold value are excluded, and from the remaining one or more candidate positions, the one candidate position is selected.
5. The inference device according to claim 4.
6. The inference unit executes inference multiple times by oversampling the image at the specific time, and selects the one candidate position from multiple candidate positions obtained by the multiple inferences.
5. The inference device according to claim 4.
7. The model outputs a heatmap with the same spatial resolution as the input tensor.
5. The inference device according to claim 4.
8. A learning method executed by a learning device that learns a neural network model for tracking an object in an image, comprising: a learning step of learning the model using first supervised answer data generated from a plurality of images in which the object is present and made of binary data, and further learning the model using second supervised answer data generated from specific images among the plurality of images and made of real-valued data; A learning method that provides:
9. An inference method executed by an inference device that tracks an object in an image using a trained neural network model, comprising: an input step of inputting a plurality of images in which the object is present; an inference step of selecting one candidate position from among a plurality of candidate positions of the object acquired using the model for an image at a specific time among the plurality of images, based on a detection result of the object for an image at a time prior to the specific time, and outputting the one candidate position as a detection result for the image at the specific time; An inference method comprising:
10. A program for causing a computer to function as each unit in the learning device according to any one of claims 1 to 3.
11. A program for causing a computer to function as each unit in the inference device according to any one of claims 4 to 7.