Tracking device, tracking method, and storage medium

The tracking device and method use inferences on time-series images to accurately assign tracking IDs by considering motion transitions, addressing errors in overlapping individuals and similar clothing, enhancing tracking robustness in industrial settings.

JP2026082276APending Publication Date: 2026-05-19NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NEC CORP
Filing Date
2024-11-07
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing tracking technologies fail to accurately identify a tracking target when individuals are close to each other or wear similar clothing, leading to erroneous assignment of tracking IDs, especially in industrial settings where visible information is not sufficient.

Method used

A tracking device and method that performs first and second inferences based on time-series images to infer the behavior of objects, using machine learning models to distinguish and assign tracking IDs accurately by considering motion transitions and historical data.

Benefits of technology

Enables accurate identification of tracking targets by considering motion transitions, even in situations where individuals overlap, thereby improving tracking robustness in industrial settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026082276000001_ABST
    Figure 2026082276000001_ABST
Patent Text Reader

Abstract

The present invention provides a tracking device, a tracking method, and a storage medium capable of accurately identifying a target from an image. [Solution] The tracking device 1X comprises a first inference means 24X, a second inference means 25X, and an identification means 26X. The first inference means 24X performs a first inference when a target image containing multiple objects is obtained at a reference time, inferring the behavior of the tracked object at the reference time, assuming for each of the multiple objects that the tracked object, tracked based on time-series images obtained before the reference time, is one of the multiple objects. The second inference means 25X performs a second inference inferring the behavior of the tracked object at the reference time based on time-series images. The identification means 26X identifies the object representing the tracked object in the target image obtained at the reference time, based on the inference results for each assumption from the first inference and the inference results from the second inference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of a tracking device, a tracking method, and a storage medium for tracking an object using time-series images.

Background Art

[0002] There is a technology for tracking an object such as a person or an object from time-series images. For example, Patent Document 1 discloses a tracking system that extracts the position of a person on an image using machine learning, performs association of the detected person using the predicted position of the person obtained by past tracking (tailgating) processing, and assigns a tracking ID to each person. Patent Document 1 also discloses processing related to detection and prediction of the behavior of a person to be tracked.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When people are close to each other, tracking may fail because a tracking ID is erroneously assigned to a wrong person. In addition, in an industrial site where workers wear the same work clothes, visible information is not useful for tracking, so such an erroneous assignment of a tracking ID is likely to occur. Therefore, it is desirable to accurately track a target regardless of visible information.

[0005] In view of the above problems, one object of the present disclosure is to provide a tracking device, a tracking method, and a storage medium capable of accurately identifying a tracking target from an image.

Means for Solving the Problems

[0006] One aspect of the tracking device is A first inference means performs a first inference to infer the behavior of the tracked object at the reference time, assuming that the tracked object tracked based on time-series images obtained before the reference time is one of the multiple objects, when a target image including multiple objects is obtained at a reference time. A second inference means performs a second inference to infer the operation of the tracked object at the reference time based on the time-series image, A means for identifying the object representing the tracking target in the target image based on the inference results for each assumption in the first inference and the inference results in the second inference, It is a tracking device that has [a certain feature].

[0007] One aspect of the tracking method is: Computers When a target image containing multiple objects is obtained at a reference time, a first inference is performed to infer the behavior of the tracked object at the reference time, assuming that the tracked object, based on time-series images obtained before the reference time, is one of the multiple objects. A second inference is performed to infer the behavior of the tracked object at the reference time based on the aforementioned time-series images. Based on the inference results for each assumption in the first inference and the inference results in the second inference, the object representing the tracking target is identified in the target image. This is a tracking method.

[0008] One aspect of the program is: When a target image containing multiple objects is obtained at a reference time, a first inference is performed to infer the behavior of the tracked object at the reference time, assuming that the tracked object, based on time-series images obtained before the reference time, is one of the multiple objects. A second inference is performed to infer the behavior of the tracked object at the reference time based on the aforementioned time-series images. This program identifies the object representing the tracking target in the target image based on the inference results for each assumption in the first inference and the inference results in the second inference. [Effects of the Invention]

[0009] One example of the effects of this disclosure is that the target being tracked can be accurately identified from the image. [Brief explanation of the drawing]

[0010] [Figure 1] The following shows a schematic configuration of the tracking system. [Figure 2] This shows the hardware configuration of the tracking device. [Figure 3] This is an example of a functional block of a tracking device. [Figure 4] (A) This is a diagram illustrating a specific example of the process for determining whether or not a distinction can be made. (B) This is a table showing the degree of overlap on the image for each combination of detected location information and predicted location information. [Figure 5] This section provides a concrete example of the input and output of the motion inference model M1. [Figure 6] (A) This shows an overview of the process for generating historical image-based operational information for the tracking target with tracking ID "z" based on the first generation example. (B) This shows an overview of the process for generating historical image-based operational information based on the second generation example. [Figure 7] An example of the relationship between each assumption and the inference results output by the operation inference model M1 is shown. [Figure 8] This section outlines the process by which the motion inference model M2 generates historical image-based motion information for each of the tracking IDs, Tracking ID 1 and Tracking ID 2. [Figure 9] This figure shows an overview of the comparison between the inference results shown in Figure 7 and the inference results shown in Figure 8. [Figure 10] This is an example of a flowchart showing the processing steps performed by a tracking device. [Figure 11] (A) This is the first example of the image display screen. (B) This is the first example of the performance score display screen. [Figure 12](A) This is a second display example of the image display screen. (B) This is a second display example of the motion score display screen. [Figure 13] This is a block diagram of the tracking device. [Figure 14] This is an example of a flowchart showing the processing procedure of the tracking device.

Mode for Carrying Out the Invention

[0011] Hereinafter, embodiments of the tracking device, the tracking method, and the storage medium will be described with reference to the drawings.

[0012] <First Embodiment> (1) System Configuration FIG. 1 shows a schematic configuration of a tracking system 100. The tracking system 100 is a system that performs tracking of an object based on time-series images, and mainly includes a tracking device 1, a storage device 2, a display device 3, an input device 4, and a camera 5. Hereinafter, the object to be tracked will be described as a general person, but instead, it may be a person having specific attributes (such as gender, age, etc.), or a specific type of moving object other than a person (such as a vehicle, a robot, etc.). Note that "motion" represents the entire movement of an object, and when the tracking target is a person, it is synonymous with "behavior".

[0013] The tracking device 1 identifies the correspondence relationship between images of a tracking target that is a subject in time-series images captured by the camera 5, and manages the tracking target by assigning common identification information (also referred to as "tracking ID") to the tracking target that is common among the images. In this case, the tracking device 1 updates the information stored in the storage device 2 based on the tracking result. Further, the tracking device 1 may present information based on the tracking result to the user of the tracking system 100 by the display device 3, or receive an input from the user by the input device 4 (so-called external input).

[0014] The storage device 2 is a memory that stores various information necessary for the processing of the tracking device 1, and functionally includes a time-series image storage unit D1 and a tracking information storage unit D2.

[0015] The time-series image storage unit D1 stores the time-series images generated by the camera 5. The images generated by the camera 5 may be supplied directly to the storage device 2, or they may be supplied to the storage device 2 via the tracking device 1 or the like.

[0016] The tracking information storage unit D2 stores tracking information, which is information generated by the tracking process performed by the tracking device 1. Tracking information is generated for each image registered in the time-series image storage unit D1 and is linked to the corresponding image. The tracking information includes a tracking ID assigned to each tracking target present in the corresponding image, location information of each tracking target within the image, and operation information representing the operation recognition result of each tracking target. The location information mentioned above represents the area of ​​the tracking target within the image, for example, information indicating the bounding box surrounding the tracking target (i.e., rectangular information). The time-series location information for each tracking ID identified by the tracking information corresponds to the trajectory information of the tracking target represented by each tracking ID. The operation information represents, for example, a score representing the likelihood (i.e., confidence level) for each operation type (i.e., operation class) representing a candidate for the expected operation.

[0017] Hereafter, the most recent image supplied from the time-series image storage unit D1 to the tracking device 1 will be referred to as the "target image," and images obtained before the target image will be referred to as "past images." That is, the target image is the image on which the tracking information is linked, while past images are images to which the tracking information has already been linked. For the sake of explanation, the target image will be assumed to be the image generated at the reference time "t," and past images will be assumed to be the images generated at times t-1, t-2, ...

[0018] Furthermore, the memory device 2 stores information about the motion inference model (a so-called motion recognition device) that performs inference about the actions of the tracked object. In this embodiment, the tracking device 1 selectively uses multiple motion inference models depending on the application. The motion inference model is, for example, a machine learning model, and may be a learning model based on a neural network, or another type of learning model such as a support vector machine, or a learning model that combines these. Examples of motion inference models having a neural network-based configuration include SlowFast and VideoMAE. For example, if the above-mentioned motion inference model has a neural network-based configuration such as a convolutional neural network, the memory device 2 stores information about various parameters such as the layer structure of the motion inference model, the neuron structure of each layer, the number and size of filters in each layer, and the weights of each element of each filter. Details of the motion inference model used by the tracking device 1 will be described later.

[0019] The storage device 2 may be an external storage device such as a hard disk connected to or built into the tracking device 1, or it may be a portable storage medium such as flash memory. Furthermore, the storage device 2 may be a server device that communicates data with the tracking device 1. Also, the storage device 2 may consist of multiple devices.

[0020] The display device 3 displays information based on the control of the tracking device 1. Examples of the display device 3 include displays, projectors, etc. When the display device 3 receives a display signal supplied from the tracking device 1, it displays information based on the received display signal.

[0021] Input device 4 is an interface that accepts user input, which is external input based on user operations using the tracking system 100. Examples include a touch panel, buttons, keyboard, and voice input device. Input device 4 supplies input signals generated based on user input to the tracking device 1. Camera 5 is one or more cameras that photograph the area to be monitored for tracking, and the generated images are stored in the time-series image storage unit D1.

[0022] The configuration of the tracking system 100 shown in Figure 1 is an example, and various modifications may be made to this configuration. For example, the tracking device 1, storage device 2, display device 3, input device 4, and camera 5 may be integrated into any combination. The tracking system 100 may also be equipped with sound output devices such as speakers. Furthermore, the tracking device 1 may be composed of multiple devices. In this case, the multiple devices constituting the tracking device 1 exchange information among themselves that is necessary to execute pre-assigned processes.

[0023] (2) Hardware configuration Figure 2 shows the hardware configuration of the tracking device 1. The tracking device 1 includes a processor 11, memory 12, and interface 13 as hardware components. The processor 11, memory 12, and interface 13 are connected via a data bus 19.

[0024] The processor 11 executes predetermined processes by running programs stored in memory 12. The processor 11 is a processor such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), or TPU (Tensor Processing Unit). The processor 11 may be composed of multiple processors. The processor 11 is an example of a computer.

[0025] Memory 12 is composed of various volatile and non-volatile memories, such as RAM (Random Access Memory) and ROM (Read Only Memory). Memory 12 also stores programs for the tracking device 1 to perform various processes. Furthermore, Memory 12 is used as working memory to temporarily store information obtained from the storage device 2. Memory 12 may also function as storage device 2. Similarly, storage device 2 may function as memory 12 of the tracking device 1. The programs executed by the tracking device 1 may be stored in a storage medium other than memory 12.

[0026] Interface 13 is an interface for electrically connecting the tracking device 1 with other devices. These interfaces may be wireless interfaces such as network adapters for wirelessly transmitting and receiving data with other devices, or they may be hardware interfaces for connecting with other devices via cables, etc.

[0027] The hardware configuration of the tracking device 1 is not limited to the configuration shown in Figure 2. For example, the tracking device 1 may include at least one of the display device 3 or the input device 4. Furthermore, the tracking device 1 may be connected to or have a built-in sound output device such as a speaker.

[0028] (3) Overview of the tracking process This section outlines the tracking process performed by tracking device 1. In general terms, tracking device 1 identifies the person corresponding to a tracking ID based on the consistency between the inference result of hypothetical behavior, which is obtained by tentatively linking each of multiple individuals whose tracking IDs cannot be distinguished to a respective tracking ID, and the inference result of behavior of the target being tracked, which is based on past images. This allows tracking device 1 to consider the transitions in behavior and achieve robust tracking even in situations such as industrial sites where apparent information is not useful and overlaps occur between tracked targets.

[0029] Figure 3 shows an example of the functional blocks of the tracking device 1. As shown in Figure 3, the processor 11 of the tracking device 1 functionally includes a tracking candidate detection unit 21, a trajectory prediction unit 22, a tracking information matching unit 23, a first motion inference unit 24, a second motion inference unit 25, a motion information matching unit 26, a third motion inference unit 27, and a tracking information management unit 28. In Figure 3, blocks where data is exchanged are connected by solid lines, but the combination of blocks where data is exchanged is not limited to this. The same applies to the diagrams of other functional blocks described later.

[0030] The tracking candidate detection unit 21 acquires the target image, which is the latest image generated at the reference time t, from the time-series image storage unit D1, detects tracking target candidates (in this case, people) from the acquired target image, and generates location information (also called "detected location information") for the detected candidates. The detected tracking target candidates will hereafter be called "tracking candidates." In this case, the tracking candidate detection unit 21 may use any object detection model to generate a bounding box surrounding the region of the person from the target image. In this case, the object detection model is, for example, a deep learning model, and is machine-trained to output location information representing the bounding box of the person in the input image when an image is input. Parameters for configuring the object detection model are determined in advance by machine learning and stored in the storage device 2, etc. The tracking candidate detection unit 21 may also use an object detection model such as instance segmentation to detect the region of a tracking candidate that has any shape other than a rectangle. The tracking candidate detection unit 21 supplies the target image along with the detected location information representing the region of each tracking candidate to the trajectory prediction unit 22. Furthermore, the tracking candidate detection unit 21 extracts past images from the time series image storage unit D1 for the period from time "t-Δ" to "t-1" (Δ is an integer of 2 or more) for use in a later block, and supplies these past images to the trajectory prediction unit 22.

[0031] The trajectory prediction unit 22 predicts the position of each person to be tracked in the target image based on the trajectory information of each tracked person to which a tracking ID has been assigned. In this case, the trajectory prediction unit 22 refers to the tracking information stored in the tracking information storage unit D2 and identifies the time-series position information of the tracked person for each tracking ID as trajectory information. In this case, the trajectory prediction unit 22 may use any object tracking algorithm, such as a Kalman filter, to predict the position of the tracked person in the target image to which a tracking ID has been assigned in past images. Examples of the object tracking algorithms mentioned above include SORT (Simple Online and Realtime Tracking) and ByteTrack. The trajectory prediction unit 22 then generates predicted position information indicating the predicted position of the tracked person in the target image for each tracking ID and supplies the predicted position information for each tracking ID to the tracking information matching unit 23. The predicted position information is information that represents the area of ​​the tracked person in the target image, for example, it represents the bounding box.

[0032] The tracking information matching unit 23 compares the predicted location information for each tracking ID generated by the trajectory prediction unit 22 with the detected location information for each tracking candidate generated by the tracking candidate detection unit 21 to identify the tracking candidate corresponding to each tracking ID. In this case, the tracking information matching unit 23 identifies the detected location information that is most similar to the estimated location information of each tracking ID, considers the identified detected location information to correspond to the tracking target at the reference time of each tracking ID, and associates it with each tracking ID. The tracking information matching unit 23 also determines whether there are tracking candidates that cannot be distinguished in correspondence with the tracking ID. That is, the tracking information matching unit 23 determines whether there are detected location information that represents multiple tracking candidates that may correspond to a single tracking ID due to location overlap. A specific example of this method for determining whether they can be distinguished will be described later. The tracking information matching unit 23 then supplies the detected location information that could be distinguished in correspondence with the tracking ID (i.e., associated with the tracking ID) to the third operation inference unit 27, along with the associated tracking ID and images (target image and past image). Meanwhile, the tracking information matching unit 23 supplies the third operation inference unit 27, along with the image, with detected location information representing tracking candidates that cannot be distinguished in correspondence with the tracking ID.

[0033] The first motion inference unit 24 assumes that for each tracking ID, an indistinguishable tracking candidate is the target of tracking, and generates motion information representing the inference result of the tracking target's motion at reference time t under each assumption, based on the target image and past images. In other words, the first motion inference unit 24 assumes that a possible tracking candidate is the target of tracking for each tracking ID, and generates motion information at reference time t under each assumption as an inference result. Hereafter, the motion information at reference time t under each assumption will also be referred to as "assumption-based motion information".

[0034] The first motion inference unit 24 generates assumption-based motion information using a machine learning-based motion inference model. In this case, the first motion inference unit 24 generates inference images for the time intervals "t-Δ" to "t" that represent the tracked target corresponding to each tracking ID, based on past images from time intervals "t-Δ" to "t-1" and the target image at reference time t. Then, the first motion inference unit 24 uses the inference images for the time intervals "t-Δ" to "t" for each tracking ID and the machine learning-based motion inference model to infer the motion at reference time t, and generates assumption-based motion information representing the inference result. The first motion inference unit 24 supplies the assumption-based motion information for each assumption and the detection position information used in each assumption to the motion information matching unit 26. Hereafter, the motion inference model used by the first motion inference unit 24 will also be referred to as "motion inference model M1". The inference images for the time intervals "t-Δ" to "t" are an example of a "sequence of images representing the tracked target".

[0035] The second motion inference unit 25 generates motion information (also called "motion information based on past images") representing the motion of the tracked object at a reference time t, based on past images, for each tracking ID. In this case, the second motion inference unit 25 extracts past images from time "t-Δ" to "t-1" from the time-series image storage unit D1 and extracts location information for each past image associated with each tracking ID from the tracking information storage unit D2. Then, based on the extracted past images and location information, the second motion inference unit 25 generates inference images for time "t-Δ" to "t-1" for each tracking ID. Then, using the inference images for time "t-Δ" to "t-1" for each tracking ID and a machine learning-based motion inference model, the second motion inference unit 25 infers the motion at a reference time t and generates motion information based on past images as an inference result. The inference images for time "t-Δ" to "t-1" are an example of a "sequence of images representing the tracked object".

[0036] The motion inference model used by the second motion inference unit 25 is a model that has been pre-trained to output an inference result of the motion of the tracked object in the next time-series image obtained (i.e., the Δ+1th image) when Δ time-series images of the tracked object are input. As will be described later, the second motion inference unit 25 may, instead of using past images, infer motion information at reference time t by extrapolation based on motion information at times "t-Δ" to "t-1" stored in the tracking information storage unit D2. The second motion inference unit 25 supplies past image-based motion information corresponding to each tracking ID to the motion information matching unit 26. Hereafter, the motion inference model used by the second motion inference unit 25 will also be referred to as "motion inference model M2".

[0037] The motion information matching unit 26 identifies the correspondence between the assumption-based motion information supplied from the first motion inference unit 24 and the past image-based motion information supplied from the second motion inference unit 25. Specifically, for each tracking ID, the motion information matching unit 26 identifies the assumption-based motion information that is consistent with (i.e., most similar to) the past image-based motion information, and determines that the detected location information used to generate the identified assumption-based motion information represents the tracking target of the tracking ID. As a result, for each tracking ID where the corresponding detected location information could not be distinguished, the motion information matching unit 26 links the identified assumption-based motion information with the detected location information. The motion information matching unit 26 then supplies the set of tracking ID, motion information, and detected location information to the tracking information management unit 28 for each tracking ID.

[0038] The third operation inference unit 27 infers the operation of the tracked object at a reference time t for each tracked ID associated with the detected location information by the tracking information matching unit 23, based on the target image, past images, and the location information of the tracked object in the target image and past images. In this case, first, the third operation inference unit 27 obtains past images from time "t-Δ" to "t-1" from the time-series image storage unit D1, and obtains location information associated with each tracked ID in those past images from the tracking information storage unit D2. Then, based on the obtained past images and location information, the third operation inference unit 27 generates a time-series inference image by cutting out the tracked object for each tracked ID from the past images from time "t-Δ" to "t-1". In addition, the third operation inference unit 27 generates an inference image by cutting out the tracked object for each tracked ID from the target image based on the target image and the detected location information of the tracked candidate corresponding to each tracked ID in the target image. As a result, the third operation inference unit 27 obtains an inference image for each tracked ID from time "t-Δ" to "t". The third motion inference unit 27 then uses the inference images from time "t-Δ" to "t" and a machine-learned motion inference model to infer the motion at reference time t for each tracking ID, and generates motion information representing the inference result. In this case, the motion inference model is a model that has been pre-machine-learned to output the inference result of the motion of a person in the last image of the input time-series images when a predetermined number (in this case, Δ+1 images) of time-series images representing a specific person are input. The third motion inference unit 27 then supplies the tracking ID, the motion information at reference time t, and the detected location information to the tracking information management unit 28. Hereafter, the motion inference model used by the third motion inference unit 27 will also be referred to as "motion inference model M3".

[0039] Furthermore, the third motion inference unit 27 may perform the following processes: assigning a new tracking ID to the detected location information of a tracking candidate that does not correspond to any of the existing tracking IDs, and inferring the motion of the tracking target of the newly assigned tracking ID based on the target image. In this case, the third motion inference unit 27 supplies the tracking information management unit 28 with a set of the newly assigned tracking ID, the detected location information, and the motion information inferred based on the target image.

[0040] The tracking information management unit 28 updates the tracking information storage unit D2 based on the information supplied by the operation information matching unit 26 and the third operation inference unit 27. Specifically, it adds operation information and detection location information at reference time t to the tracking information for each tracking ID registered in the tracking information storage unit D2. In this case, the tracking information management unit 28 updates the tracking information for tracking IDs that the tracking information matching unit 23 determines can be distinguished based on the information supplied by the third operation inference unit 27, and updates the tracking information for tracking IDs that the tracking information matching unit 23 determines cannot be distinguished based on the information supplied by the operation information matching unit 26.

[0041] Here, the tracking candidate detection unit 21, trajectory prediction unit 22, tracking information matching unit 23, first motion inference unit 24, second motion inference unit 25, motion information matching unit 26, third motion inference unit 27, and tracking information management unit 28 can be realized, for example, by the processor 11 executing a program. Alternatively, the necessary programs may be recorded on any non-volatile storage medium and installed as needed to realize each component. At least a portion of these components may be realized not only by software programs, but also by any combination of hardware, firmware, and software. Furthermore, at least a portion of these components may be realized using a user-programmable integrated circuit, such as an FPGA (Field-Programmable Gate Array) or a microcontroller. In this case, the program composed of the above components may be realized using this integrated circuit. Also, at least a portion of each component may be composed of an ASSP (Application Specific Standard Produce), ASIC (Application Specific Integrated Circuit), or quantum processor (quantum computer control chip). Thus, each component may be realized by various hardware. The same applies to other embodiments described later. Furthermore, each of these components may be realized through the collaboration of multiple computers, for example, using cloud computing technology.

[0042] (4) Determination of whether or not a distinction is possible Next, we will explain a specific example of the distinction determination, which is the determination of whether or not the detected location information associated with the tracking ID can be distinguished.

[0043] Figure 4(A) is a schematic diagram illustrating a specific example of the process for determining whether or not objects can be distinguished. Figure 4(A) shows a specific example of determining whether or not objects can be distinguished when a target image at a reference time t is obtained in which multiple people overlap in the image. Here, it is assumed that the tracking targets with tracking IDs "x" and "y" are being tracked in past images.

[0044] In this case, the tracking candidate detection unit 21 generates detection location information corresponding to three tracking candidates from the target image. Here, the detection location information corresponding to the three tracking candidates is represented by bounding boxes Pg, Ph, and Pi. The trajectory prediction unit 22 generates predicted location information for tracking IDs "x" and "y" at reference time t, based on the location information of past images of tracking IDs "x" and "y". Here, the predicted location information for tracking IDs "x" and "y" is represented by bounding boxes Px and Py.

[0045] The tracking information matching unit 23 then calculates the degree of overlap of pairs extracted from bounding boxes Pg, Ph, and Pi representing detected position information and bounding boxes Px and Py representing predicted position information. Here, the degree of overlap is the degree of overlap of the bounding boxes on the image, and an index such as IoU is used. If the position information is information based on the orientation of an object (for example, the position information of key joints), an index such as OKS (Object Keypoint Similarity) may be used as the degree of overlap.

[0046] Figure 4(B) is Table T1, which shows the degree of overlap on the image of pairs extracted from bounding boxes Pg, Ph, and Pi representing detected location information and bounding boxes Px and Py representing predicted location information. Here, the degree of overlap is represented by a range of values ​​from a minimum of 0 to a maximum of 1. For bounding box Px, the degree of overlap with bounding box Pg is 0.8, and the degree of overlap with bounding box Pi is 0.6, and these values ​​are approximate. Therefore, the tracking information matching unit 23 determines that the detected location information corresponding to bounding boxes Pg and Pi, respectively, cannot be distinguished in relation to the tracking ID "x".

[0047] In more detail, first, the tracking information matching unit 23 determines the detected location information that matches (aligns with) the predicted location information for each tracking ID, based on a matching method such as the Hungarian algorithm. Next, the degree of overlap between the detected location information and the predicted location information (in this case, the degree of overlap of the bounding boxes) is set to "O", and the degree of overlap of the detected location information that matches the predicted location information is set to "Om". Then, if the degree of overlap O is greater than or equal to a predetermined threshold, and there is detected location information with an overlap degree O that satisfies the following equation in relation to the degree of overlap Om, the tracking information matching unit 23 determines that the detected location information is indistinguishable from the matched detected location information. Om > O> Om - s, where s is a real number.

[0048] For example, suppose the predetermined threshold is "0.5", the real number s is "0.3", and the bounding box Px representing the predicted position information and the bounding box Pg representing the detected position information match. In this case, the overlap Om of the bounding box Px is 0.8, and the overlap O (=0.6) of the bounding box Pi is greater than or equal to the threshold of 0.5, and "0.8>O>0.5 (=0.8-0.3)", satisfying the above equation. Therefore, the tracking information matching unit 23 determines that the detected position information corresponding to the bounding boxes Pg and Pi cannot be distinguished.

[0049] (5) Generating assumption-based operational information Next, we will specifically explain the generation of assumption-based motion information by the first motion inference unit 24.

[0050] Figure 5 shows a specific example of the input and output of the motion inference model M1 used by the first motion inference unit 24. In Figure 5, two bounding boxes Pj and Pk representing tracking candidates are obtained as detection position information for the image at reference time t. The detection position information represented by these bounding boxes Pj and Pk may correspond to the tracking target with tracking ID "z", and the tracking information matching unit 23 has determined that they cannot be distinguished.

[0051] In this case, the first motion inference unit 24 sets a first assumption that the target being tracked at reference time t for tracking ID "z" is represented by detection position information corresponding to the bounding box Pj, and a second assumption that the target being tracked at reference time t for tracking ID "z" is represented by detection position information corresponding to the bounding box Pk. Then, the first motion inference unit 24 uses the motion inference model M1 to obtain the motion inference result for reference time t based on the first assumption and the motion inference result for reference time t based on the second assumption. In this case, the first motion inference unit 24 generates time-series inference images based on each assumption and inputs the time-series inference images to the motion inference model M1 to obtain the inference result output from the motion inference model M1. Here, the first motion inference unit 24 generates the inference image for reference time t under the first assumption based on the target image and the bounding box Pj, and generates the inference image for reference time t under the second assumption based on the target image and the bounding box Pk. Furthermore, the first motion inference unit 24 generates inference images for the time intervals "t-Δ" to "t-1" used in each assumption, based on past images from the time intervals "t-Δ" to "t-1" and the positional information associated with those past images.

[0052] Here, the motion inference model M1 outputs a sequence of scores (i.e., a score vector) representing the likelihood of a given motion type (i.e., a class of motion) as an inference result. Examples of given motion types include "cart transport," "heavy machinery operation," and "compaction work." The first motion inference unit 24 then uses the score vector output by the motion inference model M1 as assumption-based motion information. The assumption-based motion information may be the scores of all motion types output by the motion inference model M1, or it may be the scores of the top predetermined number of motion types.

[0053] Here, the motion inference model M1 is, for example, a neural network that has undergone machine learning, and examples of such neural networks include SlowFast and VideoMAE. The time-series inference images input to the motion inference model M1 may be time-series images cropped from the region of the target being tracked (for example, cropped from the bounding box), time-series images cropped from the region of surrounding objects of the target being tracked, such as tools used by the target being tracked, or time-series pose information of the target being tracked. In this case, the pose information is the position information of the joint points on the image of the person being tracked.

[0054] (6) Generation of motion information based on past images Next, we will explain the first and second generation examples, which are examples of generating motion information based on past images by the second motion inference unit 25.

[0055] Figure 6(A) shows an overview of the process for generating past image-based motion information for the tracking target with tracking ID "z" based on the first generation example.

[0056] In the first generation example, the second motion inference unit 25 generates inference images for the time range "t-Δ" to "t-1" based on past images from time range "t-Δ" to "t-1" and the corresponding location information of the tracking ID "z". Then, the second motion inference unit 25 uses the motion inference model M2 to input the time-series inference images into the motion inference model M2 and obtains the inference results output from the motion inference model M2. In this case, the inference results output by the motion inference model M2 are data in the same format as the inference results output by the motion inference model M1, for example, a score vector of the expected motion types. The second motion inference unit 25 then uses the score vector output by the motion inference model M2 as motion information based on past images. The motion information based on past images may be the scores of all motion types output by the motion inference model M2, or it may be the scores of the top predetermined number of motion types. The motion inference model M2 is, for example, a neural network that has undergone machine learning, and examples of such neural networks include SlowFast and VideoMAE. Furthermore, the time-series inference images input to the motion inference model M2 may be time-series images cropped from the area of ​​the target being tracked, time-series images cropped from the area of ​​surrounding objects of the target being tracked, such as tools used by the target being tracked, or time-series pose information of the target being tracked.

[0057] Figure 6(B) shows an overview of the process for generating motion information based on past images, using the second generation example. Specifically, Figure 6(B) shows the distribution of scores for a particular motion type.

[0058] In the second generation example, instead of using an inference image, the second motion inference unit 25 extracts motion information for tracking ID "z" from the tracking information storage unit D2 at times "t-Δ" to "t-1", and extrapolates the score at reference time t based on the time-series score represented by the extracted motion information. In this case, for example, the second motion inference unit 25 infers the score at reference time t from the scores at times "t-Δ" to "t-1" for each type of motion. In this case, the motion inference model M2 is an algorithm that implements an arbitrary extrapolation method, and when motion information (i.e., a score vector) at times "t-Δ" to "t-1" is input, it outputs motion information at reference time t.

[0059] Even with the second generation example, the second motion inference unit 25 can generate a score vector at reference time t from past score vectors. Then, the tracking device 1 can achieve robust matching of motion transitions by predicting the future motion of the tracked object using the first or second generation example.

[0060] (7) Verification of operation information Next, the matching of assumption-based operation information and past image-based operation information by the operation information matching unit 26 will be specifically explained. The operation information matching unit 26 calculates a cost according to the similarity between the score vector shown by the assumption-based operation information and the score vector shown by the past image-based operation information, and determines a matching between the assumption-based operation information and the past image-based operation information that minimizes the cost. In this case, the cost is an arbitrary index value representing the similarity between the vectors (cosine similarity, L2 norm, etc.) or its reciprocal, and is set such that, for example, the cost increases as the score vectors become more similar. The matching of assumption-based operation information and past image-based information is determined by an arbitrary matching method, such as the Hungarian algorithm. Then, for each tracking ID, the operation information matching unit 26 adopts the detection position information used in the assumption of the assumption-based operation information that matches the past image-based operation information as the detection position information at reference time t.

[0061] Figure 7 shows an example of the relationship between each set assumption and the inference results output by the motion inference model M1. Here, there are tracking candidates Cm and Cn that correspond to indistinguishable detection location information for tracking ID 1 and tracking ID 2. The first motion inference unit 24 sets assumptions x1, y1, x2, and y2, assuming that tracking candidates Cm and Cn represent the tracking targets for tracking ID 1 and tracking ID 2, respectively, and obtains inference results 1a, 1b, 2a, and 2b corresponding to each assumption from the motion inference model M1. These inference results correspond to assumption-based motion information, respectively.

[0062] Figure 8 shows an overview of the process by which the motion inference model M2 generates past image-based motion information for each of the tracking IDs 1 and 2. In the example in Figure 8, the first motion inference unit 24 inputs the time-series images of the tracked object based on past images from time "t-Δ" to "t-1" for each of the tracking IDs 1 and 2 to the motion inference model M2, thereby obtaining inference results 1c and 2c output by the motion inference model M2. The inference results 1c and 2c correspond to past image-based motion information.

[0063] Figure 9 shows an overview of the matching process between the inference results shown in Figure 7 and the inference results shown in Figure 8. In this case, the operation information matching unit 26 calculates a cost based on the similarity of the score vectors for all combinations of inference results 1a, 1b, 2a, 2b and inference results 1c, 2c. The matrix shown in Figure 9 represents the cost of the corresponding combination of inference results. Here, the cost of inference result 1c and inference result 1a, and the cost of inference result 2c and inference result 2b are both at their maximum value of 1.0, so the operation information matching unit 26 determines that inference result 1c for tracking ID "1" and inference result 1a for assumption x1, and inference result 2c for tracking ID "2" and inference result 2b for assumption y2 are matched. Therefore, the operation information matching unit 26 adopts assumptions x1 and y2, and determines that tracking candidate Cm is the tracking target for tracking ID 1 at reference time t, and tracking candidate Cn is the tracking target for tracking ID 2 at reference time t.

[0064] In this way, by incorporating movement information into the matching of tracking IDs and distinguishing people's information in more detail, it becomes possible to achieve robust tracking even in situations where people overlap. In other words, by introducing a matching method that takes movement transitions into consideration, robust tracking can be achieved even in situations such as industrial settings where apparent information is not useful and people overlap.

[0065] (8) Processing flow Figure 10 is an example of a flowchart showing the processing procedure performed by the tracking device 1.

[0066] First, the tracking device 1 detects tracking candidates from the target image at a reference time t corresponding to the current processing time, and predicts the position of the target at reference time t for each tracking ID based on past images (step S11). As a result, the tracking device 1 generates detected position information of tracking candidates present in the target image and predicted position information for each tracking ID. The processing in step S11 corresponds to the processing performed by the tracking candidate detection unit 21 and the trajectory prediction unit 22.

[0067] Next, the tracking device 1 performs a distinction determination based on the detected location information of the tracking candidate and the predicted location information for each tracking ID, identifies the tracking ID corresponding to the distinguishable detected location information, and infers the operation at reference time t for the identified tracking ID (step S12). The processing in step S12 corresponds to the processing performed by the tracking information matching unit 23 and the third operation inference unit 27.

[0068] Next, the tracking device 1 determines whether or not there are multiple indistinguishable detection location information (step S13). If there are no multiple indistinguishable detection location information (step S13; No), the tracking information at reference time t based on the processing result in step S12 is stored in the tracking information storage unit D2 (step S17).

[0069] On the other hand, if there are multiple indistinguishable detection location pieces (step S13; Yes), the tracking device 1 makes an assumption for each indistinguishable detection location piece for each tracking ID and infers the operation at reference time t for each assumption (step S14). In this way, the tracking device 1 generates assumption-based operation information. The processing in step S14 corresponds to the processing performed by the first operation inference unit 24. Then, for each tracking ID, the tracking device 1 infers the operation at reference time t based on past images (step S15). In this way, the tracking device 1 generates past image-based operation information. The processing in step S15 corresponds to the processing performed by the second operation inference unit 25.

[0070] Then, based on the matching of the assumption-based operation information and the past image-based operation information, the tracking device 1 identifies the detection location information corresponding to the tracking ID for which the corresponding detection location information could not be identified in step S12 (step S16). In this case, for each tracking ID, the tracking device 1 identifies the detection location information used to generate the assumption-based operation information that matches the past image-based operation information. The processing in step S16 corresponds to the processing performed by the operation information matching unit 26. Then, the tracking device 1 stores the tracking information at the reference time t in the tracking information storage unit D2 (step S17). The processing in step S17 corresponds to the processing performed by the tracking information management unit 28.

[0071] Next, the tracking device 1 determines whether or not to terminate the tracking process (step S18). If the tracking device 1 determines to terminate the tracking process (step S18; Yes), it terminates the flowchart process. On the other hand, if the tracking device 1 determines not to terminate the tracking process (step S18; No), it uses the newly obtained image from camera 5 as the target image obtained at reference time t and returns to step S11.

[0072] (9) Application examples According to the above embodiment, the latest tracking information based on the images generated by the camera 5 is stored in the storage device 2, and the tracking system 100 can automatically record the actions of individuals. By analyzing this tracking information, the work actions of each worker can be visualized, improving productivity and safety. Improvements in generation include automation of work recording, detection of work delays and errors, work efficiency analysis, optimization of personnel allocation, and other operational efficiency improvements. Improvements in safety include alerts for unsafe behavior, near-miss monitoring, and prevention of other industrial accidents. Furthermore, in warehousing, manufacturing, and construction industries, the actions of each worker can be accurately grasped based on the tracking information, and personnel resources can be optimized. In manufacturing and warehousing industries, the actions of each worker can be accurately grasped based on the tracking information, and this can be used for work assurance and training support. In warehousing and manufacturing industries, it is also conceivable that the time-series movements of work objects (including robots) can be grasped and used for automating item handling.

[0073] Furthermore, the tracking device 1 may display information about the tracked object in real time based on the tracking information. A specific example of real-time display processing will be described below.

[0074] Figure 11(A) is a first example of the image display screen showing the latest image generated by camera 5, and Figure 11(B) is a first example of the motion score display screen showing the transition of motion scores for each worker being tracked. The tracking device 1 generates display information by referring to the time-series image storage unit D1 and the tracking information storage unit D2, and transmits the generated display information to the display device 3, thereby displaying at least one of the image display screen and the motion score display screen on the display device 3.

[0075] In the image display screen shown in Figure 11(A), the tracking device 1 displays, for each of the tracked workers A and B, who are assigned a tracking ID, hypothetical-based operation information and scores based on past image-based operation information for the identified operation. Specifically, based on the tracking information for worker A corresponding to the reference time, the tracking device 1 displays on the image that worker A is performing "compaction work" in association with worker A. Furthermore, the tracking device 1 displays the score based on past image-based operation information (corresponding to "Prediction from trajectory: compaction work 0.7") and the score based on matched hypothetical-based operation information (corresponding to "Current: compaction work: 0.7") in association with worker A. Similarly, based on the tracking information corresponding to the reference time, the tracking device 1 displays on the image that worker B is performing "cart transport" in association with worker B. Furthermore, tracking device 1 displays a score based on past image-based operation information (corresponding to "Prediction from trajectory: Cart transport: 0.6") and a score based on matched assumption-based operation information (corresponding to "Current cart transport: 0.8"), associated with worker A.

[0076] By displaying such an image screen, the tracking device 1 can allow the user to understand in detail the inference result for the worker's current work type, along with a score representing the likelihood of the inference result.

[0077] Furthermore, the motion score display screen shown in Figure 11(B) graphically represents the time-series score indicating the likelihood of each work type for each worker. Here, the dashed graph labeled "Prediction from Trajectory" corresponds to a graph showing the time change of the score based on motion information based on past images, while the solid graph labeled "Current" corresponds to a graph showing the time change of the score based on assumption-based motion information. In Figure 11(B), the work type with the highest score is indicated above the graph along the arrow representing the time axis, and in the case of worker A, "Cart Transport" is followed by "Compaction". In addition, the motion score display screen is provided with a scroll bar 70, and by operating the scroll bar 70, it is possible to display the graph corresponding to any worker being tracked.

[0078] By displaying such an operation score screen, the tracking device 1 allows the user to confirm the time-series inference results for any worker's work type.

[0079] Furthermore, the tracking device 1 may perform a process to predict future work types and display the work type prediction results on the image display screen and the operation score display screen.

[0080] Figure 12(A) shows a second example of the image display screen, and Figure 12(B) shows a second example of the motion score display screen. In the image display screen shown in Figure 12(A), the tracking device 1 predicts the actions of worker A and worker B at a time after the reference time t (for example, time t+1) based on the detected position information of worker A and worker B at the reference time t, and displays a score based on the motion information representing the predicted actions as "Future prediction from trajectory". For example, if the motion inference model M2 is a model that predicts actions based on time-series images, the tracking device 1 generates Δ time-series inference images based on motion information and images for the past Δ time periods including the reference time t. Then, the tracking device 1 inputs the generated inference images into the motion inference model M2 to obtain the future time-series motion information output by the motion inference model M2. Furthermore, if the motion inference model M2 is an extrapolation model, the tracking device 1 acquires motion information for future times based on motion information for past Δ time points including the reference time t and the motion inference model M2. Also, similar to the image display screen in the first display example, the tracking device 1 displays on the image that worker A is performing "compaction work" and worker B is performing "cart transport" based on the tracking information corresponding to the reference time t, associating it with each worker. In addition, based on the motion information at the reference time t (assumed motion information), the tracking device 1 associates the current score of "0.7" for "compaction work" with worker A and the current score of "0.8" for "cart transport" with worker B, associating them on the image.

[0081] Furthermore, the motion score display screen shown in Figure 12(B) graphically represents the time-series score indicating the likelihood of each work type for each worker. Here, the dashed graph labeled "Prediction from Trajectory" corresponds to a graph showing the time change of the score based on past image-based motion information and predicted motion information, while the solid graph labeled "Current" corresponds to a graph showing the time change of the score based on assumption-based motion information. As shown in Figure 12(B), the graph labeled "Prediction from Trajectory" also shows predicted score values ​​beyond the reference time. By displaying the motion score display screen according to the second display example, the tracking device 1 allows the user to confirm the time-series inference results, including future predictions for any worker's work type.

[0082] <Second Embodiment> Figure 13 is a block diagram of the tracking device 1X. The tracking device 1X comprises a first inference means 24X, a second inference means 25X, and a identification means 26X. The tracking device 1X may be composed of multiple devices.

[0083] The first inference means 24X performs a first inference when a target image containing multiple objects is obtained at a reference time, by assuming that for each of the multiple objects, the tracked target, tracked based on time-series images obtained before the reference time, is one of the multiple objects, and infers the behavior of the tracked target at the reference time. That is, when the target image at the reference time contains the 1st to Nth (N is an integer greater than or equal to 2) objects, the first inference means 24X assumes that the tracked target at the reference time is the 1st to Nth object, and infers N patterns of behavior for one tracked target. The first inference means 24X can be, for example, the first behavior inference unit 24 in the first embodiment.

[0084] The second inference means 25X performs a second inference, which infers the operation of the tracked object at a reference time based on the time-series image. In this case, the second inference means 25X infers one pattern of operation for one tracked object. The second inference means 25X can be, for example, the second operation inference unit 25 in the first embodiment.

[0085] The identification means 26X identifies the object representing the tracking target in the target image obtained at the reference time, based on the inference results for each assumption from the first inference and the inference results from the second inference. In this case, the identification means 26X identifies which of the 1st to Nth objects the tracking target is. The identification means 26X can be, for example, the operation information matching unit 26 in the first embodiment.

[0086] Figure 14 is an example of a flowchart showing the processing procedure of the tracking device 1X. First, the first inference means 24X performs a first inference (step S21) inferring the behavior of the tracked object at the reference time, assuming that for each of the multiple objects, the tracked object, which was tracked based on time-series images obtained before the reference time, is one of the multiple objects, when a target image containing multiple objects is obtained at the reference time. Next, the second inference means 25X performs a second inference (step S22) inferring the behavior of the tracked object at the reference time based on the time-series images. Then, the identification means 26X identifies the object representing the tracked object in the target image obtained at the reference time (step S23) based on the inference results for each assumption from the first inference and the inference results from the second inference.

[0087] According to the second embodiment, the tracking device 1X can perform robust tracking even in situations where apparent information is not useful and objects overlap.

[0088] In each of the embodiments described above, the program can be stored using various types of non-transitory computer-readable medium and supplied to a computer, such as a processor. Non-transitory computer-readable mediums include various types of tangible storage mediums. Examples of non-transitory computer-readable mediums include magnetic storage mediums (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage mediums (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). The program may also be supplied to the computer by various types of transient computer-readable mediums. Examples of transient computer-readable mediums include electrical signals, optical signals, and electromagnetic waves. Transitory computer-readable mediums can supply the program to the computer via wired communication channels such as electric wires and optical fibers, or via wireless communication channels.

[0089] Furthermore, some or all of the above embodiments may also be described as follows, but are not limited to the following. Also, some or all of the configurations described in Appendices 2 to 9, which are dependent on Appendice 1 above, may be similarly dependent on Appendices 10 and 11. Moreover, not limited to the devices, methods, and storage media described in the appendices, some or all of the configurations described in the appendices may also be dependent on methods, various hardware, software, various recording means for recording software (including storage media), or systems, without departing from the above embodiments.

[0090] [Note 1] A first inference means performs a first inference to infer the behavior of the tracked object at the reference time, assuming that the tracked object tracked based on time-series images obtained before the reference time is one of the multiple objects, when a target image including multiple objects is obtained at a reference time. A second inference means performs a second inference to infer the operation of the tracked object at the reference time based on the time-series image, A means for identifying the object representing the tracking target in the target image based on the inference results for each assumption in the first inference and the inference results in the second inference, A tracking device having a tracking device. [Note 2] The tracking device according to Appendix 1, wherein the identifying means identifies the inference result most similar to the inference result of the second inference from the inference results for each assumption, and identifies the object representing the tracking target based on the assumption in the identified inference result. [Note 3] The system further includes determination means for determining whether it is possible to distinguish which of the plurality of objects the tracking target is, The tracking device according to Appendix 1 or 2, wherein the first inference means performs the first inference when it is determined that the distinction cannot be made. [Note 4] The system further includes object detection means for detecting the regions of the multiple objects from the target image, The determination means is a tracking device as described in Appendix 3, which determines whether or not the distinction can be made based on the degree of overlap of the regions. [Note 5] The first inference means generates a sequence of images representing the tracked object based on the time-series images and the target image for each assumption, and infers the behavior at the reference time based on the sequence of images and the machine learning model. The tracking device described in any one of the appendices 1 to 4, wherein the machine learning model is a model trained to output an inference result of the movement of an object when a time-series image representing the object is input. [Note 6] The tracking device according to any one of the appendices 1 to 4, wherein the second inference means infers the operation at the reference time by extrapolation based on operation information representing the operation of the tracked subject prior to the reference time, which is generated based on the time-series image. [Note 7] The second inference means generates a sequence of images representing the target being tracked based on the time-series images, and infers the behavior at the reference time based on the sequence of images and a machine learning model. The tracking device described in any one of the appendices 1 to 6, wherein the machine learning model is a model trained to output an inference result of the predicted movement of an object when a time-series image representing the object is input. [Note 8] A tracking device according to any one of the appendices 1 to 7, further comprising display control means for displaying information representing the movement of an object identified as a tracking target on a display device in association with the target image, based on at least one of the inference results for each of the above assumptions or the inference results from the second inference. [Note 9] A tracking device according to any one of the appendices 1 to 8, wherein the inference results for each assumption and the inference results from the second inference show a score representing the probability for each assumed type of operation. [Note 10] A tracking method performed by a computer, When a target image containing multiple objects is obtained at a reference time, a first inference is performed to infer the behavior of the tracked object at the reference time, assuming that the tracked object, based on time-series images obtained before the reference time, is one of the multiple objects. A second inference is performed to infer the behavior of the tracked object at the reference time based on the aforementioned time-series images. Based on the inference results for each assumption in the first inference and the inference results in the second inference, the object representing the tracking target is identified in the target image. Tracking method. [Note 11] When a target image containing multiple objects is obtained at a reference time, a first inference is performed to infer the behavior of the tracked object at the reference time, assuming that the tracked object, based on time-series images obtained before the reference time, is one of the multiple objects. A second inference is performed to infer the behavior of the tracked object at the reference time based on the aforementioned time-series images. Based on the inference results for each assumption in the first inference and the inference results in the second inference, the object representing the tracking target is identified in the target image. A program that instructs a computer to perform a process. [Note 12] A storage medium containing the program described in Appendix 11.

[0091] Although the present invention has been described above with reference to embodiments, the present invention is not limited to the above embodiments. Various modifications to the structure and details of the present invention can be made as understood by those skilled in the art within the scope of the present invention. That is, the present invention includes the full disclosure, including the claims, and of course, various modifications and alterations that those skilled in the art could make in accordance with the technical idea. Furthermore, the above-mentioned patent and non-patent disclosures are incorporated herein by reference. [Explanation of Symbols]

[0092] 1. 1X Tracking Device 2 Storage device 3 Display device 4 Input devices 5 Cameras 11 processors 12 memory 13 Interfaces 100 Tracking Systems

Claims

1. A first inference means performs a first inference to infer the behavior of the tracked object at the reference time, assuming that the tracked object tracked based on time-series images obtained before the reference time is one of the multiple objects, when a target image including multiple objects is obtained at a reference time. A second inference means performs a second inference to infer the operation of the tracked object at the reference time based on the time-series images, A means for identifying the object representing the tracking target in the target image based on the inference results for each assumption in the first inference and the inference results in the second inference, A tracking device having a tracking device.

2. The tracking device according to claim 1, wherein the identifying means identifies the inference result that is most similar to the inference result of the second inference from the inference results for each assumption, and identifies the object representing the tracking target based on the assumption in the identified inference result.

3. The system further includes determination means for determining whether it is possible to distinguish which of the plurality of objects the tracking target is, The tracking device according to claim 1, wherein the first inference means performs the first inference when it is determined that the distinction cannot be made.

4. The system further includes object detection means for detecting the regions of the multiple objects from the target image, The tracking device according to claim 3, wherein the determination means determines whether or not the distinction can be made based on the degree of overlap of the regions.

5. The first inference means generates a sequence of images representing the tracked object based on the time-series images and the target image for each assumption, and infers the operation at the reference time based on the sequence of images and the machine learning model. The tracking device according to claim 1, wherein the machine learning model is a model trained to output an inference result of the movement of an object when a time-series image representing the object is input.

6. The tracking device according to claim 1, wherein the second inference means infers the operation at the reference time by extrapolation based on operation information representing the operation of the tracked object prior to the reference time, which is generated based on the time-series image.

7. The second inference means generates a sequence of images representing the target being tracked based on the time-series images, and infers the behavior at the reference time based on the sequence of images and a machine learning model. The tracking device according to claim 1, wherein the machine learning model is a model trained to output an inference result of the predicted movement of an object when a time-series image representing the object is input.

8. The tracking device according to claim 1, further comprising display control means for displaying information representing the movement of an object identified as a tracking target on a display device in association with the target image, based on at least one of the inference results for each of the above assumptions or the inference results from the second inference.

9. A tracking method performed by a computer, When a target image containing multiple objects is obtained at a reference time, a first inference is performed to infer the behavior of the tracked object at the reference time, assuming that the tracked object, based on time-series images obtained before the reference time, is one of the multiple objects. A second inference is performed to infer the behavior of the tracked object at the reference time based on the aforementioned time-series images. Based on the inference results for each assumption in the first inference and the inference results in the second inference, the object representing the tracking target is identified in the target image. Tracking method.

10. When a target image containing multiple objects is obtained at a reference time, a first inference is performed to infer the behavior of the tracked object at the reference time, assuming that the tracked object, based on time-series images obtained before the reference time, is one of the multiple objects. A second inference is performed to infer the behavior of the tracked object at the reference time based on the aforementioned time-series images. Based on the inference results for each assumption in the first inference and the inference results in the second inference, the object representing the tracking target is identified in the target image. A program that instructs a computer to perform a process.