Object monitoring device and object monitoring method

The object monitoring device tracks objects by linking video information across multiple areas using attribute detection and prediction, addressing the challenge of feature amount changes to maintain target visibility.

JP2025099723APending Publication Date: 2025-07-03KONICA MINOLTA INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023216611
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing object monitoring systems struggle to track objects when their feature amounts change, leading to potential loss of the target due to altered characteristics, such as a person stopping and starting to walk.

Method used

An object monitoring device that includes an attribute detection unit, a predicted attribute estimation unit, and a tracking unit to link video information across multiple monitoring areas based on detected and predicted attributes, using a machine learning model to estimate future attributes and associate video information to track the object.

Benefits of technology

Enables effective tracking of objects even when their feature amounts change, reducing the risk of losing the target by linking video information across areas based on predicted attributes and generating captions for time-series changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025099723000001_ABST
    Figure 2025099723000001_ABST
Patent Text Reader

Abstract

To provide a target object monitoring device capable of tracking a target object even if a feature amount of the target object changes.SOLUTION: A target object monitoring device 1 that monitors a target object by video information of a plurality of monitoring areas according to the present invention, includes: an attribute detection unit 11 that detects an attribute of the target object based on the video information; a predicted attribute estimation unit 12 that estimates a predicted attribute from time-series change in the attribute; and a tracking unit 15 that tracks the target object by associating video information of another monitoring area with the target object based on the attribute detected by the attribute detection unit and the predicted attribute predicted by the predicted attribute estimation unit when the target object fades out from one monitoring area.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an object monitoring device that monitors an object moving with a plurality of cameras and a method for monitoring the object.

Background Art

[0002] It is common practice to install a plurality of cameras inside a facility, display the video data of each camera on a monitor provided in a monitoring room, and have a monitor visually monitor the inside of the facility. In this monitoring system, a monitor simultaneously monitors the displays of a plurality of monitors to monitor the inside of the facility.

[0003] For this reason, the burden on the monitor is large, and there is a risk of overlooking an object. To solve this problem, a technique for linking related cameras has been devised. For example, Patent Document 1 discloses an information processing apparatus including a detection unit that links video data from a plurality of cameras based on the similarity of image feature amounts (face and body images), integrally analyzes the data, and detects a target person image; a tracking unit that tracks the target person based on the detection result of the target person image; and an output unit that outputs the tracking result of the target person.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] According to the above prior art, the monitoring burden on a monitor based on video data from a plurality of cameras can be reduced. However, in the technique of Patent Document 1, an image feature amount of a target person image is used as a query to extract a person whose image feature amount of video data is similar. For this reason, when a person stops running and starts walking, the feature amount changes, so the target person cannot be extracted and the target may be lost.

[0006] An object of the present invention is to provide an object monitoring device capable of tracking an object even when a feature amount of the object changes.

Means for Solving the Problems

[0007] The above problems are solved as follows.

[0008] (1) In an object monitoring device that monitors an object based on video information of a plurality of monitoring areas, an attribute detection unit that detects an attribute of the object based on the video information, a predicted attribute estimation unit that estimates a predicted attribute from a time-series change of the attribute, and when the object fades out from one monitoring area, based on the attribute detected by the attribute detection unit and the predicted attribute predicted by the predicted attribute estimation unit, a tracking unit that links the video information of another monitoring area and the object to track the object.

[0009] (2) In the object monitoring device according to (1), the tracking unit quantifies the attribute of an object detected from the video information of another monitoring area based on the attribute of the object and the predicted attribute predicted from the time-series change of the attribute, and sets the attribute with the largest calculated numerical value as the attribute of the object with a high degree of coincidence, and links the video information in which the attribute is detected to the object to track the object.

[0010] (3) In the object monitoring device according to (2), the tracking unit scores each attribute, ranks the options of the predicted attribute based on the prediction probability, and weights them according to the priority order.

[0011] (4) In the object monitoring device according to (1), further comprising a caption generation unit that generates a caption from the time-series change of the attribute.

[0012] (5) In an object monitoring device that monitors an object based on video information of a plurality of monitoring areas, an attribute detection unit that detects the attributes of the object based on the video information, a caption generation unit that generates a caption from the time-series change of the attributes, a predicted caption estimation unit that estimates a predicted caption from the time-series change of the attributes, and when the object fades out from one monitoring area, based on the caption detected by the caption generation unit and the predicted caption predicted by the predicted caption estimation unit, a tracking unit that links the video information of another monitoring area and the object to track the object. An object monitoring device comprising:

[0013] (6) The object monitoring device according to (5), wherein the predicted caption estimation unit estimates a predicted caption by a machine learning model.

[0014] (7) The object monitoring device according to (5), wherein the tracking unit determines it as an abnormality when it fails to link the object and the video information of another monitoring area within a predetermined time after the object fades out from one monitoring area.

[0015] (8) The object monitoring device according to (1) or (5), further comprising a monitoring area information storage unit that stores monitoring area information indicating the positional relationship of the monitoring areas and composed of the area name, the camera that captures the monitoring area, and information on adjacent areas, and the tracking unit sets the monitoring area indicated by the information on the adjacent areas of the monitoring area information as another monitoring area to be linked to the object.

[0016] (9) The object monitoring device according to (8), wherein when the camera is a PTZ camera, the monitoring area information storage unit stores monitoring area information according to the shooting range of the camera.

[0017] (10) A target object monitoring method for monitoring a target object based on video information of a plurality of monitoring areas, comprising: detecting an attribute of the target object based on the video information; estimating a predicted attribute from a time-series change of the attribute; and when the target object fades out from one monitoring area, tracking the target object by associating video information of another monitoring area with the target object based on the attribute and the predicted attribute.

Advantages of the Invention

[0018] According to the present invention, it is possible to provide a target object monitoring device capable of tracking a target object even when a feature amount of the target object changes.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Embodiments for Carrying Out the Invention

[0020] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. FIG. 1 is a diagram showing a configuration of a monitoring system for monitoring a target object by a target object monitoring device according to an embodiment.

[0021] The monitoring system consists of a plurality of cameras 2a to 2e, a display device 4, a network 3, and an object monitoring device 1. The plurality of cameras 2a to 2e each capture images of different monitoring areas. The display device 4 is composed of a plurality of monitors and displays the images of the monitoring areas captured by each camera. The network 3 connects the cameras 2a to 2e, the display device 4, and the object monitoring device 1 and communicates video information. The object monitoring device 1 displays the video information of the cameras 2a to 2e acquired via the network 3 on the display device 4 and tracks the monitored objects in the video information.

[0022] The object monitoring device 1 includes a video information receiving unit (not shown) that receives the video information of the cameras 2a to 2e via the network 3, and a display output unit (not shown) that outputs the received video information of the cameras 2a to 2e to the display devices 4a to 4e corresponding to the cameras 2a to 2e.

[0023] Furthermore, the object monitoring device 1 includes an attribute detection unit 11, a predicted attribute estimation unit 12, a caption generation unit 13, a tracking unit 15, a monitoring area information storage unit 16, and a caption information storage unit 17. The attribute detection unit 11 detects the attributes of the target object based on the video information of the cameras 2a to 2e received by the video information receiving unit. The predicted attribute estimation unit 12 predicts the attributes based on the changes in the attributes of the target object as well. The caption generation unit 13 generates a caption for the target object based on the time-series changes in the attributes of the target object. The tracking unit 15 links the video information and the target object based on the relevance of the attributes of the target object in the video information and tracks the target object. The monitoring area information storage unit 16 stores the positional relationship of the shooting areas of the cameras 2a to 2e. The caption information storage unit 17 stores the history of the captions.

[0024] Note that the object monitoring device 1 is configured to include the processing units and storage units of the attribute detection unit 11, the predicted attribute estimation unit 12, the caption generation unit 13, the tracking unit 15, the monitoring area information storage unit 16, and the caption information storage unit 17 according to their functions.

[0025] Specifically, the target object monitoring device 1 is an information processing device composed of a ROM that stores programs, a CPU that executes the programs, a RAM that serves as a work area, and an HDD / SSD that stores programs and processing data.

[0026] The processing units of the attribute detection unit 11, the predicted attribute estimation unit 12, the caption generation unit 13, and the tracking unit 15 of the target object monitoring device 1 realize their functions when the CPU executes based on the programs in the ROM or HDD / SSD. Also, the monitoring area information storage unit 16 and the caption information storage unit 17 are configured in the HDD / SSD.

[0027] Figure 2 is a diagram showing the positional relationship of the areas of the shooting information. The target object monitoring device 1 detects the attributes of the target object based on the respective video information of the areas 5A, 5B, 5C, 5D, and 5E, and associates them with the video information (area) in which the target object is included. Here, the video information of the areas 5A, 5B, 5C, 5D, and 5E is the video information captured by the cameras 2a, 2b, 2c, 2d, and 2e, respectively.

[0028] Then, assuming that the target object is a moving object, the target object monitoring device 1 detects the target object corresponding to the detected attribute for the video information after a predetermined time has passed, and associates it with the video information (area) in which the target object is included. In this way, the target object monitoring device 1 sequentially associates the target object specified by the attribute with the video information (area) to track the target object. The attributes will be described later.

[0029] Here, the example in Figure 2 will be described. Suppose the target object monitoring device 1 detects a target object having a predetermined attribute in area 5A (video information of camera 2a). Assuming that the target object has moved, the target object monitoring device 1 determines the presence or absence of a target object corresponding to the attribute of the target object detected in area 5A (video information of camera 2a) for the video information of cameras 2a to 2e after a predetermined time has passed.

[0030] At this time, the target object monitoring device 1 determines in the order of decreasing probability of the target object moving (in the order of adjacent areas or in the order of proximity), i.e., in the order of area 5A, area 5B, area 5D, area 5C, and area 5E. Further, the target object monitoring device 1 may determine the presence or absence of the target object only in adjacent areas.

[0031] When the target object monitoring device 1 detects a target object corresponding to the attribute of the target object detected in area 5A from the video information of camera 2b, it associates the video information of camera 2b (area 5B) with the target object specified by the attribute. Thereby, the target object monitoring device 1 detects that the target object detected in area 5A has moved to area B. The target object monitoring device 1 repeats this to track the target object detected in area 5A.

[0032] FIG. 3 corresponds to FIG. 2 and shows the configuration of the monitoring area information stored in the monitoring area information storage unit 16 (FIG. 1) indicating the positional relationship of the monitoring areas. The monitoring area information is composed of the area name, the camera that photographs the area, and information on adjacent areas. The adjacent area consists of the direction in which the adjacent area is located and the area name. Also, a distance may be added to the adjacent area. For example, when the area name is area 5A, it indicates that the area is photographed by camera 2a, area 5B is located to the east, and area 5D is located to the southeast.

[0033] FIG. 3 shows a case where there is no restriction on the area movement of the target object. However, in FIG. 2, when there is a wall between area 5A and area 5B and direct movement is not possible, the adjacent area of area 5A in FIG. 3 is only "southeast: area 5D", and in area 5B, "west: area 5A" is not shown.

[0034] When the camera that photographs the video of the monitoring area is a camera whose angle of view can be changed, such as a PTZ camera having functions of pan (horizontal head swing), tilt (vertical head swing), and zoom (telephoto / wide angle), the positional relationship of the camera's shooting area changes according to the change in the angle of view. Therefore, the information indicating the positional relationship stored in the monitoring area information storage unit 16 in FIG. 3 stores a plurality of information corresponding to the shooting range of the PTZ camera. Alternatively, the monitoring area information storage unit 16 stores information indicating the positional relationship at predetermined time intervals.

[0035] Next, the operation of the target object monitoring device 1 regarding the tracking of one target object will be described with reference to the flowchart of FIG. 4.

[0036] In step S1, the attribute detection unit 11 detects the attributes of the target object from the video information of the target object. The attributes of the target object when the target object to be monitored is a person are information indicating the nature and characteristics of the target object, such as the gender of female or male, the length of hair such as long hair or short hair, the color of eyes such as blue eyes or black eyes, the action state such as walking or running, the type of clothing, and the type of shoes.

[0037] In step S2, if there is a time-series change in the attributes detected in step S1, the caption generation unit 13 generates a caption explaining the change and stores it in the caption information storage unit 17. Specifically, the caption generation unit 13 qualitatively captures the time-series change of the attributes from features such as the change in the speed, acceleration, angular velocity, and direction of the action, taking the speed, acceleration, angular velocity, and direction of the action as attributes.

[0038] In addition, when the obtained captions of "bend down" and "stand up" are repeated, the caption generation unit 13 converts them to "flexing". Also, when the captions of "carry a load from point A" and "release the load at point B" are generated, the caption generation unit 13 converts them to "carry the load from A to B". In this way, by taking action recognition into account, the caption generation unit 13 can obtain richer information as a time-series change than local attribute detection.

[0039] In step S3, the target object monitoring device 1 determines whether the target object fades out in the subsequent video information in time series of the video information in which the attributes are detected. If the target object monitoring device 1 determines that the target object fades out (Yes in S3), it proceeds to step S4. If not (No in S3), it returns to step S1. That is, the target object monitoring device 1 performs attribute detection and caption generation until the target object exits from one area.

[0040] In step S4, the prediction attribute estimation unit 12 generates a prediction attribute for tracking the target object from the time-series change of the attributes of the target object detected in step S1. That is, the target object monitoring device 1 tracks, as the movement destination of the target object, the imaging area corresponding to the video information in which an attribute with a high degree of coincidence with the generated prediction attribute (including the detected attribute) of the target object is detected among the video information of the cameras in a plurality of imaging areas.

[0041] Specifically, the prediction attribute estimation unit 12 generates a plurality of prediction attribute options with prediction probabilities based on the time-series data of the attributes of the target object by using a behavior model based on machine learning such as deep learning. At this time, the prediction attribute estimation unit 12 also trains the behavior model based on the time-series information of the attributes of the target object detected in step S1.

[0042] In step S5, the tracking unit 15 ranks the prediction attribute options obtained by the prediction attribute estimation unit 12 based on the prediction probability, and weights them according to the priority order.

[0043] In step S6, the tracking unit 15 scores the attributes of the target object to be tracked from the feature amount and variability of the attributes of the target object.

[0044] In step S7, the tracking unit 15 quantifies the attributes of the objects detected from the video information of a plurality of cameras after a predetermined time when the target object has faded out, based on the weighting of the prediction attributes obtained in step S5 and the scores of the attributes obtained in step S6. Then, the tracking unit 15 sets the attribute with the largest calculated value as the attribute of the object with a high degree of coincidence, and associates the camera (video information) that captured the video information with the detected attribute with the target object. That is, the tracking unit 15 tracks the target object after fading out on the assumption that it exists in the area captured by the camera associated with the target object.

[0045] At this time, specifically, the tracking unit 15 refers to the monitoring area information storage unit 16 and acquires, as the moving destination of the object to be tracked, the area adjacent to the furnace area where the object was located before determining the fade-out in step S3. Then, the tracking unit 15 identifies the camera that captures the acquired area and detects the attributes of the object from the video information of the identified camera.

[0046] Hereinafter, specific examples will be described. ≪Example A≫ First, when the object to be tracked (female) runs from the area 5B captured by camera 2b to the area C captured by camera 2c, the process of linking the female as the same object to be tracked between cameras 2b and 2c will be described corresponding to the flowchart of FIG. 4.

[0047] As step S1, the object monitoring device 1 detects the attributes of the object to be tracked: gender [female], height [170 cm], hair length [long hair], eye color [blue eyes], and action state [running].

[0048] As step S2, the object monitoring device 1 obtains the time-series change of the attributes of the object to be tracked based on the change in the action state of the object to be tracked, such as "running at 5 km / h", "running at 4 km / h", and "running at 3 km / h". Then, from the time-series change, captions such as "the female is running while stalling" or "the 170-cm, long-haired, blue-eyed female is running while stalling" are generated. Although the action attribute of the "action state" of the object to be tracked has not changed, it can be recognized from the time-series change of the attributes that the object is running while stalling, and by using this as a caption, the state of the object to be tracked can be shown in detail.

[0049] When the object monitoring device 1 recognizes from the video information in which the attributes and captions of the object to be tracked are generated that the object to be tracked has faded out, which corresponds to step S3, the following object tracking process is performed.

[0050] As step S4, the target object monitoring device 1 generates, from the time-series change of the behavior state of the target object, the attributes of the behavior state of the target object, namely "running as it is", "stopping from stalling", and "walking", as the predicted attributes after fade-out. That is, the target object monitoring device 1 predicts that the target object after fade-out is one of "a woman is running", "a woman is stopping", and "a woman is walking".

[0051] Next, as step S5, the target object monitoring device 1 prioritizes the predicted attributes of the target object, namely "running", "walking", and "stopping", in this order. Then, as step S7, the target object monitoring device 1 numerically values the priorities as attribute weightings, i.e., "running: 0.6", "walking: 0.3", and "stopping: 0.1".

[0052] As step S6, the target object monitoring device 1 scores according to the feature amounts and variability rates of the attributes of the target object, namely "height", "hair length", "eye color", and "behavior state". Hereinafter, for the attribute of "gender", it is fixed as female and the explanation is omitted.

[0053] Specifically, for the attribute of "height", since the variability rate is low, the score is high and it is scored as 100. For the attribute of "hair length", since the variability rate is high but the feature amount is large, the score is medium and it is scored as 70. For the attribute of "eye color", since the variability rate is low and the feature amount is high, the score is high and it is scored as 130. For the attribute of "behavior state", since the variability rate is high and the options are expanded by the predicted attributes, the score is low and it is scored as 50.

[0054] As step S7, the target object monitoring device 1 detects the attributes of the object, namely height [170 cm], hair length [long hair], eye color [blue eyes], and behavior state [walking], detected from the video information of the camera 2c that captured the area 5C a predetermined time after the target object faded out. Then, the target object monitoring device 1 numerically values based on the weighting of the predicted attributes obtained in step S5 and the scores of the attributes obtained in step S6, and calculates 100 + 70 + 130 + 50 * 0.3 = 315.

[0055] In addition, the target object monitoring device 1 detects the attributes of the object such as height [170 cm], hair length [long hair], eye color [black eyes], and action state [running] detected from the video information of the camera 2d in the area 5D, and calculates 100 + 70 + 0 + 50 * 0.6 = 200 as the quantified attribute. Furthermore, the target object monitoring device 1 detects the attributes of the object such as height [170 cm], hair length [shaved head], eye color [black eyes], and action state [running] detected from the video information of the camera 2e in the area 5E, and calculates 100 + 0 + 0 + 50 * 0.6 = 130 as the quantified attribute.

[0056] Then, the target object monitoring device 1 regards the attributes of the object detected from the video information of the camera 2c in the area 5C with the largest calculated value as the attributes of the object with the highest degree of match, and associates the camera 2c with the target object. That is, the target object monitoring device 1 assumes that the target object exists in the area 5C photographed by the associated camera 2c, and tracks the target object after fade-out.

[0057] Based on the above tracking results, the target object monitoring device 1 captions "The woman ran decelerating from area 5B towards area 5C and walked into area 5C".

[0058] ≪Case I≫ Suppose that the target object monitoring device 1 detects, in chronological order, the attributes of the target object such as gender [male], hair length [short hair], clothing [black T-shirt], shoes [red shoes], and action state [running], and the attributes of gender [male], hair length [short hair], clothing [black T-shirt], shoes [red shoes], and action state [fall into the blind spot outside the area] from the video information of the area A photographed by the camera 2a. In this case, the target object monitoring device 1 can assume from the chronological change of the action state attribute that the object was active but fell into a blind spot or tumbled for some reason.

[0059] Then, the target object monitoring device 1 generates "running again" and "being carried" as prediction attributes of the action state in this case as normal patterns, increases their priorities, and assigns larger weights. When an attribute of the action state of "lying down" other than the normal pattern is detected from the video information of cameras 2b to 2e other than camera 2a, it can be regarded as an abnormality.

[0060] If an attribute of the target object's gender [male], hair length [short hair], clothing [black T-shirt], shoes [red shoes], and action state [running] or an attribute of the target object's gender [male], hair length [short hair], clothing [black T-shirt], shoes [red shoes], and action state [being carried] is detected from the video information of camera 2b, it is regarded as the attribute of the object with the highest degree of matching, and camera 2b and the target object are linked. That is, the target object monitoring device 1 tracks the target object after fade-out on the assumption that the target object exists in the area 5B photographed by the linked camera 2b.

[0061] If the above attributes are not detected in the video information of any camera, it is detected as an abnormality, and an alarm is issued or recorded.

[0062] In the above, it has been explained that when the target object monitoring device 1 fades out from one monitoring area, the video information of other monitoring areas and the target object are linked based on the attributes detected by the attribute detection unit and the predicted attributes predicted by the predicted attribute estimation unit to track the target object. However, it is also possible to track the target object based on the caption.

[0063] FIG. 5 is a diagram for explaining the configuration of the target object monitoring device 1 that tracks the target object based on the caption. The target object monitoring device 1 of the monitoring system in FIG. 1 is different in that it includes a predicted caption estimation unit 14 that predicts the caption based on the time-series change of the attributes of the target object instead of the predicted attribute estimation unit 12.

[0064] Next, differences in the operation of the target object monitoring device 1 regarding the tracking of one target object in FIG. 4 will be described. The method of tracking based on the caption of the target object differs in steps S4 to S7 from the method of tracking based on the attributes of the target object. Details will be described below.

[0065] In step S4, the prediction caption estimation unit 14 generates a plurality of prediction caption options having prediction probabilities for tracking the target object from the time-series change of the attributes of the target object detected in step S1 by using a behavior model based on machine learning such as deep learning.

[0066] In step S5, the tracking unit 15 ranks the prediction caption options obtained by the prediction caption estimation unit 14 based on the prediction probability, and weights them according to the priority.

[0067] In step S6, the tracking unit 15 quantifies the caption of the target object from the feature amount and variability of the attributes of the target object to be tracked.

[0068] In step S7, the tracking unit 15 quantifies the caption of the object detected from the video information of a plurality of cameras after a predetermined time when the target object has faded out, based on the weighting of the prediction caption obtained in step S5 and the score of the caption obtained in step S6. Then, the tracking unit 15 sets the caption with the largest calculated numerical value as the caption of the object with a high degree of coincidence, and associates the camera that captured the video information where the caption was detected with the target object. That is, the tracking unit 15 tracks the target object after fading out on the assumption that it exists in the area captured by the camera associated with the target object.

[0069] As described above, the target object monitoring device 1 can track the target object using the caption of the object. However, if the tracking unit 15 fails to associate the target object with the video information of another monitoring area within a predetermined time after the target object fades out from one monitoring area, that is, if the caption of the object cannot be detected corresponding to the predicted caption, it may be regarded as an abnormality and an alarm may be issued.

[0070] The present invention is not limited to the above-described embodiments, and includes various modifications. The above embodiments have been described in detail for easy understanding of the present invention, and are not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment can be replaced with the configuration of another embodiment, and the configuration of another embodiment can also be added to the configuration of one embodiment.

Explanation of Reference Numerals

[0071] 1 Target object monitoring device 11 Attribute detection unit 12 Predicted attribute estimation unit 13 Caption generation unit 14 Predicted caption estimation unit 15 Tracking unit 16 Monitoring area information storage unit 17 Caption information storage unit 2a~2e Camera 3 Network 4 Display device 5A~5E Area

Claims

1. In an object monitoring device that monitors a target object based on video information of a plurality of monitoring areas, an attribute detection unit that detects an attribute of the target object based on the video information; a predicted attribute estimation unit that estimates a predicted attribute from the time-series change of the attribute; a tracking unit that, when the target object fades out from one monitoring area, associates the video information of another monitoring area with the target object based on the attribute detected by the attribute detection unit and the predicted attribute predicted by the predicted attribute estimation unit, and tracks the target object; An object monitoring device comprising:

2. In the object monitoring device according to claim 1, the tracking unit quantifies the attributes of the objects detected in the video information of other monitoring areas based on the attributes of the target object and the predicted attributes predicted from the time-series changes of the attributes, regards the attribute with the largest calculated numerical value as the attribute of the object with a high degree of coincidence, associates the video information in which the attribute is detected with the target object, and tracks the target object, Object monitoring device.

3. In the object monitoring device according to claim 2, the tracking unit scores each attribute, prioritizes the options of the predicted attributes based on the prediction probability, and weights them according to the priority, Object monitoring device.

4. In the object monitoring device according to claim 1, further, a caption generation unit that generates a caption from the time-series change of the attribute, An object monitoring device comprising:

5. In an object monitoring device that monitors a target object based on video information of a plurality of monitoring areas, an attribute detection unit that detects an attribute of the target object based on the video information; a caption generation unit that generates a caption from the time-series change of the attribute; a predicted caption estimation unit that estimates a predicted caption from the time-series change of the attribute; a tracking unit that, when the target object fades out from one monitoring area, associates the video information of another monitoring area with the target object based on the caption detected by the caption generation unit and the predicted caption predicted by the predicted caption estimation unit, and tracks the target object; An object monitoring device comprising:

6. In the object monitoring device according to claim 5, the predicted caption estimation unit estimates a predicted caption by a machine learning model, Object monitoring device.

7. In the object monitoring device according to claim 5, When the tracking unit fails to associate the target object with the video information of another monitoring area within a predetermined time after the target object fades out from one monitoring area, it is regarded as abnormal. Target object monitoring device.

8. In the target object monitoring device according to claim 1 or claim 5, A monitoring area information storage unit that stores monitoring area information indicating the positional relationship of the monitoring areas, and is composed of the area name, the camera that captures the monitoring area, and information on adjacent areas. The tracking unit uses the monitoring area indicated by the information on the adjacent area of the monitoring area information as another monitoring area associated with the target object. Target object monitoring device.

9. In the target object monitoring device according to claim 8, When the camera is a PTZ camera, the monitoring area information storage unit stores the monitoring area information according to the shooting range of the camera. Target object monitoring device.

10. A target object monitoring method for monitoring a target object by video information of a plurality of monitoring areas, comprising: Detecting an attribute of the target object based on the video information; Estimating a predicted attribute from the time-series change of the attribute; When the target object fades out from one monitoring area, tracking the target object by associating the video information of another monitoring area with the target object based on the attribute and the predicted attribute; A target object monitoring method including the above steps.

Citation Information

Patent Citations

  • Information processing device, information processing method, and program

    JP2022030832A