Learning device, data change estimation device, learning method, data change estimation method, learning program, and data change estimation program

The learning device and method address the inability of conventional techniques to detect changes in target areas by training a score calculator to analyze first and second data sets in both orders, enabling efficient change estimation.

JP7736613B2Active Publication Date: 2025-09-09SECOM CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022052590
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-28
Publication Date
2025-09-09
Estimated Expiration
2042-03-28

AI Technical Summary

Technical Problem

Conventional image analysis techniques fail to compare first and second data sets to estimate changes such as the appearance or disappearance of a target area between them.

Method used

A learning device and method that trains a score calculator to input features of first and second data in a predetermined and reverse order to calculate scores indicating the disappearance or appearance of an object, allowing for the estimation of these changes with a simple configuration.

Benefits of technology

Enables the estimation of object disappearance and appearance by calculating scores in both directions, providing a simple and effective method for change detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007736613000030
    Figure 0007736613000030
  • Figure 0007736613000031
    Figure 0007736613000031
  • Figure 0007736613000032
    Figure 0007736613000032
Patent Text Reader

Abstract

To enable training a score calculator that calculates both a score indicating disappearance of an object from first data to second data and a score indicating appearance of the object.SOLUTION: A learning device includes a learning unit which trains a score calculator such that the score calculator receives input of a feature amount of first data and a feature amount of second data in a prescribed order based on learning data, for calculating a score indicating a specific change that is disappearance or appearance of an object represented by the data, the change is from the first data to the second data, and receives input of the feature amount of the first data and the feature amount of the second data in a reverse order of the prescribed order, for calculating a score indicating the specific change in the change from the second data to the first data.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a learning device, a data change estimation device, a learning method, a data change estimation method, a learning program, and a data change estimation program. [Background technology]

[0002] 2. Description of the Related Art Conventionally, a technique is known for extracting a target region, which is a region of a target to be extracted from an image, using a neural network.

[0003] For example, in the technology described in Patent Document 1, a convolutional neural network is used to extract features for each segmented region of an input image. Then, for each segmented region that overlaps with the inspection region, an integrated feature is calculated using a weight according to the degree of overlap and the extracted feature, and a target region is extracted based on the integrated feature. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2021-114223 Summary of the Invention [Problem to be solved by the invention]

[0005] However, conventional technology only extracts a target area from an image, and is not capable of comparing first data, such as a past image, with second data, such as a current image, to estimate changes such as the appearance or disappearance of a target area between the first data and the second data.

[0006] Therefore, the present invention has been made in consideration of the above circumstances, and aims to provide a learning device, a learning method, and a learning program that can estimate the disappearance and appearance of an object by training a score calculator that calculates both a score indicating the disappearance or appearance of an object from first data to second data and a score indicating the disappearance or appearance of an object from second data to first data.

[0007] Another object of the present invention is to provide a data change estimation device, a data change estimation method, and a data change estimation program that can estimate the disappearance and appearance of an object by calculating, with a simple configuration, both a score indicating the disappearance or appearance of an object from first data to second data and a score indicating the disappearance or appearance of an object from second data to first data. [Means for solving the problem]

[0008] In order to achieve the above object, the learning device of the present invention is configured to include: a learning unit that trains a score calculator to input features of the first data and features of the second data in a predetermined order based on learning data indicating a specific change from first data to second data, which is the disappearance or appearance of an object represented by the data, and learning data indicating the specific change from the second data to the first data; the learning unit inputs features of the first data and features of the second data into the score calculator in a reverse order to the predetermined order, and calculates a score indicating the specific change from the second data to the first data.

[0009] According to the learning device of the present invention, a learning unit inputs features of the first data and features of the second data into a score calculator in a predetermined order based on learning data that indicates a specific change from first data to second data, which is the disappearance or appearance of an object represented by the data, and learning data that indicates the specific change from the second data to the first data, to calculate a score that indicates the specific change in the change from the first data to the second data, and trains the score calculator to input features of the first data and features of the second data into the score calculator in a reverse order to the predetermined order, and calculates a score that indicates the specific change in the change from the second data to the first data.

[0010] In this way, the score calculator is trained to input the feature amounts of the first data and the feature amounts of the second data into a predetermined order to calculate a score indicating a specific change from the first data to the second data, which is the disappearance or appearance of an object represented by the data, and to input the feature amounts of the first data and the feature amounts of the second data into the score calculator in the reverse order of the predetermined order to calculate a score indicating the specific change in the change from the second data to the first data. In this way, the disappearance and appearance of an object can be estimated by training the score calculator to calculate both a score indicating the disappearance or appearance of an object from the first data to the second data and a score indicating the disappearance or appearance of an object from the second data to the first data.

[0011] A data change estimation device according to the present invention includes: a score calculator that is trained to input features of the third data and the fourth data in a predetermined order and calculate a score indicating the specific change in the change from the third data to the fourth data, based on training data indicating a specific change from the fourth data to the third data, which is the disappearance or appearance of an object represented by the data, and training data indicating the specific change from the fourth data to the third data; and a score calculation unit that inputs features of the third data and the fourth data to the score calculator in a predetermined order and calculates the score indicating the specific change in the change from the fourth data to the third data, and inputs the features of the third data and the fourth data in the reverse order to the predetermined order to the score calculator; and

[0012] According to the data change estimation device of the present invention, a score calculator is trained in the same manner as the above-described learning device. A score calculation unit inputs the feature amounts of first data and the feature amounts of second data in a predetermined order to the score calculator trained by the learning unit to calculate a score indicating the specific change in the change from the first data to the second data, and inputs the feature amounts of the first data and the feature amounts of the second data to the score calculator in a reverse order to the predetermined order to calculate a score indicating the specific change in the change from the second data to the first data.

[0013] In this way, the feature amounts of the first data and the feature amounts of the second data are input to the score calculator in a predetermined order to calculate a score indicating the specific change in the change from the first data to the second data, and the feature amounts of the first data and the feature amounts of the second data are input to the score calculator in the reverse order to the predetermined order to calculate a score indicating the specific change in the change from the second data to the first data. This makes it possible to estimate the disappearance or appearance of an object by calculating, with a simple configuration, both a score indicating the disappearance or appearance of an object from the first data to the second data and a score indicating the disappearance or appearance of an object from the second data to the first data.

[0014] The data change estimation device according to the present invention includes a score calculation unit that inputs features of first data and features of second data into a score calculator in a predetermined order and calculates a score indicating a specific change from the first data to the second data, which is the disappearance or appearance of an object represented by the data, and inputs the features of the first data and features of the second data into the score calculator in a reverse order to the predetermined order and calculates a score indicating the specific change in the change from the second data to the first data.

[0015] According to the data change estimation device of the present invention, the score calculation unit inputs the features of the first data and the features of the second data to a score calculator in a predetermined order to calculate a score indicating the specific change in the change from the first data to the second data, and also inputs the features of the first data and the features of the second data to the score calculator in a reverse order to the predetermined order to calculate a score indicating the specific change in the change from the second data to the first data.

[0016] In this way, the feature amounts of the first data and the feature amounts of the second data are input to the score calculator in a predetermined order to calculate a score indicating the specific change in the change from the first data to the second data, and the feature amounts of the first data and the feature amounts of the second data are input to the score calculator in the reverse order to the predetermined order to calculate a score indicating the specific change in the change from the second data to the first data. This makes it possible to estimate both the disappearance and appearance of an object from the first data to the second data with a simple configuration.

[0017] The score calculator may also calculate the score using asymmetric operations.

[0018] In addition, the score calculation unit can further calculate a score indicating no change and a score indicating a change of subject using the score indicating the specific change from the first data to the second data and the score indicating the specific change from the second data to the first data.

[0019] In addition, the above-mentioned data change estimation device can further include a feature map extraction unit that extracts a feature map for each of the first data and the second data, a candidate area estimation unit that estimates a target likelihood map, which is a map of likelihood indicating the likelihood of a target area, from the feature map using a target area estimation model, a target area extraction unit that extracts a target area from the target likelihood map, and a feature extraction unit that extracts features of the extracted target area from the feature map.

[0020] The data change estimation device further includes a score integrating unit, wherein the target region extraction unit extracts a plurality of target regions from the target likelihood map, the feature extraction unit extracts a feature for each of the plurality of target regions from the feature map, the score calculation unit inputs the feature of the target region of the first data and the feature of the target region of the second data to a score calculator in a predetermined order for each combination of the target region of the first data and the target region of the second data to calculate a score indicating the specific change from the first data to the second data, and inputs the feature of the target region of the first data and the feature of the target region of the second data to the score calculator in a reverse order to the predetermined order to calculate a score indicating the specific change from the second data to the first data, and the score integrating unit can calculate a maximum value, a minimum value, or an average value of the score indicating the specific change from the first data to the second data, and a maximum value, a minimum value, or an average value of the score indicating the specific change from the first data to the second data.

[0021] Moreover, the above-mentioned data change estimation device further includes a score integrating unit, wherein the first data is a plurality of first data measured at different positions, and the second data is a plurality of second data measured at the same position as each of the plurality of first data but at different measurement timings, and the score calculation unit inputs the feature of the first data and the feature of the second data to a score calculator in a predetermined order for each combination of the first data and the second data to calculate a score indicating the specific change from the first data to the second data, and inputs the feature of the first data and the feature of the second data to the score calculation model in a reverse order to the predetermined order to calculate a score indicating the specific change from the second data to the first data, and the score integrating unit can calculate a maximum value or a minimum value of the score indicating the specific change from the first data to the second data, and a maximum value or a minimum value of the score indicating the specific change from the second data to the first data.

[0022] The data change estimation device further includes a score integrating unit, wherein the feature map extracting unit extracts a plurality of types of feature maps for each of the first data and the second data by a neural network, the feature amount extracting unit extracts the feature amount from each of the plurality of types of feature maps for each of the first data and the second data, and the score calculating unit inputs the feature amount of the first data and the feature amount of the second data to a score calculator in a predetermined order for each of the feature amounts extracted from each of the plurality of types of feature maps, and calculates the specific change from the first data to the second data. the score integration unit can calculate a score indicating the specific change from the first data to the second data, input the feature amounts of the first data and the feature amounts of the second data into the score calculation model in an order reverse to the predetermined order, and calculate a score indicating the specific change from the second data to the first data; and the score integration unit can integrate the scores indicating the specific change from the first data to the second data, which are calculated for each of the feature amounts extracted from each of the multiple types of feature maps, and integrate the scores indicating the specific change from the second data to the first data, which are calculated for each of the feature amounts extracted from each of the multiple types of feature maps.

[0023] In a learning method according to the present invention, a learning unit inputs features of the first data and features of the second data into a score calculator in a predetermined order based on learning data that indicates a specific change from first data to second data, the specific change being the disappearance or appearance of an object represented by the data, and learning data that indicates the specific change from the second data to the first data, and trains the score calculator so as to calculate a score that indicates the specific change in the change from the first data to the second data, and inputs features of the first data and features of the second data into the score calculator in a reverse order to the predetermined order, and calculates a score that indicates the specific change in the change from the second data to the first data.

[0024] In the data change estimation method according to the present invention, a score calculation unit inputs features of first data and features of second data into a score calculator in a predetermined order to calculate a score indicating a specific change from the first data to the second data, which is the disappearance or appearance of an object represented by the data, and inputs the features of the first data and features of the second data into the score calculator in an order reverse to the predetermined order to calculate a score indicating the specific change in the change from the second data to the first data.

[0025] The learning program of the present invention is a learning program for causing a computer to execute the following steps: based on learning data indicating a specific change from first data to second data, which is the disappearance or appearance of an object represented by the data, and learning data indicating the specific change from the second data to the first data, inputting features of the first data and features of the second data into a score calculator in a predetermined order to calculate a score indicating the specific change in the change from the first data to the second data, and inputting features of the first data and features of the second data into the score calculator in the reverse order of the predetermined order to calculate a score indicating the specific change in the change from the second data to the first data.

[0026] The data change estimation program of the present invention is a data change estimation program that causes a computer to execute the following steps: inputting the features of the first data and the features of the second data into a score calculator in a predetermined order, calculating a score that indicates a specific change from the first data to the second data, which is the disappearance or appearance of an object represented by the data; and inputting the features of the first data and the features of the second data into the score calculator in the reverse order of the predetermined order, and calculating a score that indicates the specific change in the change from the second data to the first data. [Effects of the Invention]

[0027] As described above, the learning device, learning method, and learning program of the present invention have the effect of being able to estimate the disappearance and appearance of an object by training a score calculator that calculates both a score indicating the disappearance or appearance of an object from the first data to the second data and a score indicating the disappearance or appearance of an object from the second data to the first data.

[0028] Furthermore, the data change estimation device, data change estimation method, and data change estimation program according to the present invention have a simple configuration and can provide the effect of being able to estimate the disappearance and appearance of an object by calculating both a score indicating the disappearance or appearance of an object from the first data to the second data and a score indicating the disappearance or appearance of an object from the second data to the first data. [Brief explanation of the drawings]

[0029] [Figure 1] 1 is a block diagram showing a configuration of a change detector learning device according to an embodiment of the present invention; [Figure 2] FIG. 10 is a diagram showing an example of a change in belongings. [Figure 3] 2 is a block diagram showing a configuration of a change detector learning means of the change detector learning device according to the embodiment of the present invention; FIG. [Figure 4A] 2 is a block diagram showing a configuration of an object region dividing unit of the change detector learning device according to the embodiment of the present invention; FIG. [Figure 4B] 4 is a flowchart showing the operation of a target region dividing unit of the change detector learning device according to the embodiment of the present invention. [Figure 5] 2 is a block diagram showing the configuration of a change detector learning device and a score calculation unit of an image monitoring device according to an embodiment of the present invention; FIG. [Figure 6] FIG. 10 is a block diagram showing the configuration of a score calculator. [Figure 7] 1 is a block diagram showing a configuration of an image monitoring device according to an embodiment of the present invention; [Figure 8] 1 is a block diagram showing the configuration of a change detector of an image monitoring device according to an embodiment of the present invention; [Figure 9]4 is a flowchart showing the operation of the change detector learning device according to the embodiment of the present invention. [Figure 10] 4 is a flowchart showing the operation of the image monitoring device according to the embodiment of the present invention. [Figure 11] FIG. 10 is a block diagram showing the configuration of a score calculator in a modified example. DETAILED DESCRIPTION OF THE INVENTION

[0030] An embodiment of the present invention will be described in detail below with reference to the drawings. Hereinafter, as an embodiment of the present invention, a change detector learning device 2 that constructs a change detector that estimates changes in the presence or absence of belongings of a person appearing in an input image and an image monitoring device 1 that uses the change detector created by the change detector learning device 2 to detect changes in the belongings of a person appearing in an image captured of a specified space and notify the detection results will be described. While this embodiment describes an example in which the object of change detection is a person's belongings, the present invention may also detect changes in the presence or absence of furniture. The change detector learning device 2 is an example of a learning device, and the image monitoring device 1 is an example of a data change estimation device.

[0031] <Configuration of the change detector learning device> The change detector learning device according to the embodiment of the present invention is a device that learns a change detector using learning data prepared in advance.

[0032] The change detector training device can be configured with a computer including a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), RAM (Random Access Memory), and a ROM (Read Only Memory) that stores programs and various data for executing the learning processing routine described below. Figure 1 shows a block diagram of the change detector training device 2. As shown in Figure 1, the change detector training device 2 functionally includes a training data storage means 20, a change detector training means 21, and a trained model storage means 22.

[0033] (Learning data storage means 20) The learning data storage means 20 stores learning data indicating a specific change from a past image set to a current image set, which is a change in the possessions represented by the image, namely disappearance and appearance of the possessions.

[0034] Specifically, the learning data storage means 20 stores, as learning data, a large number of past image sets of changes, current image sets, and change information (presence or absence of appearance, presence or absence of disappearance). The past images of the learning data are examples of the first data and the third data, and the current images of the learning data are examples of the second data and the fourth data. These learning data are created as follows. First, a person is photographed from multiple viewpoints, and the person area is extracted from two different times t1 and t2 (t1 < t2) of the video for each viewpoint. The image set at time t1 is used as the past image set, and the image set at time t2 that photographs the same person as the past image set is used as the current image set. If the number of viewpoints from which the person is visible at time t1 is N1 and the number of viewpoints from which the person is visible at time t2 is N2, the past image set consists of N1 images and the current image set consists of N2 images. The extraction of the person area is performed by cutting out a rectangular area that includes the entire body area of the person using an existing person detector or the like. The cut-out images are resized to a predetermined image size. Then, the change information of the possessions of the same person between time t1 and time t2 is associated with and held for this past image set and the current image set.

[0035] When comparing two images, there are four states of change in the possessions (only disappearance, only appearance, simultaneous disappearance and appearance, no change). In contrast, the following can be used as annotations for expression.

[0036] · Whether something that existed in the past has disappeared in the present: Presence or absence of disappeared items · Whether there is something that did not exist in the past but newly appeared in the present: Presence or absence of newly appeared items

[0037] Examples of these are shown in Figure 2. As in Figure 2(a), if there are no possessions in the past image but there are possessions in the current image, the change in possessions will be (no lost possessions, some objects have appeared). Conversely, as in Figure 2(b), if there are possessions in the past image but not in the current image, this will be (lost possessions, no objects have appeared). Also, as in Figure 2(c), if there are different possessions in the past and current images because the person photographed has changed possessions, this will be (lost possessions, some objects have appeared), indicating a change in possessions.

[0038] The information on whether or not an object has disappeared and whether or not an object has appeared is expressed as two binary values ​​(1 if an object has disappeared or appeared, 0 if it has not), and is stored as change information from the past image set to the current image set.

[0039] The past image set and the current image set may be taken at the same location or at different locations as long as they are of the same person. The above procedure is performed on videos of various people and changes in various possessions, and a large number of past image sets, current image sets, and change information are created and stored in the learning data storage means 20 as learning data.

[0040] (Change detector learning means 21) The change detector learning means 21 reads multiple pieces of learning data, which are a past image set, a current image set, and change information, from the learning data storage means 20, learns a change detector, and stores the resulting change detector in the learned model storage means 22. Figure 3 shows a block diagram of the change detector learning means 21. The change detector is an example of a model. The change detector learning means 21 is made up of a change detector 30 and a learning unit 31.

[0041] (Change Detector 30) The change detector 30 is a neural network that takes a past image set and a current image set as input and outputs a probability value of changes in possessions between those image sets, and is composed of a feature map extraction unit 40, a candidate area estimation unit 41, an individual feature extraction unit 42, an individual score calculation unit 43, an individual score loss calculation unit 44, an overall feature extraction unit 45, an overall score calculation unit 46, an overall score loss calculation unit 47, and a target likelihood distillation loss calculation unit 48.

[0042] Here, the global feature extraction unit 45, the global score calculation unit 46, the global score loss calculation unit 47, and the target likelihood distillation loss calculation unit 48 are mainly used to stabilize the learning of the candidate region estimation unit 41. Therefore, these, the individual score loss calculation unit 44, and the learning unit 31 are used only during learning, and are not used during inference.

[0043] (Feature map extraction unit 40) The feature map extraction unit 40 extracts feature maps from input images. Specifically, the feature map extraction unit 40 is a convolutional neural network (CNN) configured with convolution processing, a ReLU function (activation function), pooling processing, and the like. In this embodiment, ResNet is used as the network structure of the feature map extraction unit 40. The feature map extraction unit 40 inputs each image from the input past image set and current image set into ResNet. The feature map that is the final output of ResNet is then combined with several feature maps that are intermediate results within ResNet to output Nf feature maps. These feature maps are in the form of a three-dimensional numerical sequence with dimensions of height, width, and channel. The feature maps calculated by ResNet each have different height and width, but the other feature maps are resized to fit the feature map with the largest height and width and the highest resolution using bilinear interpolation or the like.

[0044] (Candidate area estimation unit 41) The candidate area estimation unit 41 uses an object area estimation model to estimate an initial object likelihood map, which is a map of likelihood indicating the object area representing a possession, from the feature map. Specifically, the object area estimation model used by the candidate area estimation unit 41 is a CNN configured with convolution processing, a ReLU function (activation function), etc. The candidate area estimation unit 41 calculates and outputs an initial object likelihood map for each image in each set of past and current images using the feature map calculated by the feature map extraction unit 40. In this embodiment, the feature map that is large in length and width and has the highest resolution among multiple feature maps output from ResNet is input, and an object likelihood map is calculated using a network structure configured with convolution processing, a ReLU function, and another convolution processing. The calculated initial object likelihood map represents the object likelihood, and the more object-like it is, the larger the value.

[0045] (Individual feature extraction unit 42) The individual feature extraction unit 42 is composed of a target region extraction unit 49 and an individual feature vector extraction unit 52. The individual feature vector extraction unit 52 outputs, as an individual feature vector, a feature vector extracted from each of a plurality of target regions for each image in the past image set and the current image set.

[0046] (Target area extraction unit 49 of individual feature extraction unit 42) The object region extracting unit 49 extracts a plurality of object regions from the initial object likelihood map. Specifically, the object region extracting unit 49 is made up of an object region object map normalizing unit 50 and an object region dividing unit 51.

[0047] (Object region object map normalization unit 50 of object region extraction unit 49) The target region object map normalization unit 50 calculates a normalized initial target likelihood map. Specifically, the target region object map normalization unit 50 normalizes the initial target likelihood map calculated by the candidate region estimation unit 41 so that the value of each pixel falls within the range from 0 to 1, as shown in the following equation. w i,j =sigmoid(αx i,j +β)

[0048] where sigmoid() is the sigmoid function, x i,j is the value of the coordinate (i, j) of the initial object likelihood map, and α and β are the scale coefficient and bias coefficient. The scale coefficient α and bias coefficient β are determined by learning. However, they may also be set to fixed values ​​determined in advance. The normalized initial object likelihood map represents the likelihood of an object, and the more likely an object is, the closer the value is to 1. The object map normalization unit 50 for object region outputs this normalized initial object likelihood map as a sigmoid-normalized object likelihood map.

[0049] (Target area division unit 51 of target area extraction unit 49) Using the feature map calculated by feature map extraction unit 40, object region division unit 51 divides the sigmoid-normalized object likelihood map calculated by object region object map normalization unit 50 into individual object regions so as to sequentially cut out the map, and outputs the divided object regions as object likelihood partial maps. The object likelihood partial map is in the form of a weight map having the same width and height as the sigmoid-normalized object likelihood map and values ​​greater than or equal to 0.

[0050] Specifically, as shown in FIG. 4A, object region dividing section 51 is made up of a position likelihood map calculation section 70, an appearance likelihood map calculation section 71, an object likelihood map calculation section 72, and a modification section 73.

[0051] (Position likelihood map calculation unit 70 of object region division unit 51) First, the position likelihood map calculation unit 70 sets the sigmoid normalized object likelihood map calculated by the object region object map normalization unit 50 as an initial object likelihood map. The position likelihood map calculation unit 70 also calculates a position likelihood map from the initial object likelihood map.

[0052] (Appearance likelihood map calculation unit 71 of object region division unit 51) The appearance likelihood map calculation unit 71 calculates an appearance likelihood map from the feature map and the initial target likelihood map. Specifically, the appearance likelihood map calculation unit 71 calculates an average feature amount from the feature map using the initial target likelihood map as a weight, and calculates an appearance likelihood map from the feature map, which indicates the likelihood of the feature amount corresponding to the average feature amount.

[0053] (Object likelihood map calculation unit 72 of object region division unit 51) The object likelihood map calculation unit 72 calculates an object likelihood map from the initial object likelihood map, the position likelihood map, and the appearance likelihood map.

[0054] (Modification unit 73 of object region division unit 51) The modifying unit 73 records the target region represented by the target likelihood map as the extraction result of the target region, and modifies the initial target likelihood map so as to change the likelihood of the target region to a predetermined value.

[0055] By repeating the calculation by the position likelihood map calculation unit 70, the calculation by the appearance likelihood map calculation unit 71, the calculation by the object likelihood map calculation unit 72, and the modification by the modification unit 73, a plurality of object regions are extracted.

[0056] (Operation of the target area division unit 51) 4B is a schematic flow diagram of the operation of the object region dividing unit 51. For the sake of explanation, the sigmoid normalized object likelihood map w i,j is the initial target likelihood map w i,j|τ=0 It is called.

[0057] First, the position likelihood map calculation unit 70 calculates a position likelihood map from the initial target likelihood map (step SA0).

[0058] Specifically, first, the center of gravity μ and covariance matrix C of the pixel position are calculated using the initial target likelihood map as weights as shown in the following equations, and then the position likelihood map is calculated using these. JPEG0007736613000001.jpg4093

[0059] where p i,j is the position vector of coordinates i, j and p i,j =[i,j], M i,j loc are the values ​​of coordinates i and j of the position likelihood map, and K and q are parameters for controlling the distance. Here, if q=2 and K=1, the position likelihood map becomes a weight map with a gentle gradient like a Gaussian distribution, but in this embodiment, q=6 and K=π are used. This makes the gradient of the weight map steeper, making it easier to separate the object from the background.

[0060] Next, appearance likelihood map calculation unit 71 calculates an appearance likelihood map from the feature map and the initial target likelihood map (step SA1).

[0061] Specifically, as shown in the following equation, the initial target likelihood map is used as a weight to find the average feature vector h (h with an overbar), and the appearance likelihood map is found by calculating the similarity map with this. JPEG0007736613000002.jpg3884

[0062] where h i,j is the feature vector of the feature map coordinates i,j, M i,j feat is the value of coordinate i, j in the appearance likelihood map.

[0063] Next, the object likelihood map calculation unit 72 calculates a new object likelihood map using the initial object likelihood map, the position likelihood map, and the appearance likelihood map according to the following equation (step SA2). JPEG0007736613000003.jpg2198

[0064] Here, SpSoftmax ij ( ) is the softmax function in the spatial direction, JPEG0007736613000004.jpg1984 Also, T SpSoftmax is a preset temperature parameter.

[0065] Next, it is determined whether or not the iteration end condition is satisfied (step SA3). If the iteration end condition is not satisfied, the process returns to step SA0 and the newly calculated target likelihood map w i,j|τ=1 Using the above, a position likelihood map and an appearance likelihood map are calculated, and based on these, a target likelihood map w i,j|τ=2 is calculated (steps SA0 to SA2).

[0066] If the iteration termination condition is met, the final target likelihood map w i,j|τ=T is stored in an output list or the like as one of the target likelihood partial maps to be output (step SA3, step SA4).

[0067] The iteration termination condition is, for example, i,j|τ It is possible to use whether the amount of change in the parameter has become smaller than a predetermined threshold value, whether a predetermined number of iterations has been reached, or the like.

[0068] In the above process, regions are extracted using two types of likelihood maps: a position likelihood map and an appearance likelihood map. Therefore, the extracted regions are regions that are close in location and have similar features (similar appearances in the image). Furthermore, by iterative processing, regions with more similar locations and features are aggregated, improving the accuracy of the extraction of belongings regions.

[0069] Next, the modification unit 73 modifies the initial target likelihood map w i,j|τ=0 From the target likelihood map w i,j|τ=T The area of ​​the initial target likelihood map w is removed, and the value of that area of ​​the initial target likelihood map is updated to a predetermined value (step SA5). i,j|τ=0 From the target likelihood map w i,j|τ=T As a method for removing the region and changing it to a predetermined value, for example, a method for removing the region equal to or greater than a threshold and changing it to a predetermined value as shown in the following formula can be used.

[0070] JPEG0007736613000005.jpg16109 where γ is a predetermined scale factor.

[0071] Then, it is determined whether or not the iteration end condition is satisfied (step SA6). If the iteration end condition is not satisfied, the process returns to step SA0 and the newly calculated w i,j|τ=0 new the new initial target likelihood map w i,j|τ=0 By repeating the above process, a plurality of target likelihood partial maps are extracted (steps SA0 to SA5).

[0072] If the iteration end condition is satisfied, the object region dividing unit 51 outputs these as object likelihood partial maps (steps SA6 and SA7).

[0073] The iteration termination condition is, for example, the newly calculated w i,j|τ=0 new It is possible to use whether the number of pixels equal to or greater than the threshold has become smaller than a predetermined threshold, or whether a predetermined number of iterations has been reached.

[0074] The above process divides the object likelihood map into multiple regions in the form of an object likelihood partial map. This allows a region to be extracted for each item, even if multiple items are present in the image. This allows the individual feature vector extraction unit 52, described below, to extract the features of each item region without mixing them, improving the accuracy of change detection.

[0075] (Individual feature vector extraction unit 52 of individual feature extraction unit 42) The individual feature vector extraction unit 52 calculates a feature vector for each object likelihood partial map calculated by the object region division unit 51, and outputs it as an individual feature vector. The feature vector is calculated by taking a weighted average of the feature maps, with the object partial map used as the weight, and the feature vector is calculated for each of the multiple feature maps calculated by the feature map extraction unit 40. Specifically, it is calculated using the following equation: JPEG0007736613000006.jpg2260 where w r,i,j is the value of the coordinates i, j of the rth target submap, and h f,i,j is the feature vector at coordinates i,j of the fth feature map.

[0076] (Individual score calculation unit 43) The individual score calculation unit 43 calculates an individual score for each of the plurality of target regions based on the feature vector of the target region.

[0077] Specifically, for each combination of a target area of ​​the past image set and a target area of ​​the current image set, when the individual feature vectors of the target area of ​​the past image set calculated by the individual feature extraction unit 42 and the individual feature vectors of the target area of ​​the current image set are input, an integrated individual score is calculated by integrating the individual scores through the score calculation unit 53 and the score integration unit 54.

[0078] (Score calculation unit 53 of individual score calculation unit 43) For each combination of a target area in the past image set and a target area in the current image set, the score calculation unit 53 inputs the individual feature vector of the target area in the past image set and the individual feature vector of the target area in the current image set into the score calculator in a predetermined order to calculate an individual score indicating a change in the amount of lost belongings from the past image set to the current image set, which is referred to as a loss score indicating the loss of belongings. Specifically, the score calculation unit 53 calculates individual scores for various combinations of the individual feature vectors of the past image set and the current image set. For example, the score calculation unit 53 calculates individual scores for combinations of a target area in the past image set and a target area in the current image set, and combinations of an image in the past image set and an image in the current image set.

[0079] Figure 5 shows a block diagram of the score calculation unit 53. The score calculation unit 53 uses Nf score calculators 60 (the same number as the number of feature maps). First, the score calculator 60 calculates individual scores for all combinations of feature vectors of past images and feature vectors of the current image, and sets these as loss scores (scores that represent the possibility that something that was previously present may no longer exist, when comparing the past and current images). Specifically, this is expressed as follows: JPEG0007736613000007.jpg14123 where calcScore f ( , ) is the score calculator 60 for the fth feature map, h f,i1,r1 t1 is the feature vector calculated from the r1th subregion of the fth feature map of the i1th past image, and h f,i2,r2 t2 is the feature vector calculated from the r2-th subregion of the f-th feature map of the i2-th current image, s f,i1,i2,r1,r2 disappear is the extinction score calculated from them.

[0080] Next, the score calculation unit 53 uses the same score calculator 60 to input the feature vectors to the score calculator 60 in the reverse order (as in the calculation of the disappearance score described above) to calculate an individual score indicating a change in the disappearance of an item from the current image set to the past image set, i.e., a change in the appearance of an item from the past image set to the current image set, and uses this as an appearance score (a score that indicates the possibility that something that was not present in the past has now appeared, when comparing the past image and the current image). For example, if the feature vector of the past image and the feature vector of the current image are input to the score calculator 60 in this order when calculating the disappearance score, then the feature vector of the current image and the feature vector of the past image are input to the score calculator 60 in this order when calculating the appearance score. Specifically, this is expressed as follows: JPEG0007736613000008.jpg10109

[0081] The score calculator 60 is configured by a neural network. Figure 6 shows a block diagram of the score calculator 60 of this embodiment. The first feature vector and the second feature vector are input to the score calculator 60, which then applies a common common conversion unit 61 to each of them, and then performs asymmetric calculations on them, and finally calculates an individual score (disappearance score or appearance score) through a scoring process. Specifically, these are expressed as follows:

[0082] Common conversion unit 61: JPEG0007736613000009.jpg1181 Asymmetric calculation unit 62: JPEG0007736613000010.jpg2179 Scoring processing unit 63: s = sigmoid(FC(v))

[0083] Here, h1 is the first feature vector, h2 is the second feature vector, L( ) is the common conversion unit 61 configured by fully connected processing and ReLU function, etc., and h1 ’ and h2 ’ is a D-dimensional vector, FC() is a fully connected process, and sigmoid() is a sigmoid function. Note that in this embodiment, since v is a scalar value, FC() is essentially a linear transformation αv+β with a scale coefficient α and a bias coefficient β.

[0084] The asymmetric calculation unit 62 subtracts the second feature vector from the first feature vector, performs threshold processing using the ReLU function, and outputs the average value of those elements.

[0085] Here, the advantages of the score calculator 60 of this embodiment will be explained. A conventional method for comparing images is the Siamese Network, which is a neural network that takes two images as input and calculates the distance between the images.

[0086] On the other hand, in the task of change detection, the value to be output differs depending on the order of images input to the change detector 30. For example, in FIG. 2(a), only the current image contains an object, so it is desirable for the change detector 30 to output (no missing object, object appears). On the other hand, if the past image and current image in FIG. 2(a) are swapped, it becomes FIG. 2(b), in which case it is desirable for the change detector 30 to output (object disappears, object does not appear), which is different from FIG. 2(a). To achieve this asymmetry, the score calculator 60 described above needs to be able to output different values ​​for calcScore(h1,h2) and calcScore(h2,h1).

[0087] However, in the Siamese Network mentioned above, the output distance does not change even if the order of the images is changed, so this kind of asymmetry cannot be realized.

[0088] In the score calculator 60 of this embodiment, by using subtraction, an asymmetric operation, different values ​​can be output for calcScore(h1, h2) and calcScore(h2, h1), thereby realizing asymmetry and improving the accuracy of change detection.

[0089] (Score integration unit 54 of individual score calculation unit 43) The score integration unit 54 calculates the maximum, minimum, or average value of the individual scores indicating a change, which is the disappearance of an item from the past image set to the current image set, as an integrated individual score, and calculates the maximum, minimum, or average value of the individual scores indicating a change, which is the disappearance of an item from the current image set to the past image set, i.e., the appearance of an item from the past image set to the current image set, as an integrated individual score.

[0090] Specifically, the score integration unit 54 integrates the disappearance scores and appearance scores calculated by the score calculation unit 53 in various combinations, and outputs the integrated individual score. The score integration unit 54 integrates the scores for each target region and then for each image. For feature maps, the change detector learning device 2 does not integrate scores, and the individual score loss calculation unit 44 calculates the loss for each.

[0091] We will now explain how to integrate the loss scores for each target area. First, for each target area in the past image, select the score of the target area with the least change (smallest score) from the target areas in the current image. This is equivalent to matching the target area that represents the most similar possession. Next, select the score of the target area with the most change (largest score) from the target areas in the past image. This selects the score of the target area that is determined to have changed the most. This can be written as follows: JPEG0007736613000011.jpg1281

[0092] The integration of the disappearance scores for each image is performed as in the case of the integration of the scores for each target region, as shown in the following equation. JPEG0007736613000012.jpg1385

[0093] The integration of the appearance scores is performed in the same manner as the integration of the disappearance scores, as shown in the following equation. Score integration for target domain axes: JPEG0007736613000013.jpg1583 Image axis score integration: JPEG0007736613000014.jpg1676

[0094] The score integration unit 54 integrates the disappearance scores s for each feature map obtained by the above procedure. f disappear , occurrence score s f appear is output as the integrated individual score.

[0095] (Individual score loss calculation unit 44) The individual score loss calculation unit 44 calculates the loss using the integrated individual score calculated by the individual score calculation unit 43 and the change information (presence / absence of a disappeared object, presence / absence of an appeared object) that is the learning data. Specifically, it calculates the error for each disappearance score and appearance score for each feature map as shown in the following equation, and outputs the sum of these as the individual score loss. JPEG0007736613000015.jpg3479

[0096] where s f disappear , s f appear is the integrated individual score calculated by the individual score calculation unit 43, t disappear is a binary value indicating whether or not an object is missing, and t appear is a binary value indicating the presence or absence of an occurrence, and CE( ) is the cross-entropy error function.

[0097] The individual score loss calculation unit 44 is L partial is output as the individual score loss.

[0098] (Global feature extraction unit 45) The global feature extraction unit 45 is composed of a global region object map normalization unit 55 and a global feature vector extraction unit 56. The global feature vector extraction unit 56 outputs, as a global feature vector, a feature vector extracted using the object likelihood map for each image in the past image set and the current image set.

[0099] (Global region object map normalization unit 55 of global feature extraction unit 45) The whole region object map normalization unit 55 normalizes the initial object likelihood map calculated by the candidate region estimation unit 41 by performing SpSoftmax ij Normalize it with ( ) and output it as a softmax normalized target map.

[0100] (Global feature vector extraction unit 56 of global feature extraction unit 45) Global feature vector extraction unit 56 calculates a feature vector by taking a weighted average of the feature maps, using the softmax-normalized object map calculated by global region object map normalization unit 55 as the weight, and outputs the result as a global feature vector. As with individual feature vector extraction unit 52, a feature vector is calculated for each of the multiple feature maps calculated by feature map extraction unit 40. Specifically, this is calculated using the following equation: JPEG0007736613000016.jpg2159 where wi,j is the value of coordinate i,j of the softmax normalized target map, and h f,i,j is the feature vector at coordinates i,j of the fth feature map.

[0101] (Overall score calculation part 46) When the global feature vectors of the past image set and the current image set calculated by the global feature extraction unit 45 are input, the global score calculation unit 46 calculates a change detection score through the score calculation unit 57 and the score integration unit 58.

[0102] (Score calculation unit 57 of overall score calculation unit 46) The score calculation unit 57 of the overall score calculation unit 46 calculates an overall score using various combinations of the overall feature vectors of the past image set and the current image set, similar to the score calculation unit 53 of the individual score calculation unit 43. The score calculation unit 57 uses Nf score calculators 60 (the same number as the number of feature maps) to calculate disappearance scores and appearance scores as follows: JPEG0007736613000017.jpg2692 where calcScore f ( , ) is the score calculator 60 for the fth feature map, h i1 t1 is the global feature vector calculated from the fth feature map of the i1th past image, h i2 t2 is the global feature vector calculated from the f-th feature map of the i2-th current image, s f,i1,i2 disappear , s f,i2,i1 appear are the disappearance score and appearance score calculated from them. The score calculator 60 used in the score calculation unit 57 of the overall score calculation unit 46 is configured by a neural network and has the same structure as the score calculation unit 57 of the overall score calculation unit 46.

[0103] (Score integration unit 58 of overall score calculation unit 46) The score integration unit 58 of the overall score calculation unit 46 calculates the maximum, minimum, or average value of the overall score indicating a change that is the disappearance of a possession from the past image set to the current image set as the integrated overall score, and calculates the maximum, minimum, or average value of the overall score indicating a change that is the disappearance of a possession from the current image set to the past image set, i.e., the appearance of a possession from the past image set to the current image set, as the integrated overall score.

[0104] Specifically, the score integration unit 58 integrates the disappearance scores and appearance scores calculated for various combinations by the score calculation unit 57, and outputs the integrated overall score. This score integration unit 58 integrates scores only for each image. For feature maps, the change detector learning device 2 does not integrate scores, and the overall score loss calculation unit 47 calculates losses for each.

[0105] The integration method is the same as that of the score integration unit 54 of the individual score calculation unit 43, and specifically, is as follows: Score integration of disappearance scores: JPEG0007736613000018.jpg1385 Score integration of occurrence scores: JPEG0007736613000019.jpg1382

[0106] The score integration unit 58 integrates the loss scores s for each feature map obtained by the above procedure. f disappear , occurrence score s f appear is output as the integrated overall score.

[0107] (Total score loss calculation unit 47) Similar to the individual score loss calculation unit 44, the overall score loss calculation unit 47 calculates the overall score loss based on the integrated overall score calculated by the overall score calculation unit 46 and the change information (with / without loss, with / without appearance) that is the learning data. Specifically, it calculates the error for each loss score and appearance score for each feature map as shown in the following equation, and outputs the sum of these as the overall score loss. JPEG0007736613000020.jpg1582 JPEG0007736613000021.jpg2678

[0108] where s f disappear , s f appear is the integrated overall score calculated by the overall score calculation unit 46, t disappear is a binary value indicating whether or not an object is missing, and t appear is a binary value indicating the presence or absence of an appearance, and CE( ) is the cross entropy error function. The overall score loss calculation unit 47 calculates L whole is output as the overall score loss.

[0109] Here, the advantages of using the global feature extraction unit 45 and the like during learning will be described. First, in this embodiment, parameters for the candidate area estimation unit 41 and the like are initialized with random numbers, so in the early stages of learning, the target likelihood map calculated by the candidate area estimation unit 41 is often extracted from all areas and does not necessarily appear only in the possession area. Furthermore, because no learning data related to candidate areas is provided and the candidate area estimation unit 41 is trained automatically through learning of the change detector, the learning of the candidate area estimation unit 41 is prone to instability.

[0110] Furthermore, since the processing of the individual feature vector extraction unit 52 uses iterative processing and the neural network becomes deep, there is a problem that the loss is difficult to propagate to the candidate region estimation unit 41, and learning becomes even less stable.

[0111] Therefore, in this embodiment, the change detection problem is solved through a simple process of weighted averaging using a softmax normalized target map in the global feature vector extraction unit 56, and the loss is propagated to the candidate area estimation unit 41 and the like, thereby promoting the learning of the candidate area estimation unit 41. This makes it possible to stabilize the learning of the candidate area estimation unit 41 and the like.

[0112] (Target likelihood distillation loss calculation unit 48) Object likelihood distillation loss calculation unit 48 measures the similarity between the sigmoid-normalized object likelihood map calculated by object region object map normalization unit 50 and the softmax-normalized object map calculated by entire region object map normalization unit 55, and outputs the object likelihood distillation loss. This prevents the scale coefficients and bias coefficients used in object region object map normalization unit 50 from becoming excessively large or small during learning, resulting in the sigmoid-normalized object likelihood map becoming an inappropriate map for the object likelihood.

[0113] Specifically, the target likelihood distillation loss is calculated as follows: JPEG0007736613000022.jpg1982 where w i,j sig is the value of the coordinates i, j of the sigmoid normalized target likelihood map, and w i,j sm is the value of coordinate i,j of the softmax normalized target map.

[0114] In this embodiment, the error backpropagation method is used for learning, but the error information is not propagated to the softmax normalized object map, but only to the sigmoid normalized object likelihood map, which makes it possible to more stably learn the scale and bias coefficients used in the object region object map normalizer 50.

[0115] (Learning Section 31) The learning unit 31 learns the change detector 30, which includes a feature map extraction unit 40, a candidate region estimation unit 41, an overall feature extraction unit 45, an overall score calculation unit 46, an individual feature extraction unit 42, and an individual score calculation unit 43, based on the loss calculated from the overall score and the learning data, and the loss calculated from the individual scores and the learning data.

[0116] In addition, the learning unit 31 further uses a loss calculated from the normalized target likelihood map obtained by the overall feature extraction unit 45 and the normalized target likelihood map obtained by the individual feature extraction unit 42 to train the overall feature extraction unit 45 and the individual feature extraction unit 42 of the change detector 30.

[0117] More specifically, the learning unit 31 calculates the individual score loss L calculated by the individual score loss calculation unit 44. partial , the overall score loss L calculated by the overall score loss calculation unit 47 whole , the target likelihood distillation loss L calculated by the target likelihood distillation loss calculation unit 48 sim The sum of the three losses is L=L partial +L whole +L sim The parameters of the change detector 30 are updated so as to minimize

[0118] A stochastic gradient method is used to minimize the loss. In learning using the stochastic gradient method, parameters of the feature map extraction unit 40, candidate region estimation unit 41, target region object map normalization unit 50, score calculation unit 53 of the individual score calculation unit 43, and score calculation unit 57 of the overall score calculation unit 46 are first initialized with random values ​​or the like. Then, the gradient of the parameters is calculated using the backpropagation method based on the loss, and the parameters are updated based on the gradient. The change detector learning device 2 learns the change detector 30 by repeating the above processes of reading training data, calculating the loss, and updating the parameters. The iteration termination condition can be, for example, whether the amount of change in the loss has become smaller than a predetermined threshold, or whether a predetermined number of iterations has been reached.

[0119] In training the change detector 30 described above, the learning unit 31 trains the score calculator 60 so that it inputs the individual feature vectors of the past image set and the individual feature vectors of the current image set into the score calculator 60 in a predetermined order to calculate an individual score indicating a change that is the disappearance of an item from the past image set to the current image set, and also inputs the individual feature vectors of the past image set and the individual feature vectors of the current image set into the score calculator 60 in the reverse order to calculate an individual score indicating a change that is the appearance of an item from the past image set to the current image set.

[0120] Furthermore, in training the change detector 30 described above, the learning unit 31 trains the score calculator 60 so that it inputs the overall feature vector of the past image set and the overall feature vector of the current image set into the score calculator 60 in a predetermined order to calculate an overall score indicating a change that is the disappearance of an item from the past image set to the current image set, and also inputs the overall feature vector of the past image set and the overall feature vector of the current image set into the score calculator 60 in the reverse order to calculate an overall score indicating a change that is the appearance of an item from the past image set to the current image set.

[0121] The change detector learning device 2 stores the finally determined parameters of the change detector 30 together with the network structure and the like as a learned model in the learned model storage means 22.

[0122] [Configuration of image monitoring device] An image monitoring device is a device that detects changes in belongings of people captured in an image of a predetermined space and notifies the user of the detection results.

[0123] The image monitoring device can be configured with a computer including a CPU, a GPU, a RAM, and a ROM that stores programs and various data for executing the image monitoring processing routine described below. Fig. 7 shows a block diagram of the image monitoring device 1. As shown in Fig. 7, the image monitoring device 1 functionally comprises an image acquisition means 10, an image storage means 11, a detection means 12, a display means 13, and a trained model storage means 14.

[0124] Image acquisition means 10 acquires color images from a plurality of surveillance cameras capturing images of a predetermined space from different viewpoints, and outputs the images to image holding means 11.

[0125] The image storage means 11 uses an existing person detector or the like to cut out rectangular person areas from input images and stores them sequentially along with information on the time of shooting. When storing these images, it uses existing face recognition technology or person identification technology to determine whether they are the same person as images already stored. If they are the same person, the images are linked and stored along with the person ID. If there are more than a certain number of images of the same person stored at two different times that are more than a certain distance apart, the image storage means 11 separates the images by time, and outputs the older images as a past image set and the newer images as a current image set, along with the person ID. Images that are more than a certain distance older than the current time are discarded.

[0126] The trained model storage means 14 stores trained change detectors obtained in advance by the change detector training device 2.

[0127] The detection means 12 first reads the trained change detector from the trained model storage means 14. Next, when the past image set, current image set, and person ID are input from the image storage means 11, the change detector is used to estimate the probability value of change. Figure 8 shows a block diagram of the change detector 130 in the image monitoring device 1. Note that parts that have the same configuration as the change detector 30 in the change detector training device 2 are assigned the same reference numerals and detailed explanations will be omitted. The past image is an example of the first data, and the current image is an example of the second data.

[0128] First, the past image set and the current image set input to the detection means 12 are input to the feature map extraction unit 40, which outputs a feature map of each image.

[0129] The candidate region estimation unit 41 uses the feature map to output an initial target likelihood map. When the feature map and target likelihood map of each image are input, the individual feature extraction unit 42 calculates and outputs the individual feature vector of each image via the target region extraction unit 49 and the individual feature vector extraction unit 52.

[0130] The object region extracting unit 49 extracts a plurality of object regions from each image through the object region object map normalizing unit 50 and the object region dividing unit 51, and outputs the object regions as object likelihood partial maps.

[0131] The individual score calculation unit 43 calculates the individual integrated change score through the score calculation unit 53 and the score integration unit 54 .

[0132] The score calculation unit 53 inputs the individual feature vectors of the past image set and the individual feature vectors of the current image set in a predetermined order to the score calculator 60, and calculates an individual score indicating the change, that is, the disappearance of an item, from the past image set to the current image set, which is the loss score.

[0133] The score calculation unit 53 uses the score calculator 60 to input the feature vectors into the score calculator 60 in reverse order to calculate individual scores indicating the change, that is, the appearance of an item, from the past image set to the current image set, and uses these as the appearance scores.

[0134] The score calculation unit 53 may further use the disappearance score and the appearance score to calculate a score indicating no change and a score indicating a change in belongings. For example, the score calculation unit 53 may calculate a score indicating no change that is larger as both the disappearance score and the appearance score are smaller, and calculate a score indicating a change in belongings that is larger as both the disappearance score and the appearance score are larger.

[0135] Furthermore, the score integration unit 54 in the image monitoring device 1 integrates scores for each feature amount in addition to the score integration process of the score integration unit 54 in the change detector learning device 2. For example, it integrates the individual scores for each combination of a target area in the past image set and a target area in the current image set, integrates the individual scores for each combination of an image in the past image set and an image in the current image set, and integrates the individual scores for each feature map. Therefore, the score integration process is expressed as follows:

[0136] Score integration on the target area axis: JPEG0007736613000023.jpg1282 JPEG0007736613000024.jpg1379 Image axis score integration: JPEG0007736613000025.jpg2479 Feature axis score integration: JPEG0007736613000026.jpg4282

[0137] Disappearance score s disappear , occurrence score s appear are the probability values ​​of the likelihood of disappearance and the likelihood of appearance, respectively, and the individual score calculation unit 43 outputs these as individual scores.

[0138] The detection means 12 outputs the calculated individual scores as detection results together with the past image set, the current image set, the person ID, and the like.

[0139] The display means 13 displays the change detection results, the past image set, the current image set, and the person ID on a display.

[0140] [Operation of change detector learning device 2] FIG. 9 is a flow diagram of a learning processing routine related to the operation of the change detector learning device 2.

[0141] When the learning operation is started, the change detector learning device 2 sets the change detector 30 to a predetermined network structure, and initializes the parameters of the network with random numerical values ​​or the like (step S10).

[0142] The change detector learning device 2 reads the past image set, the current image set, and change information as learning data from the learning data storage means 20 (step S11).

[0143] The change detector learning device 2 inputs the past image set and the current image set into the change detector 30 in the change detector learning means 21, calculates an individual score, which is a probability value, and calculates a loss using this and the change information, which is learning data (step S12).

[0144] The change detector learning device 2 calculates the gradient of each parameter by the backpropagation method based on the calculated loss in the change detector learning means 21, and updates each parameter by the stochastic gradient method using the gradient (step S13).

[0145] The change detector learning device 2 determines whether the iteration termination condition is satisfied (step S14). If the iteration termination condition is satisfied, the change detector 30 is stored in the trained model storage means 22 (step S15) and the process ends. If the iteration termination condition is not satisfied, the operations from step S11 to step S14 are repeated until the iteration termination condition is satisfied. The iteration termination condition may be, for example, whether the amount of error variation has become smaller than a predetermined threshold value or whether a predetermined number of iterations has been reached.

[0146] [Image monitoring device 1 operation] FIG. 10 is a flow diagram of an image monitoring processing routine related to the operation of the image monitoring device 1.

[0147] When the operation starts, the image monitoring device 1 reads the trained model of the change detector 130 from the trained model storage means 14 (step S20).

[0148] Next, the image monitoring device 1 acquires a color image by the image acquisition means 10 (step S21). The acquired photographed image is sent to the image holding means 11.

[0149] The image monitoring device 1 uses the image storage means 11 to cut out the person area from the received captured image, determine whether it is the same person as the image stored, and if it is the same person, link the images together and store them together with the person ID (step S22).

[0150] Next, it is determined whether there are any images to be output (step S23). If a certain number or more of images of the same person are stored at two different times that are at least a certain distance apart, it is determined that there are images to be output, and the image storage means 11 separates those images by time, and outputs the older images as a past image set and the newer images as a current image set together with the person ID. If there are no images to be output, the process returns to image acquisition by the image acquisition means 10 (step S21).

[0151] The detection means 12 inputs the received past image set and current image set into the change detector 130, and as a result of the detection process, calculates individual scores, namely, a disappearance score and an appearance score, which are output to the display means 13 together with the past image set, the current image set, and the person ID (step S24).

[0152] The image monitoring device 1 displays the acquired image and the detection results on the display using the display means 13 (step S25). Then, a monitor visually checks the displayed detection results to determine the situation, and if there is an abnormality, dispatches a response person as necessary.

[0153] After the recognition result is displayed, the process returns to image acquisition by the image acquisition means 10 (step S21), and continues until the image monitoring device 1 is stopped.

[0154] As described above, the change detector learning device and image monitoring device according to this embodiment calculate a target likelihood map from the initial target likelihood map, record the target regions represented by the target likelihood map as the target region extraction results, and repeatedly change the initial target likelihood map to change the likelihood of the target regions to a predetermined value. This makes it possible to accurately extract multiple target regions, even if the number of target regions is arbitrary.

[0155] Furthermore, the learning device according to this embodiment trains a change detector based on the loss calculated from the overall score and training data, and the loss calculated from the individual scores and training data. In this way, even if the number of target regions is arbitrary, it is possible to train a change detector that can accurately extract multiple target regions.

[0156] Furthermore, according to the change detector learning device of this embodiment, the score calculator is trained to input the feature vectors of the past image set and the feature vectors of the current image set into a predetermined order to calculate a score indicating a change in which an item has disappeared from the past image set to the current image set, and to input the feature vectors of the past image set and the feature vectors of the current image set into the score calculator in the reverse order of the predetermined order to calculate a score indicating a change in which an item has disappeared from the current image set to the past image set, i.e., a change in which an item has appeared from the past image set to the current image set. This makes it possible to train the score calculator to calculate both a score indicating the disappearance of an item from the past image set to the current image set and a score indicating the appearance of an item.

[0157] Furthermore, according to the data change estimation device of this embodiment, the feature vectors of the past image set and the feature vectors of the current image set are input to a score calculator in a predetermined order to calculate a score indicating a change in the disappearance of an item from the past image set to the current image set, and the feature vectors of the past image set and the feature vectors of the current image set are input to the score calculator in the reverse order to the predetermined order to calculate a score indicating a change in the disappearance of an item from the current image set to the past image set, i.e., a change in the appearance of an item from the past image set to the current image set. This makes it possible to estimate both the disappearance and appearance of an item from the past image set to the current image set with a simple configuration.

[0158] <Modification> The present invention is not limited to the device configuration and operation of the above-described embodiment, and various modifications and applications are possible within the scope of the gist of the present invention.

[0159] (Variation 1) In the above embodiment, an example has been described in which images captured by a camera are used as data to detect changes in belongings as a target, but the present invention can also be applied to detecting changes in a scene. That is, it is possible to detect whether an object in a space has disappeared or a new object has appeared from past and present images of the space. The present invention can also be applied to change detection using not only visible images but also thermal images, distance images, etc. Change detection can also be performed using data other than images. For example, the present invention can be applied to comparing sounds obtained from a microphone as data to detect the appearance or disappearance of abnormal sounds as a target.

[0160] (Variation 2) In the above embodiment, the feature map extraction unit is configured with a neural network, but it may also be configured to extract features such as HOG (Histograms of Oriented Gradients) features, LBP (Local Binary Patterns) features, color histograms, etc. Also, HOG features and CNN features may be used in combination.

[0161] (Variation 3) In the above embodiment, a past image set and a current image set taken from multiple viewpoints are used, but a single viewpoint may also be used.

[0162] Alternatively, instead of images taken from multiple viewpoints, an image set may be a series of images of the same person obtained by tracking a person in images using existing person tracking technology, etc. Furthermore, an image set may be constructed using both multiple viewpoints and person tracking technology.

[0163] (Variation 4) In the above embodiment, the candidate area estimation unit is trained by propagating the change detection loss more directly to the candidate area estimation unit through the overall score loss rather than the individual score loss. However, this is not limited to this. For example, a mask image of the belongings area in the image may be prepared in advance, and the error between the mask image and the initial target likelihood map calculated by the candidate area estimation unit may be added to the loss during training to train the change detector. In this case, the overall feature extraction unit, overall score calculation unit, overall score loss calculation unit, and target likelihood distillation loss calculation unit may not be used.

[0164] (Variation 5) In the above embodiment, the object region dividing unit calculates a plurality of object likelihood partial maps by dividing the object into a plurality of object regions through iterative processing, but the present invention is not limited to this.

[0165] One method is to group similar pixels. Specifically, first, the sigmoid-normalized target likelihood map is searched for local maxima whose centers are largest in local regions, such as 3x3 pixels, and these are stored in a list or similar. Next, pixels that are similar to these maxima in terms of position and feature value are grouped together. If two local maxima are similar, the one with the smaller value in the sigmoid-normalized target likelihood map is deleted from the list of local maxima. Then, a mask image is created for each local maxima, in which pixels in the same group as the local maxima are assigned a value of 1 and other pixels are assigned a value of 0, and the target region segmentation unit outputs these as the target likelihood partial map.

[0166] (Variation 6) In the above embodiment, a score calculator using asymmetric calculation is used as the score calculator, but this is not limiting. For example, a score calculator using a different feature transformation can be used. Figure 11 shows a block diagram of a score calculator 60A using a different feature transformation.

[0167] The first feature vector and the second feature vector are input to score calculator 60A, and first conversion unit 61A and second conversion unit 61B perform different conversion processes on them, and then symmetrical calculation unit 62A performs symmetrical calculations on them, and finally scoring processing unit 63 calculates a change score (disappearance score or appearance score). Specifically, these are expressed by the following equations.

[0168] First conversion unit 61A, second conversion unit 61B: JPEG0007736613000027.jpg1175 Symmetrical arithmetic unit 62A: JPEG0007736613000028.jpg1038 Scoring processing unit 63: s = sigmoid(FC(v))

[0169] Here, h1 is the first feature vector, h2 is the second feature vector, L1( ) and L2( ) are the first and second transformations, which are composed of fully connected processes and ReLU functions. ’ and h2 ’ is a D-dimensional vector, FC() is a fully connected process, and sigmoid() is a sigmoid function. In this embodiment, since v is a scalar value, FC() is essentially a linear transformation αv+β with a scale coefficient α and a bias coefficient β.

[0170] In this score calculator 60A, asymmetry is achieved by performing different transformations on the first feature vector and the second feature vector.

[0171] (Variation 7) In the above embodiment, only one score calculator is used, but this is not limiting. For example, there is a method in which multiple score calculators are prepared, each calculates a change score, and the average of these is used as a new change score. In this case, by initializing the parameters of each score calculator with different values ​​during learning, each score calculator can be made diverse, thereby improving the accuracy of change detection.

[0172] (Variation 8) In the above embodiment, the score integration unit of the individual score calculation unit of the image monitoring device 1 integrates the scores for each image using the same max-min calculation as the change detector learning device 2, but this is not limited to this. For example, the change detector learning device 2 may learn using the max-min calculation, while the image monitoring device 1 may use an integration method different from that used by the change detector learning device 2, such as the mean-min calculation shown in the following equation: JPEG0007736613000029.jpg2685

[0173] In max-min operations, if the past or current image sets contain outlier images that are significantly different from the images expected as input, such as when a person's region is not properly cropped, the change score may fluctuate significantly due to the influence of those images.Mean-min operations have the advantage of being less susceptible to the influence of outlier images, making the change score more robust.

[0174] (Variation 9) In the above embodiment, the change detection results are displayed on the display means 13 of the image monitoring device 1, but it is also possible to display which locations have changed in more detail. In the change detector, the candidate area estimation unit calculates candidate areas that are likely to have changed, and calculates a score based on those areas. The final score is then calculated from pairs of areas, one in multiple images included in the past image set and one in multiple images included in the current image set. Therefore, by storing the scores of these pairs and tracing which areas contributed to the final score, it is possible to display which areas of which images are different using rectangles or the like on the images.

[0175] (Variation 10) In the above embodiment, the object region dividing unit 51 is described as including the position likelihood map calculating unit 70, the appearance likelihood map calculating unit 71, the object likelihood map calculating unit 72, and the modifying unit 73. However, the present invention is not limited to this. Either the position likelihood map calculating unit 70 or the appearance likelihood map calculating unit may be omitted.

[0176] (Variation 11) In the above embodiment, the image monitoring device 1 and the change detector learning device 2 are separate devices, but this is not limiting. The image monitoring device 1 may also be configured to further include each of the means of the change detector learning device 2.

[0177] As described above, those skilled in the art can make various modifications to the embodiments within the scope of the present invention. [Explanation of symbols]

[0178] 1. Image monitoring equipment 2. Change detector trainer 12 Detection Methods 14. Learned model storage means 20 Learning data storage means 21 Change detector training means 22 Trained model storage means 30, 130 Change detector 31 Learning Department 40 Feature map extraction unit 41 Candidate region estimation unit 42 Individual feature extraction unit 43 Individual score calculation section 44 Individual score loss calculation part 45 Global feature extraction unit 46 Overall score calculation section 47 Overall score loss calculation section 48 Target likelihood distillation loss calculation part 49 Target region extraction unit 50 Object map normalization unit for object region 51 Object region division unit 52 Individual feature vector extraction unit 53, 57 Score calculation section 54, 58 Score Integration Section 55 Global region object map normalization unit 56 Global feature vector extraction unit 60, 60A score calculator 70 Position likelihood map calculation unit 71 Appearance likelihood map calculation unit 72 Object likelihood map calculation unit 73 Changes

Claims

1. 1. A learning device including a learning unit that causes a score calculator to learn based on learning data indicating a change from first data to second data, the change being the disappearance of an object represented by the data, and learning data indicating a change from the second data to the first data, the change being the appearance of an object represented by the data, The learning unit the score calculator calculates a score indicating the change that is the disappearance when receiving the feature amount of the first data and the feature amount of the second data as input in a predetermined order; and causing the score calculator to learn so as to calculate a score indicating the change that is the occurrence when the feature amount of the first data and the feature amount of the second data are received as input in a reverse order to the predetermined order; Learning device.

2. 1. A learning device including a learning unit that trains a score calculator based on learning data indicating a change from first data to second data, the change being an appearance of an object represented by the data, and learning data indicating a change from the second data to the first data, the change being a disappearance of the object represented by the data, The learning unit the score calculator calculates a score indicating the change that is the occurrence when the feature amount of the first data and the feature amount of the second data are received as input in a predetermined order; and causing the score calculator to learn so as to calculate a score indicating the change that is the disappearance when the feature amount of the first data and the feature amount of the second data are received as input in a reverse order to the predetermined order; Learning device.

3. A data change estimation device including a score calculation unit that calculates a score using a score calculator trained based on learning data indicating a change from third data to fourth data, the change being a disappearance of an object represented by the data, and learning data indicating a change from the fourth data to the third data, the change being an appearance of an object represented by the data, The score calculator calculating a score indicating the change that is the disappearance when the feature amount of the third data and the feature amount of the fourth data are received as input in a predetermined order; and the feature amount of the third data and the feature amount of the fourth data are received as input in a reverse order to the predetermined order, and the feature amount of the fourth data is learned to calculate a score indicating the change that is the occurrence; The score calculation unit inputting the feature amount of the first data and the feature amount of the second data into the score calculator in a predetermined order to calculate a score indicating the change that is the disappearance; and inputting the feature amount of the first data and the feature amount of the second data into the score calculator in a reverse order to the predetermined order, and calculating a score indicating the change that is the occurrence; Data change estimation device.

4. A data change estimation device including a score calculation unit that calculates a score using a score calculator trained based on learning data indicating a change from third data to fourth data, the change being an appearance of an object represented by the data, and learning data indicating a change from the fourth data to the third data, the change being a disappearance of the object represented by the data, The score calculator calculating a score indicating the change that is the occurrence when the feature amount of the third data and the feature amount of the fourth data are received as input in a predetermined order; and the feature amount of the third data and the feature amount of the fourth data are received as input in a reverse order to the predetermined order, and the feature amount of the fourth data is learned to calculate a score indicating the change that is the disappearance; The score calculation unit inputting the feature amount of the first data and the feature amount of the second data into the score calculator in a predetermined order, and calculating a score indicating the change that is the occurrence; and inputting the feature amount of the first data and the feature amount of the second data into the score calculator in a reverse order to the predetermined order, and calculating a score indicating the change that is the disappearance; Data change estimation device.

5. inputting the feature amount of the first data and the feature amount of the second data into a score calculator in a predetermined order, and calculating a score indicating a change from the first data to the second data, the change being the disappearance of an object represented by the data; and a score calculation unit that inputs the feature amounts of the first data and the feature amounts of the second data into the score calculator in a reverse order to the predetermined order, and calculates a score that indicates a change from the second data to the first data, the change being an appearance of an object represented by the data; A data change estimation device comprising:

6. inputting the feature amount of the first data and the feature amount of the second data into a score calculator in a predetermined order, and calculating a score indicating a change from the first data to the second data, the change being the appearance of an object represented by the data; and a score calculation unit that inputs the feature amounts of the first data and the feature amounts of the second data to the score calculator in a reverse order to the predetermined order, and calculates a score that indicates a change from the second data to the first data, which is a disappearance of an object represented by the data; A data change estimation device comprising:

7. 7. The data change estimation device according to claim 3, wherein the score calculator calculates the score indicating the change that is a disappearance or the score indicating the change that is an appearance using an asymmetric operation.

8. The data change estimation device according to any one of claims 5 to 7, wherein the score calculation unit further uses the score indicating the change that is the disappearance and the score indicating the change that is the appearance to calculate a score indicating no change and a score indicating a replacement of the target.

9. a feature map extraction unit that extracts a feature map for each of the first data and the second data; a candidate region estimation unit that estimates a target likelihood map, which is a map of likelihood indicating a target region likelihood, from the feature map using a target region estimation model; a target region extraction unit that extracts a target region from the target likelihood map; 9. The data change estimation device according to claim 5, further comprising: a feature extraction unit that extracts features of the extracted target region from the feature map.

10. A learning method in which a learning unit trains a score calculator based on learning data indicating a change from first data to second data, the change being a disappearance of an object represented by the data, and learning data indicating a change from the second data to the first data, the change being an appearance of an object represented by the data, the score calculator calculates a score indicating the change that is the disappearance when receiving the feature amount of the first data and the feature amount of the second data as input in a predetermined order; and The score calculator is trained so as to calculate a score indicating the change that is the occurrence when the feature amount of the first data and the feature amount of the second data are received as input in a reverse order to the predetermined order. How to learn.

11. A learning method in which a learning unit trains a score calculator based on learning data indicating a change from first data to second data, the change being an appearance of an object represented by the data, and learning data indicating a change from the second data to the first data, the change being a disappearance of the object represented by the data, the score calculator calculates a score indicating the change that is the occurrence when the feature amount of the first data and the feature amount of the second data are received as input in a predetermined order; and The score calculator is trained so as to calculate a score indicating the change that is the disappearance when the feature amount of the first data and the feature amount of the second data are received as input in a reverse order to the predetermined order. How to learn.

12. a score calculation unit inputs the feature amount of the first data and the feature amount of the second data to a score calculator in a predetermined order, and calculates a score indicating a change from the first data to the second data, the change being the disappearance of an object represented by the data; and The feature amount of the first data and the feature amount of the second data are input to the score calculator in a reverse order to the predetermined order, and a score indicating a change from the second data to the first data, which is an appearance of an object represented by the data, is calculated. Data change estimation method.

13. a score calculation unit inputs the feature amount of the first data and the feature amount of the second data to a score calculator in a predetermined order, and calculates a score indicating a change from the first data to the second data, the change being an appearance of an object represented by the data; and The feature amount of the first data and the feature amount of the second data are input to the score calculator in a reverse order to the predetermined order, and a score indicating a change from the second data to the first data, which is a disappearance of an object represented by the data, is calculated. Data change estimation method.

14. A learning program for training a score calculator based on learning data indicating a change from first data to second data, the change being the disappearance of an object represented by the data, and learning data indicating a change from the second data to the first data, the change being the appearance of an object represented by the data, the score calculator calculates a score indicating the change that is the disappearance when receiving the feature amount of the first data and the feature amount of the second data as input in a predetermined order; and The score calculator is trained so as to calculate a score indicating the change that is the occurrence when the feature amount of the first data and the feature amount of the second data are received as input in a reverse order to the predetermined order. A learning program that makes a computer do things.

15. A learning program for training a score calculator based on learning data indicating a change from first data to second data, the change being an appearance of an object represented by the data, and learning data indicating a change from the second data to the first data, the change being a disappearance of the object represented by the data, the score calculator calculates a score indicating the change that is the occurrence when the feature amount of the first data and the feature amount of the second data are received as input in a predetermined order; and The score calculator is trained so as to calculate a score indicating the change that is the disappearance when the feature amount of the first data and the feature amount of the second data are received as input in a reverse order to the predetermined order. A learning program that makes a computer do things.

16. inputting the feature amount of the first data and the feature amount of the second data into a score calculator in a predetermined order, and calculating a score indicating a change from the first data to the second data, the change being the disappearance of an object represented by the data; and The feature amount of the first data and the feature amount of the second data are input to the score calculator in a reverse order to the predetermined order, and a score indicating a change from the second data to the first data, which is an appearance of an object represented by the data, is calculated. A data change estimation program that allows a computer to perform this task.

17. inputting the feature amount of the first data and the feature amount of the second data into a score calculator in a predetermined order, and calculating a score indicating a change from the first data to the second data, the change being the appearance of an object represented by the data; and The feature amount of the first data and the feature amount of the second data are input to the score calculator in a reverse order to the predetermined order, and a score indicating a change from the second data to the first data, which is a disappearance of an object represented by the data, is calculated. A data change estimation program that allows a computer to perform this task.

Citation Information

Patent Citations

  • Identification of spurious regions within a video frame

    JP2008535041A

  • Object monitoring device, object monitoring program, and monitoring system

    JP2021111033A

  • Recognition device, learning device, recognition system, recognition method, recognition program, learning method, and learning program

    JP2021114223A

  • Target acquisition and tracking

    US5912980A