Association device and association method

The described device and method address the challenge of associating target regions across multiple images by selecting the most suitable feature quantities based on reliability and certainty, improving the accuracy of target tracking despite variations in image resolution and occlusions.

JP7695911B2Active Publication Date: 2025-06-19SECOM CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022055579
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-06-19
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

Existing image processing technologies face challenges in associating target regions across multiple images due to variations in image resolution and occlusions, which affect the suitability of feature quantities for accurate tracking.

Method used

A device and method that extract multiple types of feature quantities for target regions in images, calculate their reliability for identifying targets, and determine the certainty of association using these feature quantities to preferentially select the most suitable type for association determination.

Benefits of technology

This approach allows for appropriate selection of feature quantities based on their suitability, enhancing the accuracy of target region association across varying image conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695911000002
    Figure 0007695911000002
  • Figure 0007695911000003
    Figure 0007695911000003
  • Figure 0007695911000004
    Figure 0007695911000004
Patent Text Reader

Abstract

To appropriately select the type of a feature amount used for association, in a tracking device which associates a plurality of first object areas showing an object in a first image with a plurality of second object areas showing the object in a second image on the basis of the feature amount of the object area.SOLUTION: A tracking device 1 comprises: a feature amount extraction unit 12 which extracts the plurality of types of feature amounts, for each of the plurality of first object areas and the plurality of second object areas, and calculates a reliability degree indicating a degree to which the feature amount is suitable for identification of the object, for each of the extracted feature amounts; a certainty degree calculation unit 13 which calculates a certainty degree indicating an accuracy degree of the association when associating the first object area with the second object area by using the feature amount of the type, on the basis of the reliability degree, for each type of the feature amount; and an association determination unit 14 which preferentially selects the feature amount of the type having a higher certainty degree, and performs association determination processing of the first object area and the second object area on the basis of the selected type of the feature amount.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a matching device and a matching method for associating target regions representing targets in each of a plurality of images with each other.

Background Art

[0002] Patent Document 1 below describes an image processing apparatus that tracks a tracking target object captured in a partial image of an input image. As a technique for estimating and tracking the positions of a plurality of objects captured in an image, a method called multiple object tracking (MOT) is known. In multiple object tracking, a plurality of target regions, which are regions of partial images each containing a tracking target, are detected from an input image at time t = n, and feature amounts are extracted for each of the detected target regions. Then, by comparing the feature amounts of the plurality of target regions detected from the input image at time t = n - 1 before time t = n, and associating the target region at time t = n with a similar feature amount with the target region at time t = n - 1, a time series of target regions representing the same tracking target is obtained.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] For the association between target regions, it is preferable to use multiple types. This is because the types of feature quantities suitable for the association vary depending on the input image. For example, when the tracking target is a person, if the person's face can be photographed with high resolution, face feature quantities are suitable. However, when the resolution of the face photograph is low, for example, head feature quantities (such as head movement and line of sight) may be more suitable than face feature quantities. Also, when the resolution of the face photograph is low or there is occlusion in part of the face or head due to an occluding object, for example, whole-body feature quantities (such as body shape, height, gender, age, etc.) may be more suitable than face feature quantities.

[0005] On the other hand, when tracking a plurality of target regions such as multi-object tracking, it is not always possible to extract all types of feature quantities for each of these plurality of target regions. The present invention has been made in view of the above problems, and in an association device that associates a plurality of first target regions representing targets in a first image and a plurality of second target regions representing targets in a second image based on the feature quantities of the target regions, an object is to appropriately select the type of feature quantity used for the association.

Means for Solving the Problems

[0006] According to one aspect of the present invention, there is provided an association device that associates each of a plurality of first target regions representing targets in a first image and a plurality of second target regions representing targets in a second image. The tracking device includes: a feature quantity extraction unit that extracts a plurality of types of feature quantities for each of the plurality of first target regions and the plurality of second target regions, and calculates a reliability indicating the degree to which the feature quantity is suitable for identifying the target for each of the extracted feature quantities; a certainty calculation unit that calculates, for each type of feature quantity, a certainty indicating the accuracy of the association when the first target region and the second target region are associated using the feature quantity of the type based on the reliability; and an association determination unit that preferentially selects a feature quantity of a type with higher certainty and performs an association determination process between the first target region and the second target region based on the selected feature quantity.

[0007] According to another aspect of the present invention, there is provided a method of associating a plurality of first target regions representing a target in a first image with each of a plurality of second target regions representing a target in a second image. In the association method, for each of the plurality of first target regions and the plurality of second target regions, at least one type of feature amount among a plurality of types is extracted, and a reliability indicating the degree to which the feature amount is suitable for identifying the target is calculated for each of the extracted feature amounts. For each type of feature amount, a certainty indicating the accuracy of the association when the first target region and the second target region are associated using the feature amount of the type is calculated based on the reliability calculated for at least one of the plurality of first target regions and the plurality of second target regions. A feature amount of a type with higher certainty is preferentially selected, and an association determination process between the first target region and the second target region is performed based on the selected feature amount.

Effect of the Invention

[0008] According to the present invention, in an association device that associates a plurality of first target regions representing a target in a first image and a plurality of second target regions representing a target in a second image based on the feature amounts of the target regions, the type of feature amount used for the association can be appropriately selected.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Embodiments for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the embodiments of the present invention shown below exemplify devices and methods for embodying the technical idea of the present invention, and the technical idea of the present invention does not specify the structure, arrangement, etc. of the components as follows. The technical idea of the present invention can be variously modified within the technical scope defined by the claims described in the claims.

[0011] (First Embodiment) (Configuration) FIG. 1 is a schematic diagram showing a hardware configuration example of the tracking device 1 according to an embodiment of the present invention. The tracking device 1 according to an embodiment of the present invention is a tracking device that detects a target region (that is, the position where the tracking target appears) where the tracking target appears from input images sequentially acquired at different times, and tracks the same tracking target by associating the target regions of the same tracking target with each other. The tracking device 1 is an example of the "association device" described in the claims. For example, when tracking a specific person, the detection target is "person" and the tracking target is "specific person". In the following description, a tracking device that tracks a plurality of persons as tracking targets will be exemplified and described, but the tracking targets in the present invention are not limited to persons. The present invention is widely applicable to tracking devices that track a plurality of objects (for example, a plurality of vehicles, etc.). In addition, the present invention is widely applicable not only to a tracking device that tracks the same tracking target from input images sequentially acquired at a plurality of times, but also to applications for associating images or target regions representing the same object with each other. For example, the images or target regions to be associated with each other do not have to be those taken at a plurality of different times. For example, in object identification for recognizing an object shown in an image, it may be applied to the association between a set of images in which a plurality of objects to be recognized are shown and a set of images in which known objects are shown. For example, in person identification, it may be applied to the association between each target area in which a person is detected in an image showing a plurality of persons, such as a surveillance image of a security camera, and a set of search images in which known persons are shown. Also, for example, when a plurality of objects are shown in each of a plurality of viewpoint images captured at the same time by a plurality of cameras with different viewpoints, it may be applied to the association of the same objects between the plurality of viewpoint images. By obtaining a plurality of viewpoint images of the same object, 3D shape estimation, more precise pose estimation, and behavior estimation become possible.

[0012] The tracking device 1 includes a photographing unit 2, a communication unit 3, a storage unit 4, an image processing unit 5, an output unit 6, and an operation input unit 7. The photographing unit 2 is a security camera installed for the purpose of monitoring a predetermined area, and is attached at a position where a person staying in the area can be photographed. The image photographed by the photographing unit 2 is transmitted to the image processing unit 5 via the communication unit 3.

[0013] The communication unit 3 performs data transmission and reception among the photographing unit 2, the image processing unit 5, the output unit 6, and the operation input unit 7. A LAN (Local Area Network) or a public line such as the Internet can be used. The storage unit 4 is composed of an HDD (Hard Disk Drive) or an SSD (Solid State Drive) or the like, and stores various programs including an operating system and various data. The image processing unit 5 is composed of a CPU, an MPU, peripheral circuits, terminals, various memories, etc., and transmits the result of performing image processing on the image photographed by the photographing unit 2 to the output unit 6 via the communication unit 3.

[0014] The output unit 6 is a display, a projector, a printer, a removable drive device, a USB (Universal Serial Bus) interface, a network interface, etc. that output various information generated by the tracking device 1. The operation input unit 7 is operated by the user and is a mouse, keyboard, etc. for receiving input of specifying a template and a search range.

[0015] FIG. 2 is an explanatory diagram of the outline of the association method of the embodiment. The input image Im n is an image generated by the imaging unit 2 at time t = n and input to the image processing unit 5. The input image Im n-1 is an image generated by the imaging unit 2 at time t = n - 1, which is one processing cycle before time t = n, and input to the image processing unit 5. The image processing unit 5 n detects a target area A 21 in which a person is shown in the input image Im 22 A 23 and A 24 by an object detection technique. For example, the image processing unit 5 can use a known method such as a method based on background difference processing and template matching as the object detection technique.

[0016] On the other hand, the target areas A n-1 in which a person is shown in the input image Im 11 A 12 and A 13 are also detected by the image processing unit 5 one processing cycle before the detection of the target areas A n ~A 21 in the input image Im 24 . The target areas A 11 ~A 13 are associated with the target areas detected before time t = n - 1 to form time series T1, T2, T3 (which may be referred to as "tracks" in the following description) of the target areas. The input image Im n-1 is an example of the "first image" described in the claims, the input image Im n is an example of the "second image" described in the claims, and the target areas A 11 ~A 13 are examples of the "first target area" described in the claims, and the target areas A 21 ~A 24 are examples of the "second target area" described in the claims.

[0017] Note that the input image Im n 's target area A 21 and the input image Im n-1 's target area A 11 are the areas of the partial images in which the same person P a appears. The target area A 22 and the target area A 12 are the areas of the partial images in which the same person P b appears. The target area A 23 and the target area A 13 are the areas of the partial images in which the same person P c appears. The target area A n of the input image Im 24 is the area of the partial image in which a person P n appears in the input image Im but does not appear in the input image Im n-1 . d

[0018] The image processing unit 5 compares all combinations between the target areas A n ~A 21 of the input image Im 24 and the target areas A n-1 ~A 11 ~A 13 of the input image Im 21 ~A 24 and determines the associations m1, m2, and m3. That is, it determines the association between the target areas A The image processing unit 5 extracts the feature amounts of each of the target areas A n ~A 21 ~A 24 of the input image Im. On the other hand, the feature amounts of each of the target areas A n-1 ~A 11 ~A 13 of the input image Im are extracted by the image processing unit 5 one processing cycle before the extraction of the feature amounts of the target areas A 21 ~A 24 and stored in the storage unit 4. The image processing unit 5 determines the associations m1, m2, and m3 between the target areas A 21 ~A 24 and the tracks T1 to T3 based on the extracted feature amounts.​

[0019] That is, among tracks T1 to T3, the target region A 21 with a feature amount similar to the feature amount of the target region A 11 of track T1 is associated with the target region A 21 Similarly, the target region A 22 with a feature amount similar to the feature amount of the target region A 12 of track T2 is associated with the target region A 22 Similarly, the target region A 23 with a feature amount similar to the feature amount of the target region A 13 of track T3 is associated with the target region A 23 That is, among tracks T1 to T3, the target region A

[0020] When determining the association between tracks T1 to T3 and the target regions A 21 ~A 24 optimization processing of the association may be performed. For example, for all patterns of the association between tracks T1 to T3 and the target regions A 21 ~A 24 the total similarity between the feature amounts may be calculated for each pattern, and the association may be determined so that the total similarity becomes high. The image processing unit 5 generates the time series of the target region A 11 ~A 13 that has not been associated with any of them as a new track. 24 That is, among tracks T1 to T3, the target region A

[0021] When associating tracks T1 to T3 with the target regions A 21 ~A 24 it is preferable to use a plurality of types. This is because the types of feature amounts suitable for the association vary depending on the input image. For example, if the face of a person P a ~P d can be photographed with high resolution, the face feature amount is suitable. However, when the resolution of the face photograph is low, for example, the head feature amount may be more suitable than the face feature amount. Also, when the resolution of the face photograph is low or there is occlusion by furniture or the like, for example, the whole body feature amount may be more suitable than the face feature amount. On the other hand, for a plurality of target regions A21 ~A 24 For each of them, it is not always possible to extract all types of feature quantities.

[0022] Therefore, the image processing unit 5 extracts a plurality of types of feature quantities for the target region A 11 ~A 13 and the target region A 21 ~A 24 and calculates the reliability of each of the extracted feature quantities. In this specification, the "reliability" of a feature quantity may be an index value indicating how suitable the feature quantity is for identifying the tracking target shown in the target region. That is, it may be an index value indicating the probability of correctly identifying the tracking target shown in the target region when the feature quantity is used to identify it. In other words, it may be an index value indicating the probability of correctly identifying (i.e., being able to associate) whether it is the same target when the tracking target shown in the target region is identified using the feature quantity. For example, it may be an index value indicating the probability of correctly identifying whether it is the same target without actually using the feature quantity to identify the target. Since the association of the target region by the feature quantity is equivalent to identifying whether two feature quantities are the same target, the reliability of the feature quantity can be regarded as an index value indicating the probability of being able to correctly associate. Next, for each type of feature quantity, the certainty indicating the accuracy of the association when the target region A 11 ~A 13 and the target region A 21 ~A 24 are associated using the feature quantity of that type is calculated based on the reliability. Then, the feature quantity of the type with a higher certainty is preferentially selected, and the association determination process between the target region A 11 ~A 13 and the target region A 21 ~A 24 is performed based on the selected feature quantity. Thereby, the tracking device 1 can appropriately select the feature quantity to be used for the association from among a plurality of types of feature quantities.

[0023] FIG. 3 is a block diagram of a functional configuration example of the tracking device 1 according to the embodiment. The tracking device 1 includes an input image acquisition unit 10, a target region detection unit 11, a feature amount extraction unit 12, a certainty calculation unit 13, a correlation determination unit 14, a track storage unit 15, and a tracking result output unit 16. The imaging unit 2 in FIG. 1 functions as the input image acquisition unit 10, the storage unit 4 functions as the track storage unit 15, the image processing unit 5 functions as the target region detection unit 11, the feature amount extraction unit 12, the certainty calculation unit 13, and the correlation determination unit 14, and the output unit 6 functions as the tracking result output unit 16.

[0024] The input image acquisition unit 10 sequentially acquires, as input images, images obtained by photographing a predetermined monitoring area at each of a plurality of different times t = 1, 2,.... For example, the input image acquisition unit 10 may acquire, as an input image, an image photographed by the imaging unit 2. The target region detection unit 11 detects, by an object detection technique, a plurality of target regions in which a person appears in the input image acquired by the input image acquisition unit 10. For example, the target region detection unit 11 can use a known method such as a method based on background difference processing and template matching as the object detection technique.

[0025] In the following description, a plurality of target regions detected from the input image acquired at time t = n are denoted as "target region A i " (i = 1, 2,...). On the other hand, in the track storage unit 15, information on a plurality of tracks T k which are a time series of target regions detected from input images acquired before time t = n is stored (k = 1, 2,...). In the following description, among the target regions constituting the track T k the target regions detected from the input images acquired in several time instants before time t = n - 1 are denoted as "target region A k ".

[0026] For each of the plurality of target regions A i the feature amount extraction unit 12 extracts, for each of a plurality of feature amount types j, among the target regions A iExtract at least one type of feature quantity that can be extracted from. For example, as the feature quantities of a plurality of feature quantity types j, face feature quantities, head feature quantities, and whole body feature quantities may be extracted. For example, the feature quantity extraction unit 12 inputs the target area A into a feature quantity extraction model modeled by a CNN composed of a multi-layer network such as those used in Deep Learning i to extract the feature quantity of the target area A i .

[0027] In this specification, the case where the feature quantity extraction unit 12 extracts at least one type of feature quantity that can be extracted among the L types of feature quantity types j is exemplified (j = 1, 2,..., L). In the first embodiment described below, a set C of feature quantity types is determined by selecting one or more types from among the L types of feature quantity types j, and the correspondence between the target area A i and the track T k is determined based on the feature quantities of the set C.

[0028] When the number of feature quantity types j included in the set C is 1, the correspondence between the target area A i and the track T k is determined based on a single feature quantity. When the number of a plurality of feature quantity types j included in the set C is 2 or more, the correspondence between the target area A i and the track T k is determined based on the feature quantities f1, f2,... of the plurality of feature quantity types j included in the set C. For example, for each of the feature quantities f1, f2,..., the similarities s1, s2,... between the target area A k of the track T k and the target area A i are calculated, and the total (s1 + s2 +...) or product (s1 × s2 ×...) of the similarities s1, s2,... is obtained as the overall similarity s t , and the correspondence between the target area A t and the track T i is determined according to the overall similarity s k . At this time, the overall similarity s between the feature quantities tTarget area A that is below the threshold i and the track T k are not associated with each other either.

[0029] Furthermore, the feature extraction unit 12 calculates the reliability r of the feature of feature type j extracted from the target area A i For example, the feature extraction unit 12 may calculate the mean and variance of the features of the target area A in the feature space using a generative model such as a variational autoencoder (VAE) or a Gaussian process. The feature extraction unit 12 may obtain the reciprocal of the variance of the features as the reliability r j i However, the present invention is not limited to this, and any method for calculating some index related to the reliability of the features can be widely used as long as it calculates the features and the index. i j i

[0030] For each set C of feature types selected from one or more of the L feature types j, the certainty calculation unit 13 calculates the certainty R indicating the accuracy of the association when the target area A i is associated with the track T k based on the reliability r C calculated for the feature types j included in the set C. j i The association determination unit 14 preferentially selects the set C of feature types j with a higher certainty R, and performs an association determination process of associating the target area A C with the track T i based on the features of the selected set C. k

[0031] Here, it should be noted that not all features included in the selected set C can be successfully extracted for all target areas A i For example, although face features can be successfully extracted for a certain target area, they may not be successfully extracted for other target areas due to low resolution or occlusion.​​​​ Therefore, the association determination unit 14 determines that for the target region A j i or the track T i where the reliability r k of the feature amount included in the selected set C is equal to or less than the threshold value, no association is made.

[0032] In the following description, the target region A k for which the association with the track T i by the association determination unit 14 has been completed is denoted as the "assigned target region", and the target region A k that has not yet been associated with the track T i is denoted as the "unassigned target region". Also, the track T i for which the association with the target region A k has been completed is denoted as the "assigned track", and the track T i that has not yet been associated with the target region A k is denoted as the "unassigned track".

[0033] The association determination unit 14 repeatedly executes the association determination process of associating the unassigned target region with the assigned track while reselecting the set C of the feature amount type j in the processing cycle at time t = n (that is, within each single processing cycle among a plurality of processing cycles at times t = 1, 2,...).

[0034] Specifically, before performing the first association determination process in the processing cycle at time t = n, the association determination unit 14 sets all the target regions A i as the "unassigned target region" and sets all the tracks T k as the "unassigned track". Each time the association determination process is performed once, the certainty calculation unit 13 calculates the certainty R i indicating the accuracy of the association when the target region A j i set in the unassigned target region is associated, based on the reliability r i of the target region A C set in the unassigned target region. The association determination unit 14 determines the certainty R CPrioritize and select a set C of higher feature type j, and determine the association between the unassigned target region and the unassigned track based on the features of the selected set C.

[0035] The association determination unit 14, for a certain target region A i associates it with a certain track T k When it is associated with, for the time series of track T k add target region A i For example, for the information of track T k stored in the track storage unit 15, add the information (position, width, height, etc.) of target region A i to update the information of track T k stored in the track storage unit 15. Also, the association determination unit 14 adds all the feature type j features and reliability r i extracted for target region A j i to the information of track T k to update the information of track T k stored in the track storage unit 15. The association determination unit 14 sets the associated target region A i and track T k as the assigned target region and the assigned track, respectively.

[0036] Each time the association determination unit 14 performs the association determination process, it determines whether a predetermined association end condition is satisfied. For example, when any of the following conditions (A1) to (A3) is satisfied, the association end condition is satisfied. (A1) All target regions A i are set as "assigned target regions". (A2) All tracks T k are set as "assigned tracks". (A3) For all sets C, all certainty degrees Rc calculated based on the reliability r j i of the unassigned target region are below the threshold. Instead of or in addition to (A3), it may be determined that the association end condition is satisfied when the following condition (A4) is satisfied. (A4) The reliability r of all unassigned target regions j i are all below the threshold value.

[0037] When the association end condition is satisfied, the association determination unit 14 determines whether the target region A set in the unassigned target region i remains. The association determination unit 14 generates a time series starting from the target region A set in the unassigned target region i as a new track. For example, information on the unassigned target region (position, width, height, etc.), and the feature amounts and reliability r of all feature amount types j extracted for the unassigned target region j i are stored in the track storage unit 15. When the above processing is completed, the processing cycle at time t = n is completed, and the processing of the next processing cycle at time t = n + 1 is started.

[0038] Next, an example of the calculation method of the certainty R C will be described. When the number of feature amount types j included in the set C is 1, the certainty calculation unit 13 may calculate the certainty R C based on the following formula (1). R C = a1×R j + b1 …(1)

[0039] In formula (1), R j is the reliability r calculated for the target region A i for which the feature amount of feature amount type j can be calculated among the unassigned target regions. j i The statistic R j may be, for example, a moment such as the average, variance, skewness of the reliability r j i , and may also be the median, maximum value, minimum value, quantile, etc. of the reliability r j i . Also, a1 is a predetermined correction coefficient, and b1 is a predetermined correction addition value. By setting different values for the correction coefficient a1 and the correction addition value b1 for each feature type j, priorities can be set for the association based on the feature of feature type j, or the reliability r j i of different value ranges may be normalized.

[0040] When the number of feature types j included in the set C is plural, the certainty calculation unit 13 calculates, for each feature type j, the reliability r i calculated for each of the target regions A j i set in the unassigned target region, which is the reliability distribution r j . FIG. 4 is a schematic diagram of the reliability distribution r j calculated by the certainty calculation unit 13. The certainty calculation unit 13 selects a pair of feature types α and β from the set C, and calculates the difference Dist(r α and r β between them. FIG. 5 is a schematic diagram of a pair of reliability distributions r α , r β for which the difference Dist(r α , r β ) is calculated in the certainty calculation unit 13. For example, the certainty calculation unit 13 may calculate the Kullback-Leibler (KL) divergence between the reliability distributions or the distance based on the inner product of the reliability sequences as the difference Dist(r α , r β ). α , r β ).

[0041] When the difference Dist(r α , r β ) is equal to or less than a predetermined threshold for all pairs of the feature types selected from the set C, the certainty R j is calculated based on the sum ΣR j of the statistic R C according to the following formula (2).

Equation

[0042] On the other hand, when the dissimilarity Dist(r α , r β ) for any pair α, β of the feature amount types selected from set C is not less than a predetermined threshold value, the association determination unit 14 does not perform the association determination process between the target region A i based on the feature amounts of set C and the track T k . For example, when the dissimilarity Dist(r α , r β ) for any pair α, β of the feature amount types selected from set C is not less than a predetermined threshold value, the certainty R C is set to 0 or a very small predetermined value. Thereby, the priority of set C is lowered so that the association determination unit 14 does not use the feature amounts of set C in the association determination process between the target region A i and the track T k .

[0043] The reason will be described below with reference to FIGS. 6(a) and 6(b). FIG. 6(a) is a schematic diagram when the dissimilarity Dist(r α , r β ) between the reliability distributions r α , r β is small, and FIG. 6(b) is a schematic diagram when the dissimilarity Dist(r α , r β ) is large. In the ranges respectively surrounded by the broken lines 20 and 21 in FIG. 6(a), the reliability r of the feature amount type α αi and the reliability r of the feature type β β i are both high. In this way, when the dissimilarity Dist(r α , r β ) is small, a target area A α i where the reliability r β i and the reliability r i are both high is likely to occur. Since the reliability represents an index value indicating the probability of correctly associating the target area using the feature quantity (i.e., being able to correctly identify it), for the target area A α i where the reliability r β i and the reliability r i are both similarly high, when associating using both feature types α and β, the probability of more correctly associating is higher compared to the case of using only one feature type. Therefore, there exists a target area Ai that can be more reliably associated based on the combination of feature types α and β.

[0044] On the other hand, when the dissimilarity Dist(r α , r β ) is large, as shown in Fig. 6(b), a target area A α i where the reliability r β i and the reliability r i are both high is unlikely to occur. Therefore, when performing the association based on such a combination of feature types, it is not known whether the association can be reliably performed, and there is a risk of incorrect association. Therefore, when the dissimilarity Dist(r α , r β ) for any pair α, β of the feature types selected from the set C does not become less than a predetermined threshold value, the association determination process between the target area A i based on the features in the set C and the track T k is not performed.

[0045] In the above description, for the target area A i set in the unassigned target area, the reliability rj i Based on the certainty R C An example of calculating it was described. Instead of this, the track T k Among the target areas that make up the unassigned track, the target area A detected from the input images acquired during several time instants before time t = n - 1 k reliability r j k Based on the certainty R C may be calculated. As described above, the association determination unit 14 pairs the target area A i with the track T k When making the association, for the target area A i the feature amounts of all feature amount types j extracted for the target area A and the reliability r j i are added to the information of the track T stored in the track storage unit 15. k Therefore, in the track storage unit 15, for the target area A of the track T k the feature amounts of all feature amount types j extracted for the target area A and the reliability r k are stored. j k

[0046] The certainty calculation unit 13 may read out the reliability r j k from the track storage unit 15 and calculate the certainty R j k based on the reliability r C The calculation of the certainty R based on the reliability r j k is the same as the calculation method of the certainty R C based on the above-described reliability r j i based on the reliability r C is the same. The association determination unit 14 preferentially selects the set C of feature amount types j with a higher certainty R j k based on the reliability r C and determines the association between the unassigned target area and the unassigned track based on the feature amounts of the selected set C.

[0047] ​In addition, the certainty calculation unit 13 calculates the certainty R i based on the reliability r j i of the target area A C (hereinafter sometimes referred to as "certainty R1" C ), and combines it with the certainty R k based on the reliability r k of the target area A j k of the track T C (hereinafter sometimes referred to as "certainty R2" C ) to calculate a combined certainty R3 C . For example, the certainty calculation unit 13 calculates the combined certainty R3 j i as the product of the certainty R1 C based on the reliability r j k and the certainty R2 C based on the reliability r C =R1 C ×R2 C .

[0048] For example, the certainty calculation unit 13 calculates the combined certainty R3 j i as the sum of the certainty R1 C based on the reliability r j k and the certainty R2 C based on the reliability r C =R1 C +R2 C . At this time, when either one of the certainty R1 C or the certainty R2 C is 0 or a very small predetermined value, the certainty calculation unit 13 sets the combined certainty R3 C to 0 or a very small predetermined value. The association determination unit 14 preferentially selects the set C of feature type j with a higher combined certainty R3 C and determines the association between the unassigned target area and the unassigned track based on the features of the selected set C.

[0049] Referring to FIG. 3. The association determination unit 14 determines the target area Ai The information (position, width, height, etc.) and the target area A i are associated with the track T k and its identification information Id are output to the tracking result output unit 16. Also, when a new track is generated, the information of the target area A i constituting the generated track and the identification information Id of the newly generated track are output to the tracking result output unit 16.

[0050] The tracking result output unit 16 displays the input images acquired by the input image acquisition unit 10 at a plurality of different times t = 1, 2,... on the output unit 6 at the respective different times. For example, it is displayed on the display of the output unit 6. At that time, based on the information of the target area A i output from the association determination unit 14, the position information of the target area A i detected from the input image is output. For example, a figure (e.g., a rectangle) representing the target area A i may be superimposed on the input image and displayed on the display.

[0051] Furthermore, the tracking result output unit 16, based on the identification information Id output from the association determination unit 14, outputs from the output unit 6 information for specifying the target area A i in which the same tracking target is detected from the input images at a plurality of different times t = 1, 2,.... For example, figures representing the target area A i in which the same tracking target is detected are displayed in the same color and brightness, and figures representing the target area A i in which different tracking targets are detected may be displayed in different colors or different brightnesses. Also, for example, the same symbol, mark, figure, character (e.g., identification information, etc.) may be added to and displayed for the target area A i in which the same tracking target is detected.

[0052] (Operation) FIG. 7 is a flowchart of an example of the association method according to the first embodiment. In step S1, the input image acquisition unit 10 acquires the input image Im n captured at time t = n. In step S2, the target area detector 11 detects a plurality of target areas A in the input image Im n in which a person is depicted i . In step S3, the feature extractor 12 extracts at least one type of feature that can be extracted from each of the plurality of target areas A i from among a plurality of feature type categories j i .

[0053] In step S4, the feature extractor 12 calculates the reliability r i of the feature of the feature type category j extracted from the target area A j i . In step S5, the association determination unit 14 sets all the target areas A i as "unassigned target areas" and sets all the tracks T k as "unassigned tracks". In step S6, the certainty calculation unit 13 calculates the certainty R j i based on the reliability r C of the unassigned target area. The certainty calculation unit 13 may read the reliability r k of the feature of the feature type category j extracted from the target area A k of the unassigned track T from the track storage unit 15 and calculate the certainty R j k based on the reliability r j k of the unassigned track. The certainty calculation unit 13 may calculate the combined certainty R3 C . C

[0054] In step S7, the association determination unit 14 selects a set C of feature type categories j according to the priority based on the certainty R C or the combined certainty R3 C . In step S8, the association determination unit 14 determines the association between the unassigned target area and the unassigned track based on the features of the selected set C ​In step S9, the association determination unit 14 sets the target region A associated in step S8 i and the track T k as the assigned target region and the assigned track, respectively.

[0055] In step S10, the association determination unit 14 determines whether a predetermined association end condition is satisfied. If the association end condition is satisfied (step S10: Y), the process proceeds to step S11. If the association end condition is not satisfied (step S10: N), the process returns to step S6. In step S11, the association determination unit 14 determines whether the target region A set in the unassigned target region i remains. If the target region A set in the unassigned target region i remains (step S11: Y), the process proceeds to step S12. If the target region A set in the unassigned target region i does not remain (step S11: N), the process ends. In step S12, the association determination unit 14 generates a time series starting from the target region A set in the unassigned target region i as a new track. Then the process ends.

[0056] (Second Embodiment) In the second embodiment described below, one type is selected from among L feature amount types j, and based on the feature amount of the selected one type, the target region A i and the track T k are determined for association. The certainty calculation unit 13 calculates, for each of the L feature amount types j, the certainty R i indicating the accuracy of the association when the target region A k is associated with the track T C based on the reliability r j i For example, when the number of feature amount types j included in the set C is 1, the certainty calculation unit 13 may calculate the certainty R C of the above formula (1).

[0057] The association determination unit 14 sets the priority of the feature type j according to the confidence level R. C For example, the higher the confidence level R, C the higher the priority may be set. In the processing cycle at time t = n (that is, within each single processing cycle among a plurality of processing cycles at times t = 1, 2,...), the association determination unit 14 repeatedly executes an association determination process of associating an unassigned target area with an assigned track while sequentially selecting from the features of the feature type j with the highest priority based on the selected features.

[0058] At this time, the association determination unit 14 does not associate the target area A j i whose reliability r is below the threshold value i or the track T. k Also, the combination of the target area A i and the track T k whose similarity between features is below the threshold value is not associated either. Each time the association determination unit 14 performs the association determination process, it determines whether a predetermined association end condition is satisfied. For example, when any of the following conditions (B1) to (B4) is satisfied, the association end condition is satisfied. (B1) Features of all feature types j have been used in the association process. (B2) All target areas A i have been set as "assigned target areas". (B3) All tracks T k have been set as "assigned tracks". (B4) The reliability r of all unassigned target areas j i is all below the threshold value. (B5) The confidence level R of the selected feature type j C is below the threshold value.

[0059] When the association end condition is satisfied, the association determination unit 14 sets the target area A set in the unassigned target area. iDetermine whether it remains. The association determination unit 14 sets the target region A set in the unassigned target region i The time series starting from is generated as a new track When the above processing is completed, the processing cycle at time t = n is completed, and the processing of the next processing cycle at time t = n + 1 is started

[0060] Note that the confidence calculation unit 13, similar to the first embodiment, uses the track T set in the unassigned track k Of the target regions constituting the target region A detected from the input image acquired during several time instants before time t = n - 1 k Confidence r j k Based on this, the confidence R C May be calculated The association determination unit 14 uses the confidence r j k Based on the confidence R C May set the priority of the feature type j accordingly Also, the confidence calculation unit 13, similar to the first embodiment, uses the confidence r of the target region A i Confidence r j i Based on this, the confidence R C (Hereinafter sometimes referred to as "confidence R1 C ") and the confidence r of the target region A of the track T k Of the target region A k Confidence r j k Based on this, the confidence R C (Hereinafter sometimes referred to as "confidence R2 C ") and the combined confidence R3 C May be calculated. The association determination unit 14 may set the priority of the feature type j according to the combined confidence R3 C

[0061] (Operation) FIG. 8 is a flowchart of an example of the association method according to the second embodiment The processing of steps S20 to S23 is the same as the processing of steps S1 to S4 in FIG. 7 ​In step S24, the certainty calculation unit 13 calculates, for each of the L feature quantity types j, a certainty R indicating the accuracy of the association when the target region A i is associated with the track T k based on the reliability of the feature quantity of the feature quantity type j. C is calculated. In step S25, the association determination unit 14 sets all the target regions A i as "unassigned target regions" and sets all the tracks T k as "unassigned tracks". Also, the priority of the feature quantity type j is set according to the certainty R C .

[0062] In step S26, the association determination unit 14 sets a variable P for designating the priority to 1. In step S27, the association determination unit 14 selects the feature quantity of the feature quantity type j whose priority based on the certainty R C is the P-th. In step S28, the association determination unit 14 determines the association between the unassigned target region and the unassigned track based on the selected type of feature quantity. In step S29, the association determination unit 14 sets the target region A i associated in step S28 and the track T k as an assigned target region and an assigned track, respectively.

[0063] In step S30, the association determination unit 14 increases the value of the variable P by 1 (that is, increments the variable P). In step S31, the association determination unit 14 determines whether or not a predetermined association end condition is satisfied. If the association end condition is satisfied (step S: Y), the process proceeds to step S. If the association end condition is not satisfied (step S: N), the process proceeds to step S. The processes of steps S32 and S33 are the same as the processes of steps S11 and S12 in FIG. 7.

[0064] (Effect of the Embodiment) (1) The tracking device 1 associates each of a plurality of first target regions representing a target in a first image with each of a plurality of second target regions representing the target in a second image. The tracking device 1 extracts at least one type of feature amount out of a plurality of types for each of the plurality of first target regions and the plurality of second target regions, and calculates a reliability indicating the degree to which the feature amount is suitable for identifying the target for each of the extracted feature amounts. A feature amount extraction unit 12, for each type of feature amount, calculates a certainty indicating the accuracy of the association when the first target region and the second target region are associated using the feature amount of the type, based on the reliability calculated for at least one of the plurality of first target regions and the plurality of second target regions. A certainty calculation unit 13, and a correspondence determination unit 14 that preferentially selects a type of feature amount with a higher certainty and performs a correspondence determination process between the first target region and the second target region based on the selected type of feature amount.

[0065] Thereby, when associating the plurality of first target regions with the plurality of second target regions, a more suitable type of feature amount can be appropriately selected and used from among the plurality of types of feature amounts. As a result, the association can be performed using a type of feature amount suitable for that time. In addition, the certainty can be calculated based on the assumption that there is a correlation between the reliability of the feature amount and the accuracy of the association based on the feature amount. Note that the plurality of first target regions may be regions set within a single image, or may be regions set in each of the plurality of images. Similarly, the plurality of second target regions may be regions set within a single image, or may be regions set in each of the plurality of images.

[0066] (2) The certainty calculation unit 13 may recalculate the certainty based on the reliability calculated for at least one of a first unassigned target region that is a first target region that cannot be associated based on the selected type of feature amount and a second unassigned target region that is a second target region that cannot be associated based on the selected type of feature amount. The correspondence determination unit 14 may preferentially select a type of feature amount with a higher recalculated certainty and perform a correspondence determination process between the first unassigned target region and the second unassigned target region based on the selected type of feature amount. When there is a target area that cannot be appropriately associated based on the feature amount of the selected type, if the feature amount of a type suitable for this target area is reselected and the association determination process is performed, the target areas that can be associated with an appropriate type of feature amount can be increased. As a result, since the association can be performed using the feature amount of the type suitable for each target area, a more appropriate association becomes possible.

[0067] (3) The certainty calculation unit 13 may set a set in which one or more types are selected from a plurality of types, and for each set, calculate a certainty indicating the accuracy of the association when the first target area and the second target area are associated using the feature amounts of the types included in this set. The association determination unit 14 may select a set with a higher certainty and perform the association determination process using the feature amounts of the types included in the selected set. This makes it possible to appropriately select the combination of feature amounts used for the association.

[0068] (4) The certainty calculation unit 13 calculates a reliability distribution that is a distribution of reliabilities calculated for each of a plurality of target areas that are at least one of the plurality of first target areas or the plurality of second target areas, and when the difference between the reliability distributions calculated for the reliabilities of the types included in a set including two or more types is equal to or less than a threshold value, the certainty of the set including two or more types may be calculated based on the reliabilities of the types included in this set. This makes it possible to select a set of feature amounts that is advantageous for the association when performing the association determination process based on a set of feature amounts of two or more types.

[0069] (5) The certainty calculation unit 13 may perform the association determination process between the first target area that cannot be associated based on the selected type of feature amount and the second target area based on a type of feature amount having a lower certainty than the certainty of the selected type. If there is a target area that cannot be appropriately associated based on the feature amount of the selected type, the type of feature amount suitable for this target area is reselected and the association determination process is performed. As a result, the target areas that can be associated with the appropriate type of feature amount can be increased.

[0070] (6) The confidence calculation unit 13 may calculate a statistical amount of the confidence calculated for at least one of the plurality of first target areas and the plurality of second target areas, and calculate the confidence based on the calculated statistical amount. By calculating the statistical amount of the confidence calculated for each of the plurality of target areas, a scalar value summarizing the characteristics of these confidences can be calculated as the confidence.

[0071] (7) The confidence calculation unit 13 may calculate the confidence by correcting the statistical amount according to the type of the feature amount. Thereby, a priority can be set for the association by the feature amount of the type of the feature amount, or the confidence values with different value ranges can be normalized. (8) The association determination unit 14 may not associate a first target area or a second target area whose calculated confidence is lower than the threshold value with the first target area or the second target area, respectively. Thereby, it is possible to suppress inappropriate association by a feature amount with low confidence.

Description of Reference Numerals

[0072] 1... Tracking device, 2... Photographing unit, 3... Communication unit, 4... Storage unit, 5... Image processing unit, 6... Output unit, 7... Operation input unit, 10... Input image acquisition unit, 11... Target area detection unit, 12... Feature amount extraction unit, 13... Confidence calculation unit, 14... Association determination unit, 15... Track storage unit, 16... Tracking result output unit

Claims

1. A correlation device for correlating each of a plurality of first target regions representing a target in a first image with each of a plurality of second target regions representing the target in a second image, For each of the plurality of first target regions and the plurality of second target regions, at least one type of feature amount among a plurality of types is extracted, and a reliability indicating the degree to which the feature amount is suitable for identifying the target is calculated for each of the extracted feature amounts, a feature amount extraction unit; For each type of the feature amounts, a certainty indicating the accuracy of the association when the first target region and the second target region are associated using the feature amount of the type is calculated based on the reliability calculated for at least one of the plurality of first target regions and the plurality of second target regions, a certainty calculation unit; A correlation determination unit that preferentially selects a feature amount of a type with a higher certainty and performs a correlation determination process between the first target region and the second target region based on the selected type of feature amount; A correlation device characterized by comprising:

2. The certainty calculation unit recalculates the certainty based on the reliability calculated for at least one of a first unassigned target region, which is a first target region that cannot be associated based on the selected type of feature amount, and a second unassigned target region, which is a second target region that cannot be associated based on the selected type of feature amount, The correlation determination unit preferentially selects a feature amount of a type with a higher recalculated certainty and performs a correlation determination process between the first unassigned target region and the second unassigned target region based on the selected type of feature amount. The correlation device according to claim 1, characterized in that.

3. The certainty calculation unit sets a set of one or more types selected from the plurality of types, and for each set, calculates the certainty indicating the accuracy of the association when the first target region and the second target region are associated using the feature amounts of the types included in the set, The association determination unit selects the set with a higher degree of certainty, and performs the association determination process using the feature amounts of the types included in the selected set. The association device according to claim 1 or 2, characterized in that.

4. The confidence level calculation unit calculates a confidence distribution, which is a distribution of the confidence levels calculated for each of a plurality of target regions, where at least one of the plurality of first target regions or the plurality of second target regions is included, when the difference degree between the confidence distributions calculated for the confidence levels of the types included in the set including two or more types is equal to or less than a threshold value, calculates the certainty of the set including the two or more types based on the confidence levels of the types included in the set. The association device according to claim 3, characterized in that.

5. The association determination process for the first target region and the second target region that cannot be associated based on the selected type of feature amount is performed by the association determination unit based on a type of feature amount with a lower degree of certainty than the degree of certainty of the selected type, according to the association device of claim 1.

6. The confidence level calculation unit calculates a statistic of the confidence levels calculated for at least one of the plurality of first target regions and the plurality of second target regions, and calculates the certainty based on the calculated statistic, according to the association device of any one of claims 1 to 5.

7. The confidence level calculation unit calculates the certainty by correcting the statistic according to the type of the feature amount, according to the association device of claim 6.

8. The association determination unit does not associate the first target region or the second target region, for which the calculated confidence level is lower than the threshold value, with the first target region or the second target region, respectively, according to the association device of any one of claims 1 to 7.

9. A method for associating a plurality of first target regions representing a target in a first image with each of a plurality of second target regions representing the target in a second image, For each of the plurality of first target regions and the plurality of second target regions, at least one type of feature amount among a plurality of types is extracted, and a reliability indicating the degree to which the feature amount is suitable for identifying the target is calculated for each of the extracted feature amounts, For each type of the feature amounts, a certainty indicating the accuracy of the association when the first target region and the second target region are associated using the feature amount of the type is calculated based on the reliability calculated for at least one of the plurality of first target regions and the plurality of second target regions, The feature amount of the type with higher certainty is preferentially selected, and an association determination process between the first target region and the second target region is performed based on the selected type of feature amount. An association method characterized by the above.

Citation Information

Patent Citations

  • Apparatus for selecting feature information applied for image recognition processing, and image recognition processing apparatus

    JP2011060024A

  • Human image processing apparatus, and human image processing method

    JP2013196034A

  • Image processor

    JP2019075051A