Object Position Estimation Device, Object Position Estimation Method, and Recording Medium

By using convolutional processing in object position estimation equipment to generate feature maps and estimate object probability, the problems of limited computer processing speed and difficulty in estimating overlap position of objects are solved, and a robust and high-precision object position estimation is achieved.

CN115720664BActive Publication Date: 2025-06-17NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080102293.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-23
Publication Date
2025-06-17
Estimated Expiration
2040-06-23

AI Technical Summary

Technical Problem

Computer processing speed is limited, making it difficult to continuously and comprehensively change the position and size of some areas in the image, especially when objects overlap with each other, it is difficult to accurately estimate the position of each object.

Method used

Using feature extraction units and likelihood map estimation units, a feature map is generated by performing convolution processing on the target image, and using these feature maps to estimate the probability of an object present at each position, achieving a robust and high-precision object position estimation.

Benefits of technology

Even if the objects overlap each other in the image, the position of each object can be estimated robustly and with high accuracy, solving the problem of position estimation in the case of overlapping objects in traditional techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115720664B_ABST
    Figure CN115720664B_ABST
Patent Text Reader

Abstract

The present invention estimates the corresponding position of an object robustly and with high precision even when the objects overlap in an image. An object position estimation device (1) is provided with: a feature extraction unit (10) including a first feature extraction unit (21) and a second feature extraction unit (22), the first feature extraction unit (21) generating a first feature map by subjecting a target image to convolutional calculation processing, and the second feature extraction unit (22) generating a second feature map by further subjecting the first feature map to convolutional calculation processing; and a likelihood map estimation unit (20) including a first position likelihood estimation unit (23) and a second position likelihood estimation unit (24), the first position likelihood estimation unit (23) estimating a first likelihood map indicating the probability that an object having a first size exists at each position of the target image by using the first feature map, and the second position likelihood estimation unit (24) estimating a second likelihood map indicating the probability that an object having a second size larger than the first size exists at each position of the target image by using the second feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an object position estimation device, an object position estimation method, and a recording medium, and more particularly, to an object position estimation device, an object position estimation method, and a recording medium for estimating the position of an object in an image. Background Art

[0002] Related techniques for estimating the position of an object in an image are known (PTL 1 and PTL2). In the related technique described in NPL 1, an estimator learns the recognition of an object by using a sample image showing the entire object. The estimator thus trained scans the image in order to estimate the position of the object in the image. Specifically, in the related technique described in NPL 1, for example, the estimator estimates the Haar-Like feature amount of the object in the image and estimates the object region of the recognized object. At this time, the estimator scans each partial region while changing the position and size of each partial region in the image.

[0003] [Citation List]

[0004] [Patent Documents]

[0005] [PTL 1] JP 2019-096072 A

[0006] [PTL 2] JP 2018-147431 A

[0007] [Non-Patent Documents]

[0008] [NPL 1] "Rapid Object Detection Using a Boosted Cascade of Simple Features", P. Viola et al., CVPR (Conference on Computer Vision and Pattern Recognition), pp. 511-518. Summary of the Invention

[0009] Technical Problem

[0010] The processing speed of a computer is limited. Therefore, when the estimator scans the image, it is difficult to continuously and comprehensively change the position and size of a partial region in the image. In the case where part or all of an object in the image is occluded by another object, it may be difficult to specify the object region in the image and accurately estimate the position of each object.

[0011] The present invention has been made in view of the above problems, and an object of the present invention is to provide an object position estimation device, an object position estimation method, and a recording medium that can robustly and highly accurately estimate the position of each object even when objects in an image overlap with each other.

[0012] Solution to the problem

[0013] An object position estimation device according to an aspect of the present invention includes: a feature extraction unit including a first feature extraction unit and a second feature extraction unit, the first feature extraction unit being configured to generate a first feature map by performing a convolution process on a target image, and the second feature extraction unit being configured to generate a second feature map by further performing a convolution process on the first feature map; and a likelihood map estimation unit including a first position likelihood estimation unit and a second position likelihood estimation unit, the first position likelihood estimation unit being configured to estimate a first likelihood map indicating the probability of the presence of an object having a first size at each position of the target image by using the first feature map, and the second position likelihood estimation unit being configured to estimate a second likelihood map indicating the probability of the presence of an object having a second size larger than the first size at each position of the target image by using the second feature map.

[0014] A position estimation method according to an aspect of the present invention includes: generating a first feature map by performing a convolution process on a target image, and generating a second feature map by further performing a convolution process on the first feature map; and estimating a first likelihood map indicating the probability of the presence of an object having a first size at each position of the target image by using the first feature map, and estimating a second likelihood map indicating the probability of the presence of an object having a second size larger than the first size at each position of the target image by using the second feature map.

[0015] A recording medium according to an aspect of the present invention causes a computer to execute: generating a first feature map by performing a convolution process on a target image, and generating a second feature map by further performing a convolution process on the first feature map; and estimating a first likelihood map indicating the probability of the presence of an object having a first size at each position of the target image by using the first feature map, and estimating a second likelihood map indicating the probability of the presence of an object having a second size larger than the first size at each position of the target image by using the second feature map.

[0016] Advantageous effects of the invention

[0017] According to one aspect of the present invention, even when objects overlap with each other in an image, the position of each object can be robustly and highly accurately estimated. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a block diagram showing the configuration of an object position estimation device according to Example Embodiment 1.

[0019] Figure 2 is a block diagram showing the configuration of a system including an object position estimation device according to Example Embodiment 2.

[0020] Figure 3 is a flowchart showing the processing flow executed by each unit of the object position estimation device according to Example Embodiment 2.

[0021] Figure 4 is a block diagram showing the configuration of a modified object position estimation device according to Example Embodiment 2.

[0022] Figure 5 is a block diagram showing the configuration of an object position estimation device according to Example Embodiment 3.

[0023] Figure 6 is a block diagram showing the configuration of an object position estimation device according to Example Embodiment 4.

[0024] Figure 7 is a block diagram showing the configuration of an object position estimation device according to Example Embodiment 5.

[0025] Figure 8 is a block diagram showing the configuration of an object position estimation device according to Example Embodiment 6.

[0026] Figure 9 is a flowchart showing the processing flow executed by each unit of the object position estimation device according to Example Embodiment 6.

[0027] Figure 10 is a block diagram showing the configuration of a modified object position estimation device according to Example Embodiment 6.

[0028] Figure 11 is a diagram for explaining the processing flow in which the training data generation unit of the modified object position estimation device according to Example Embodiment 6 generates the first correct likelihood map / the second correct likelihood map.

[0029] Figure 12 is a block diagram showing the configuration of an object position estimation device according to Example Embodiment 7.

[0030] Figure 13 is a block diagram showing the configuration of a modified object position estimation device according to Example Embodiment 7.

[0031] Figure 14 is a diagram showing the hardware configuration of the object position estimation device according to any one of Example Embodiments 1 to 7. Detailed Description

[0032] [Example Embodiment 1]

[0033] Example embodiment 1 will be described with reference to Figure 1 .

[0034] (System)

[0035] The system according to Example embodiment 1 will be described with reference to Figure 1 . Figure 1 The configuration of the system according to Example embodiment 1 is schematically shown. As Figure 1 shown, the system according to Example embodiment 1 includes an image acquisition device 90 and an object position estimation device 1. The image acquisition device 90 acquires one or more images. For example, the image acquisition device 90 acquires a still image output from a video device such as a camera, or an image frame of a moving image output from a video device such as a video recorder.

[0036] The image acquisition device 90 sends the acquired one or more images (for example, a still image or an image frame of a moving image) to the object position estimation device 1. Hereinafter, the image sent from the image acquisition device 90 to the object position estimation device 1 is referred to as a target image 70. The operation of the object position estimation device 1 is controlled by, for example, a computer program.

[0037] (Object position estimation device 1)

[0038] As Figure 1 shown, the object position estimation device 1 includes a feature extraction unit 10 and a likelihood map estimation unit 20. The likelihood map estimation unit 20 is an example of a likelihood map estimation device.

[0039] The feature extraction unit 10 includes a first feature extraction unit 21 and a second feature extraction unit 22. The likelihood map estimation unit 20 includes a first position likelihood estimation unit 23 and a second position likelihood estimation unit 24. The object position estimation device 1 may include three or more feature extraction units and three or more position likelihood estimation units. The first feature extraction unit 21 and the second feature extraction unit 22 are examples of a first feature extraction device and a second feature extraction device. The first position likelihood estimation unit 23 and the second position likelihood estimation unit 24 are examples of a first position likelihood estimation device and a second position likelihood estimation device.

[0040] The first feature extraction unit 21 generates a first feature map indicating the features of an object by performing a convolution process on the target image 70. Specifically, the first feature extraction unit 21 applies a first filter to a matrix in which the target image 70 is represented by pixel values, while sliding the first filter by a predetermined amount of movement. The first filter is a matrix (kernel) that multiplies a part (referred to as a partial region) of the matrix representing the target image 70 by pixel values. The first feature extraction unit 21 outputs the sum of the values obtained through matrix operations between a part of the matrix representing the target image 70 by pixel values and the matrix representing the first filter as an element of the first feature map. The first feature extraction unit 21 outputs the first feature map including a plurality of elements to the first position likelihood estimation unit 23 of the likelihood map estimation unit 20.

[0041] The second feature extraction unit 22 further performs a convolution process on the first feature map to generate a second feature map indicating the features of the object. Specifically, the second feature extraction unit 22 applies a second filter to the first feature map while sliding the second filter by a predetermined amount of movement, and outputs the sum of the values obtained through matrix operations between a part of the matrix of the first feature map and the matrix representing the second filter as an element of the second feature map. Specifically, the second filter is a matrix that multiplies a part of the first feature map. The second feature extraction unit 22 outputs the second feature map including a plurality of elements to the second position likelihood estimation unit 24 of the likelihood map estimation unit 20.

[0042] Using the first feature map received from the first feature extraction unit 21, the first position likelihood estimation unit 23 estimates a first likelihood map that indicates the probability of the presence of an object having a first size at each position of the target image 70. Specifically, as the first position likelihood estimation unit 23, an estimation unit (in one example, a CNN; Convolutional Neural Network). The trained estimation unit estimates the position (likelihood map) of an object having a first size in the target image 70 based on the first feature map. The first size indicates any shape and size within a first predetermined range (described later) included in the target image 70.

[0043] The first position likelihood estimation unit 23 calculates the likelihood of an object of the first size for each partial region of the target image 70, that is, the probability of being an object having the first size. The first position likelihood estimation unit 23 estimates the first likelihood map that represents, in terms of likelihood, the likelihood of an object of the first size calculated for each partial region of the target image 70. The likelihood at each coordinate of the first likelihood map indicates the probability of the presence of an object having the first size at the relevant position in the target image 70. The first position likelihood estimation unit 23 outputs the first likelihood map estimated in this way.

[0044] The second position likelihood estimation unit 24 uses the second feature map to estimate a second likelihood map that indicates the probability of the presence of an object having a second size at each relevant position in the target image 70. Specifically, the second feature extraction unit 22 calculates the likelihood of an object having a second size for each partial region of the target image 70, that is, the probability of being an object having a second size. The second feature extraction unit 22 estimates a second likelihood map that represents, in terms of likelihood, the likelihood of an object having a first size for each partial region of the target image 70. The likelihood at each coordinate of the second likelihood map indicates the probability of the presence of an object having a second size at a relevant position in the target image 70. The second position likelihood estimation unit 24 outputs the second likelihood map estimated in this way. The second size indicates any size within a second predetermined range (described later) in the target image 70.

[0045] Hereinafter, the object may be referred to as an "object having a first size", which has the same meaning as an "object of the first size". There are cases where an "object having a second size" is mentioned with the same meaning as a "second-sized object".

[0046] Alternatively, the first position likelihood estimation unit 23 and the second position likelihood estimation unit 24 estimate the positions of objects having different attributes for each attribute of the pre-classified objects. Then, the first position likelihood estimation unit 23 and the second position likelihood estimation unit 24 estimate a first likelihood map / second likelihood map for each attribute of the objects and output the first likelihood map / second likelihood map for each attribute of the objects. The first position likelihood estimation unit 23 and the second position likelihood estimation unit 24 may have different network configurations or may have a single network configuration for each attribute. In this case, both the first position likelihood estimation unit 23 and the second position likelihood estimation unit 24 output a plurality of likelihood maps in the channel direction of the attributes.

[0047] (Effect of this exemplary embodiment)

[0048] According to the configuration of this exemplary embodiment, the first feature extraction unit 21 of the feature extraction unit 10 generates a first feature map indicating the features of an object by performing a convolution process on the target image 70. The second feature extraction unit 22 of the feature extraction unit 10 further performs a convolution process on the first feature map to generate a second feature map indicating the features of the object. The first position likelihood estimation unit 23 of the likelihood map estimation unit 20 uses the first feature map to estimate a first likelihood map that indicates the probability of the presence of an object having a first size at each position of the image. The second position likelihood estimation unit 24 of the likelihood map estimation unit 20 uses the second feature map to estimate a second likelihood map that indicates the probability of the presence of an object having a second size greater than the first size at each position of the image.

[0049] As described above, the object position estimation device 1 uses the first feature map and the second feature map to estimate the positions of the objects having the first size and the objects having the second size in the target image 70, respectively. Therefore, even if the objects overlap each other in the image, the positions of each object can be estimated robustly and with high accuracy.

[0050] [Example Embodiment 2]

[0051] Example Embodiment 2 will be described with reference to Figure 2 and Figure 3 .

[0052] (Object Position Estimation Device 2)

[0053] As Figure 2 shown, the object position estimation device 2 includes a first feature extraction unit 21, a second feature extraction unit 22, a first position likelihood estimation unit 23, and a second position likelihood estimation unit 24.

[0054] The object position estimation device 2 acquires the target image 70 from the image acquisition device 90. The object position estimation device 2 estimates the position of a predetermined type of object (hereinafter, simply referred to as an object) included in the target image 70. For example, the object position estimation device 2 estimates the position of a person, a car, a tree, an animal, an umbrella, or a part thereof. Hereinafter, an example in which the object is a human head will be described.

[0055] In this Example Embodiment 2, the likelihood at each coordinate of the first likelihood map / second likelihood map output by the object position estimation device 2 indicates the probability that a human head (as an example of an object) having the first size / second size exists at each relevant position in the target image 70. The likelihoods in the first likelihood map / second likelihood map are normalized such that the sum of the likelihoods in each of the first likelihood map / second likelihood map matches the number of human heads having the first size / second size appearing in the target image 70. As a result, the sum of all the likelihoods in each of the first likelihood map / second likelihood map is related to the total number of human heads having the first size / second size appearing in the target image 70. Normalization of the likelihoods in the first likelihood map / second likelihood map is not necessary.

[0056] The first feature extraction unit 21 performs a convolution process on the target image 70 to generate a first feature map 80 indicating object features. For example, the first feature extraction unit 21 may be a convolutional neural network (CNN). The first feature extraction unit 21 outputs the first feature map 80 to each of the first position likelihood estimation unit 23 and the second feature extraction unit 22.

[0057] The first feature map 80 is input from the first feature extraction unit 21 to the first position likelihood estimation unit 23. The first position likelihood estimation unit 23 estimates a first likelihood map by performing a convolution process on the first feature map 80. For example, the first position likelihood estimation unit 23 can be used as a convolutional neural network separately or integrated with the first feature extraction unit 21. As described above, the likelihood at each coordinate of the first likelihood map indicates the probability that an object of the first size exists at each relevant position in the target image 70. As described above, the first size indicates any shape and size within a first predetermined range (described later) included in the target image 70. The first position likelihood estimation unit 23 outputs the estimated first likelihood map.

[0058] The second feature extraction unit 22 obtains the first feature map 80 from the first feature extraction unit 21. The second feature extraction unit 22 further performs a convolution process on the first feature map 80 to generate a second feature map 81 indicating object features. The data size of the second feature map 81 is smaller than the data size of the first feature map 80. The second feature extraction unit 22 outputs the second feature map 81 to the second position likelihood estimation unit 24.

[0059] As described above, the data size of the first feature map 80 is relatively larger than the data size of the second feature map 81. That is, each element of the first feature map 80 is related to the features of a small partial region of the target image 70. Therefore, the first feature map 80 is suitable for capturing fine features of the target image 70. On the other hand, each element of the second feature map 81 is related to the features of a large local area of the target image 70. Therefore, the second feature map 81 is suitable for capturing rough features of the target image 70.

[0060] In Figure 2 the first feature extraction unit 21 and the second feature extraction unit 22 of the object position estimation device 2 are shown as separate functional blocks. However, the first feature extraction unit 21 and the second feature extraction unit 22 can form an integrated network. In this case, the first half of the integrated network corresponds to the first feature extraction unit 21, and the second half of the integrated network corresponds to the second feature extraction unit 22.

[0061] The second feature map 81 is input from the second feature extraction unit 22 to the second position likelihood estimation unit 24. The second position likelihood estimation unit 24 estimates a second likelihood map by performing a convolution process on the second feature map 81. As described above, the likelihood at each coordinate of the second likelihood map indicates the probability that an object of the second size exists at each relevant position in the target image 70. As described above, the second size indicates any size within a second predetermined range (described later) in the target image 70.

[0062] Alternatively, the second feature extraction unit 22 may generate a second feature map based on the target image 70 itself. In this case, the second feature extraction unit 22 acquires the target image 70 instead of the first feature map 80. The second feature extraction unit 22 generates the second feature map 81 by performing a convolution process on the target image 70.

[0063] In Figure 2 FIG., the first feature extraction unit 21, the second feature extraction unit 22, the first position likelihood estimation unit 23, and the second position likelihood estimation unit 24 of the object position estimation device 2 are shown as separate functional blocks. However, the first feature extraction unit 21, the second feature extraction unit 22, the first position likelihood estimation unit 23, and the second position likelihood estimation unit 24 may form an integrated network.

[0064] The first position likelihood estimation unit 23 estimates the position of an object of a first size within a first predetermined range. In other words, in the case where an object existing in the target image 70 has a first size, the first position likelihood estimation unit 23 estimates a first likelihood map.

[0065] On the other hand, the second position likelihood estimation unit 24 estimates the position of an object of a second size within a second predetermined range. That is, in the case where an object existing in the target image 70 has a second size, the position of the object is estimated by the second position likelihood estimation unit 24. The second size is larger than the first size. The first predetermined range defining the first size and the second predetermined range defining the second size are determined in advance so as not to overlap each other.

[0066] For example, the first predetermined range and the second predetermined range are determined based on the data sizes of the relevant first feature map 80 and the relevant second feature map 81, respectively. For example, the reference size of an object in the target image 70 (hereinafter referred to as the first reference size) is first determined using the first feature map 80. Next, another reference size of an object in the target image 70 (hereinafter referred to as the second reference size) is determined using the second feature map 81.

[0067] Specifically, the first reference size is T1, and the second reference size is T2. At this time, the first predetermined range is determined as a*T1 < k ≤ b*T1 using the first reference size T1 and constants a and b (0 < a < b). Here, k represents the size of the object. On the other hand, the second predetermined range is determined as c*T2 < k ≤ d*T2 using the second reference size T2 and constants c and d (0 < c < d).

[0068] The constants (a, b) for defining the first predetermined range and the constants (c, d) for defining the second predetermined range may be equal to or different from each other. Preferably, the condition b*T1 = c*T2 is satisfied so that there is no gap between the first predetermined range and the second predetermined range.

[0069] The supplementary reference size and the predetermined range will be added. As described above, each reference size is determined based on the data size of each feature map, and specifically, each reference size is determined to be a size proportional to the reciprocal of the data size of each feature map. The reference size has a proportional relationship with the predetermined range. Therefore, each predetermined range is determined to have a size proportional to the reciprocal of the data size of each feature map.

[0070] The training method of each unit (i.e., the first feature extraction unit 21, the second feature extraction unit 22, the first position likelihood estimation unit 23, and the second position likelihood estimation unit 24) included in the object position estimation device 2 according to the present exemplary embodiment 2 will be described in exemplary embodiment 6 to be described later. The training function may be provided in the object position estimation device 2 or may be provided in another device other than the object position estimation device 2. In the latter case, the object position estimation device 2 acquires each unit pre-trained by another device.

[0071] "Acquiring each trained unit" described herein may be acquiring the network itself related to each unit (i.e., the program in which the training parameters are set), or may be acquiring only the trained parameters. In the latter case, the object position estimation device 2 acquires the trained parameters from another device and sets the trained parameters in the program prepared in advance in the recording medium of the object position estimation device 2.

[0072] As described above, the first feature map 80 is suitable for capturing the fine features of the target image 70. The first position likelihood estimation unit 23 uses the first feature map 80 to estimate the position of an object having a first size (an object that appears small in the image) in the target image 70. On the other hand, the second feature map 81 is suitable for capturing the rough features of the target image 70. The second position likelihood estimation unit 24 uses the second feature map 81 to estimate the position of an object having a second size larger than the first size (an object that appears large in the image).

[0073] The object position estimation device 2 according to the present exemplary embodiment 2 can effectively estimate the positions of an object having a first size and an object having a second size in the target image 70 by using the first feature map 80 and the second feature map 81 in combination.

[0074] The first position likelihood estimation unit 23 can calculate the total number of objects with the first size in the target image 70 by summing the overall likelihood of the normalized first likelihood map. The second position likelihood estimation unit 24 can calculate the total number of objects with the second size by summing the overall likelihood of the normalized second likelihood map. In addition, the object position estimation device 2 can calculate the total number of objects with the first size or the second size in the target image 70 by adding the total number of objects with the first size and the total number of objects with the second size obtained by the above method.

[0075] (Operation of the object position estimation device 2)

[0076] Reference will be made to Figure 3 to describe in detail the operation of the object position estimation device 2 according to the second exemplary embodiment. Figure 3 is a flowchart showing the operation of the object position estimation device 2.

[0077] As Figure 3 shown, the first feature extraction unit 21 acquires the target image 70 from the image acquisition device 90 (step S10).

[0078] The first feature extraction unit 21 generates a first feature map 80 by performing a convolution process on the target image 70 (step S11). The first feature extraction unit 21 outputs the first feature map 80 to the first position likelihood estimation unit 23 and the second feature extraction unit 22.

[0079] The first position likelihood estimation unit 23 estimates a first likelihood map indicating the positions of objects with the first size by performing a convolution process on the first feature map 80 (step S12). The first position likelihood estimation unit 23 outputs the estimated first likelihood map.

[0080] The second feature extraction unit 22 acquires the first feature map 80 from the first feature extraction unit 21 and generates a second feature map 81 by performing a convolution process on the first feature map 80 (step S13).

[0081] The second position likelihood estimation unit 24 estimates a second likelihood map indicating the positions of objects with the second size by performing a convolution process on the second feature map 81 (step S14). The second position likelihood estimation unit 24 outputs the estimated second likelihood map.

[0082] The above steps S12, S13, and S14 can be executed sequentially. The order between the processes of steps S12, S13, and S14 can be changed. However, the process in step S14 needs to be executed after the process in step S13.

[0083] Therefore, the operation of the object position estimation device 2 ends.

[0084] The configuration in which the object position estimation device 2 includes two feature extraction units (i.e., the first feature extraction unit 21 and the second feature extraction unit 22) and two likelihood map estimation units (i.e., the first position likelihood estimation unit 23 and the second position likelihood estimation unit 24) has been described above. However, the object position estimation device 2 may include three or more feature extraction units and three or more position likelihood estimation units (Modified Type 1).

[0085] [Modified Type 1]

[0086] Figure 4 The configuration of the object position estimation device 2a according to Modified Type 1 is shown. As Figure 4 shown, the object position estimation device 2a includes n (n is an integer of 3 or more) feature extraction units and n position likelihood estimation units. The first feature map is obtained by the first feature extraction unit performing convolution processing on the target image. The second feature map, the third feature map, …, and the n-th feature map are obtained by the i-th feature extraction unit performing convolution processing on the previous-stage feature map. Here, i is an integer from 2 to n.

[0087] Specifically, the i-th feature extraction unit of the object position estimation device 2a generates the i-th feature map by performing convolution processing on the (i - 1)-th feature map. In Figure 4 the shown Modified Type 1, the network in which the first feature extraction unit to the n-th feature extraction unit are connected can be regarded as an integrated feature extraction unit 10.

[0088] The i-th feature map (i = 1 to n) is input to the i-th position likelihood estimation unit. The i-th position likelihood estimation unit estimates the position of an object having the i-th size by performing convolution processing on the i-th feature map. Then, the i-th position likelihood estimation unit estimates and outputs the i-th likelihood map indicating the position of the object having the i-th size. In Figure 4 the shown Modified Type 1, all the feature extraction units and all the likelihood estimation units can be used as an integrated neural network.

[0089] According to the configuration of Modified Type 1, three or more likelihood maps indicating the positions of objects having three or more different sizes from each other can be estimated and output based on the target image. That is, the object position estimation device 2a according to Modified Type 1 can estimate the positions of objects having three or more different sizes from each other.

[0090] [Modified Type 2]

[0091] In Modification 2, the first position likelihood estimation unit 23 and the second position likelihood estimation unit 24 estimate the position of an object for each attribute of the pre-classified object. Then, the first position likelihood estimation unit 23 and the second position likelihood estimation unit 24 estimate the first likelihood map / second likelihood map for each attribute of the object, and output the estimated first likelihood map / second likelihood map.

[0092] For example, if the object is a person or a part of a person, the attributes may be related to the person himself, such as the age of the person, the gender of the person, the orientation of the face, the speed at which the person moves, or the affiliation of the person (social person, student, family member, etc.). Alternatively, the attributes may be related to the group formed by the objects, such as the queue or stay of a crowd including people, or the state of a crowd including people (e.g., panic).

[0093] In one example, the attributes of a person (object) are classified into two categories: children and adults. In this case, the first position likelihood estimation unit 23 estimates the positions of children and adults of a first size in the target image 70. On the other hand, the second position likelihood estimation unit 24 estimates the positions of children and adults of a second size in the target image 70.

[0094] The first position likelihood estimation unit 23 and the second position likelihood estimation unit 24 may be configured as a neural network that outputs the positions of children and adults to each channel. In this case, the first position likelihood estimation unit 23 estimates the positions of children of a first size and adults of a first size in the target image 70, and outputs the estimated positions as channels. The second position likelihood estimation unit 24 estimates the positions of children of a second size and adults of a second size in the target image 70, and outputs the estimated positions as channels.

[0095] According to the second modification, the first position likelihood estimation unit 23 and the second position likelihood estimation unit 24 set the attributes of the object (in the above example, children and adults) as channels of the neural network, and estimate the position of the object having a size determined by each position likelihood estimation unit as a likelihood map for each attribute. As a result, the first position likelihood estimation unit 23 and the second position likelihood estimation unit 24 can estimate the position of the object for each size of the object and for each attribute.

[0096] (Effect of this exemplary embodiment)

[0097] According to the configuration of the present exemplary embodiment, the first feature extraction unit 21 of the feature extraction unit 10 generates a first feature map indicating the features of an object by performing a convolution process on the target image 70. The second feature extraction unit 22 of the feature extraction unit 10 further performs a convolution process on the first feature map to generate a second feature map indicating the features of the object. The first position likelihood estimation unit 23 of the likelihood map estimation unit 20 uses the first feature map to estimate a first likelihood map, which indicates the probability of the existence of an object having a first size at each position of the image. The second position likelihood estimation unit 24 of the likelihood map estimation unit 20 uses the second feature map to estimate a second likelihood map, which indicates the probability of the existence of an object having a second size greater than the first size at each position of the image.

[0098] As described above, the object position estimation device 1 estimates the positions of the objects having the first size and the objects having the second size in the target image 70 using the first feature map and the second feature map, respectively. Therefore, even if the objects overlap each other in the image, the object position estimation device 1 can robustly and highly accurately estimate the positions of the objects.

[0099] According to the configuration of the present exemplary embodiment, when scanning the target image 70, it is not necessary to change the size and position of a partial region where an object is detected as in the related art. Therefore, the object position estimation device 2 can accurately estimate the position of the object without depending on the arrangement of the partial region.

[0100] Furthermore, according to the configuration of the present exemplary embodiment, the first likelihood map / second likelihood map is normalized such that the sum of the overall likelihoods of the first likelihood map / second likelihood map is equal to each total number of the objects having the first size / second size in the target image 70. Therefore, the object position estimation device 2 can obtain the total number of the objects having the first size and the total number of the objects having the second size in the target image 70, and the total number of the objects included in the image 70, from the sum of the likelihoods in the entire first likelihood map and the sum of the likelihoods in the entire second likelihood map.

[0101] [Exemplary Embodiment 3]

[0102] Reference will be made to Figure 5 to describe Exemplary Embodiment 3.

[0103] (Object Position Estimation Device 3)

[0104] Figure 5 is a block diagram showing the configuration of an object position estimation device 3 according to the present Exemplary Embodiment 3. As Figure 5As shown, the object position estimation device 3 includes a first feature extraction unit 21, a second feature extraction unit 22, a first position likelihood estimation unit 23, and a second position likelihood estimation unit 24. In addition, the object position estimation device 3 further includes a first counting unit 25 and a second counting unit 26. Similar to the object position estimation device 2a of the modified example according to Example Embodiment 2, the object position estimation device 3 of the modified example according to this Example Embodiment 3 may include three or more feature extraction units and three or more position likelihood estimation units. In that case, the same number of counting units as the number of feature extraction units and position likelihood estimations are added. The first counting unit 25 and the second counting unit 26 are examples of the first counting component and the second counting component.

[0105] The first feature extraction unit 21 generates a first feature map 80 based on the target image 70, and the second feature extraction unit 22 generates a second feature map 81 based on the first feature map 80 generated by the first feature extraction unit 21.

[0106] Alternatively, the second feature extraction unit 22 may generate the second feature map from the target image 70 itself. In this case, the second feature extraction unit 22 acquires the target image 70 instead of the first feature map 80. The second feature extraction unit 22 generates the second feature map 81 by performing a convolution process on the target image 70 itself.

[0107] The first counting unit 25 acquires the first feature map 80 from the first feature extraction unit 21 and uses the first feature map 80 to calculate the total number of objects having a first size in the target image 70. Specifically, the first counting unit 25 is trained so as to be able to identify the features of the objects having the first size. The trained first counting unit 25 detects each object having the first size in the target image 70 and counts the objects to calculate the total number of objects having the first size.

[0108] The second counting unit 26 acquires the second feature map 81 from the second feature extraction unit 22 and uses the second feature map 81 to calculate the total number of objects having a second size in the target image 70. Specifically, the second counting unit 26 is trained so as to be able to identify the features of the objects having the second size. The trained second counting unit 26 detects each object having the second size in the target image 70 and counts the objects to calculate the total number of objects having the second size. For example, the first counting unit 25 / second counting unit 26 is a convolutional neural network having trained parameters. Then, the first feature extraction unit 21, the second feature extraction unit 22, the first position likelihood estimation unit 23, the second position likelihood estimation unit 24, the first counting unit 25, and the second counting unit 26 may be configured as one neural network. Examples of the training method of the first counting unit 25 and the second counting unit 26 will be described in the following example embodiments.

[0109] (Effect of this exemplary embodiment)

[0110] According to the configuration of this exemplary embodiment, the first feature extraction unit 21 generates a first feature map 80 indicating the features of an object by performing convolutional processing on the target image 70. The second feature extraction unit 22 further performs convolutional processing on the first feature map 80 to generate a second feature map 81 indicating the object features. The first position likelihood estimation unit 23 uses the first feature map 80 to estimate a first likelihood map, which indicates the probability of the presence of an object having a first size at each position of the target image 70. Using the second feature map 81, the second position likelihood estimation unit 24 estimates a second likelihood map, which indicates the probability of the presence of an object having a second size greater than the first size at each position of the target image 70.

[0111] As described above, since the object position estimation device 3 uses the first feature map 80 and the second feature map 81 to estimate the positions of the object having the first size and the object having the second size, even if the objects overlap each other in the target image 70, the positions of each object can be robustly and highly accurately estimated.

[0112] In addition, according to the configuration of this exemplary embodiment, the first counting unit 25 counts the objects having the first size in the target image 70 using the first feature map 80. The second counting unit 26 counts the objects having the second size in the target image 70 using the second feature map 81. As a result, the object position estimation device 3 can more accurately estimate the total number of objects having the first size / objects having the second size included in the target image 70.

[0113] [Exemplary Embodiment 4]

[0114] will be described with reference to Figure 6 Exemplary Embodiment 4 will be described.

[0115] (Object position estimation device 4)

[0116] Figure 6 is a block diagram of the configuration of the object position estimation device 4 according to Exemplary Embodiment 4 of the present invention. As Figure 6As shown, the object position estimation device 4 includes a first feature extraction unit 21, a second feature extraction unit 22, a first position likelihood estimation unit 23, and a second position likelihood estimation unit 24. In addition, the object position estimation device 4 further includes a first position specifying unit 27 and a second position specifying unit 28. Similar to the object position estimation device 2a of the modified example according to Example Embodiment 2, the object position estimation device 4 of the modified example according to the present Example Embodiment 4 may include three or more feature extraction units and three or more position likelihood estimation units. In that case, as many position specifying units as the number of feature extraction units and position likelihood estimations are added. The first position specifying unit 27 and the second position specifying unit 28 are examples of the first position specifying component and the second position specifying component.

[0117] The first position specifying unit 27 specifies the position of an object having a first size in the target image 70 according to a first likelihood map indicating the position of an object having a first size, which is obtained from the first position likelihood estimation unit 23.

[0118] Specifically, the first position specifying unit 27 extracts the coordinates indicating the local maximum of the likelihood from the first likelihood map. After obtaining the coordinates indicating the local maximum of the likelihood from the first likelihood map, the first position specifying unit 27 may integrate a plurality of coordinates indicating the local maximum of the likelihood into one based on the distance between the coordinates indicating the local maximum of the likelihood or the Mahalanobis distance in which the spread of the likelihood around the coordinates indicating the local maximum of the likelihood is a variance value.

[0119] For example, in the case where the Mahalanobis distance between the coordinates indicating the local maximum of the likelihood is less than a threshold, the first position specifying unit 27 integrates these local maxima. In this case, the first position specifying unit 27 may set the average value of the plurality of local maxima as the integrated local maximum. Alternatively, the first position specifying unit 27 may set the intermediate position of the coordinates of the plurality of local maxima indicating the local maximum as the coordinates of the integrated local maximum.

[0120] Thereafter, the first position specifying unit 27 calculates the total number of objects having a first size in the target image 70 (hereinafter, referred to as the first number of objects) by summing all the likelihoods in the first likelihood map.

[0121] When the first quantity of objects in the target image 70 is not zero, the first position specifying unit 27 further extracts, from among the coordinates indicating the local maxima of likelihoods in the first likelihood map, the same number of coordinates as the first quantity of objects in the target image 70 in descending order of likelihood. As a result, even when a large number of local maxima caused by noise appear in the first likelihood map, the first position specifying unit 27 can exclude the local maxima unrelated to the objects having the first size. When one or more of the coordinates thus extracted correspond to the positions of the objects having the first size, the first position specifying unit 27 generates a first object position map. The first position specifying unit 27 may output the coordinates themselves instead of the object position map. The first object position map indicates the positions where the objects having the first size exist in the target image 70.

[0122] The first position specifying unit 27 may also extract, from among the coordinates indicating the local maxima of likelihoods extracted from the first likelihood map, the coordinates having a likelihood equal to or greater than a predetermined value. As a result, the first position specifying unit 27 can exclude the local maxima unrelated to the objects having the first size. The first position specifying unit 27 specifies that an object having the first size exists at the position in the target image 70 associated with the coordinates thus extracted.

[0123] Specifically, the second position specifying unit 28 uses the second likelihood map to specify the positions of the objects having the second size in the target image 70. For example, the second position specifying unit 28 extracts the coordinates indicating the local maxima of likelihoods from the second likelihood map. After obtaining the coordinates indicating the local maxima of likelihoods from the second likelihood map, the second position specifying unit 28 may integrate a plurality of the coordinates indicating the local maxima of likelihoods into one based on the distance between the coordinates indicating the local maxima of likelihoods or the Mahalanobis distance whose expansion of likelihoods around the coordinates indicating the local maxima of likelihoods is a variance value.

[0124] For example, when the Mahalanobis distance between the coordinates indicating the local maxima of likelihoods is less than a threshold value, the second position specifying unit 28 integrates these local maxima. In this case, the second position specifying unit 28 may set the average value of the plurality of local maxima as the integrated local maximum. Alternatively, the second position specifying unit 28 may set the intermediate position of the plurality of coordinates indicating the local maxima as the coordinates of the integrated local maximum.

[0125] Thereafter, the second position specifying unit 28 calculates the total number of the objects having the second size in the target image 70 (hereinafter referred to as the second quantity of objects) by summing all the likelihoods in the second likelihood map.

[0126] When the second quantity of the objects in the target image 70 is not 0, the second position specifying unit 28 further extracts, from among the coordinates indicating the local maxima of likelihood in the second likelihood map, the same number of coordinates as the second quantity of the objects in the target image 70 in descending order of likelihood. In a case where one or more of the coordinates thus extracted correspond to the positions of the objects of the second size, the second position specifying unit 28 generates a second object position map. The second position specifying unit 28 may output the coordinates themselves instead of the object position map. The second object position map indicates the positions where the objects of the second size exist in the target image 70.

[0127] The second position specifying unit 28 may also extract, from among the coordinates indicating the local maxima of likelihood extracted from the second likelihood map, the coordinates having a likelihood equal to or greater than a predetermined value. As a result, the second position specifying unit 28 can exclude the local maxima unrelated to the objects of the second size. The second position specifying unit 28 specifies that the objects of the second size exist at the positions in the target image 70 associated with the coordinates extracted in this way.

[0128] The first position specifying unit 27 / second position specifying unit 28 may perform image processing such as blurring on the first likelihood map / second likelihood map as preprocessing for generating the first object position map / second object position map. As a result, noise can be removed from the first likelihood map / second likelihood map. As postprocessing for generating the first object position map / second object position map, the first position specifying unit 27 / second position specifying unit 28 may use, for example, the distance between the coordinates indicating the positions of the objects of the first size / second size, or the Mahalanobis distance having the spread of the likelihood around the coordinates indicating the positions of the objects of the first size / second size as a variance value, to integrate the coordinates indicating the positions of the objects of the first size / second size.

[0129] The first position specifying unit 27 / second position specifying unit 28 may output the coordinates indicating the positions of the objects of the first size / second size estimated as described above by any method. For example, the first position specifying unit 27 / second position specifying unit 28 may cause a display device to display a graph representing the coordinates indicating the object positions, or may store the data of the coordinates indicating the object positions in a storage device (not shown).

[0130] (Effect of the present exemplary embodiment)

[0131] According to the configuration of the present exemplary embodiment, the first feature extraction unit 21 generates a first feature map 80 indicating the features of an object by performing a convolution process on the target image 70. The second feature extraction unit 22 further performs a convolution process on the first feature map 80 to generate a second feature map 81 indicating the object features. The first position likelihood estimation unit 23 uses the first feature map 80 to estimate a first likelihood map, which indicates the probability of the presence of an object having a first size at each position of the target image 70. Using the second feature map 81, the second position likelihood estimation unit 24 estimates a second likelihood map, which indicates the probability of the presence of an object having a second size greater than the first size at each position of the target image 70.

[0132] As described above, since the object position estimation device 4 uses the first feature map 80 and the second feature map 81 to estimate the positions of the object having the first size and the object having the second size, even if the objects overlap each other in the target image 70, the positions of each object can be robustly and highly accurately estimated.

[0133] According to the configuration of the present exemplary embodiment, the first likelihood map / second likelihood map is converted into a first object position map / second object position map indicating the determined object positions. Then, as a result of estimating the object positions, the first object position map / second object position map or information based on the first object position map / second object position map is output. As a result, the object position estimation device 4 can provide information indicating the estimation result of the object positions in a form that can be easily processed by another device or another application.

[0134] [Exemplary Embodiment 5]

[0135] Reference will be made to Figure 7 to describe Exemplary Embodiment 5.

[0136] (Object Position Estimation Device 5)

[0137] Figure 7 is a block diagram showing the configuration of an object position estimation device 5 according to the present Exemplary Embodiment 5. As Figure 7 shown, similar to Exemplary Embodiment 3, the object position estimation device 5 includes a first feature extraction unit 21, a second feature extraction unit 22, a first position likelihood estimation unit 23, a second position likelihood estimation unit 24, a first counting unit 25, and a second counting unit 16. In addition, the object position estimation device 5 further includes a first position specifying unit 29 and a second position specifying unit 30. The object position estimation device 5 may include three or more feature extraction units, three or more position likelihood estimation units, and three or more counting units. In that case, the same number of position specifying units as the number of feature extraction units, position likelihood estimations, and counting units are added.

[0138] The first position specifying unit 29 acquires a first likelihood map indicating the probability of the presence of an object having a first size from the first position likelihood estimating unit 23. The first position specifying unit 29 acquires a first number of objects (which is the total number of objects having the first size) from the first counting unit 25. The first position specifying unit 29 specifies coordinates indicating a local maximum of the likelihood from the first likelihood map. The first position specifying unit 29 extracts the same number of coordinates as the total number of objects indicated by the first number of objects from among the coordinates indicating the local maximum of the likelihood in the first likelihood map in descending order of likelihood. Then, the first position specifying unit 29 generates a first object position map indicating the positions of the objects having the first size.

[0139] The second position specifying unit 30 acquires a second likelihood map indicating the probability of the presence of an object having a second size from the second position likelihood estimating unit 24. The second position specifying unit 30 acquires a second number of objects (which is the total number of objects having the second size) from the second counting unit 26. The second position specifying unit 30 specifies coordinates indicating a local maximum of the likelihood from the second likelihood map. The second position specifying unit 30 extracts the same number of coordinates as the total number of objects indicated by the second number of objects from among the coordinates indicating the local maximum of the likelihood in the second likelihood map in descending order of likelihood. Then, in the case where the extracted coordinates correspond to the positions of the objects having the second size, the second position specifying unit 30 generates a second object position map.

[0140] Alternatively, the first position specifying unit 29 and the second position specifying unit 30 may also have the functions of the first position specifying unit 27 and the second position specifying unit 28 described in Example Embodiment 4.

[0141] Specifically, the first likelihood map / second likelihood map may include noise. Therefore, as preprocessing for generating the first object position map / second object position map, the first position specifying unit 29 / second position specifying unit 30 may perform image processing such as blurring processing on each of the first likelihood map / second likelihood map. As a result, the noise included in the first likelihood map / second likelihood map can be made less obvious.

[0142] As postprocessing, the first position specifying unit 29 / second position specifying unit 30 may acquire coordinates indicating a local maximum of the likelihood from the first object position map / second object position map, and then integrate a plurality of coordinates indicating the local maximum of the likelihood into one based on the distance between the coordinates indicating the local maximum of the likelihood, or the Mahalanobis distance having the spread of the likelihood around the coordinates indicating the local maximum of the likelihood as a variance value.

[0143] For example, when the Mahalanobis distance between the coordinates indicating the local maxima of likelihood is less than a threshold value, the first position specifying unit 29 / second position specifying unit 30 integrates these local maxima. In this case, the first position specifying unit 29 / second position specifying unit 30 may set the average value of the multiple local maxima as the integrated local maximum. Alternatively, the first position specifying unit 29 / second position specifying unit 30 may set the intermediate position of the multiple coordinates indicating the local maxima as the coordinates of the integrated local maximum.

[0144] The first position specifying unit 29 / second position specifying unit 30 may output the first object position map / second object position map, or information based on the first object position map or second object position map, by any method. For example, the first position specifying unit 29 / second position specifying unit 30 controls a display device to display the first object position map / second object position map, or information based on the first object position map / second object position map, on the display device. Alternatively, the first position specifying unit 29 / second position specifying unit 30 may store the first object position map / second object position map in a storage device accessible from the object position estimation device 5. In addition, the first position specifying unit 29 / second position specifying unit 30 may send the first object position map / second object position map, or information based on the first object position map / second object position map, to another device accessible from the object position estimation device 5.

[0145] (Effect of this exemplary embodiment)

[0146] According to the configuration of this exemplary embodiment, the first feature extraction unit 21 generates a first feature map 80 indicating the features of an object by performing a convolution process on the target image 70. The second feature extraction unit 22 further performs a convolution process on the first feature map 80 to generate a second feature map 81 indicating the object features. The first position likelihood estimation unit 23 uses the first feature map 80 to estimate a first likelihood map that indicates the probability of the presence of an object having a first size at each position of the target image 70. Using the second feature map 81, the second position likelihood estimation unit 24 estimates a second likelihood map that indicates the probability of the presence of an object having a second size greater than the first size at each position of the target image 70.

[0147] As described above, since the object position estimation device 5 uses the first feature map 80 and the second feature map 81 to estimate the positions of the objects having the first size / second size, even if these objects overlap each other in the target image 70, the positions of each object can be estimated robustly and with high accuracy.

[0148] According to the configuration of the present exemplary embodiment, the first position specifying unit 29 / second position specifying unit 30 converts the first likelihood map / second likelihood map into a first object position map / second object position map indicating the determined object position. Then, as a result of estimating the object position, the first object position map / second object position map or information based on the first object position map / second object position map is output. As a result, the object position estimation device 5 can provide information indicating the estimation result of the object position in a form that can be easily processed by another device or another application.

[0149] In addition, the first position specifying unit 29 / second position specifying unit 30 obtains the same number of coordinates as the total number of objects having the first size / second size counted by the first counting unit 25 and the second counting unit 26 from among the coordinates of the local maxima of the likelihood indicated in the likelihood map in descending order of likelihood. Therefore, even when a large number of local maxima of likelihood caused by noise appear on the first likelihood map / second likelihood map, the object position estimation device 5 can accurately obtain the coordinates of the objects having the first size / second size that appear in the target image 70.

[0150] [Exemplary Embodiment 6]

[0151] Reference will be made to Figure 8 and Figure 9 to describe Exemplary Embodiment 6.

[0152] (Object Position Estimation Device 6)

[0153] Figure 8 is a block diagram showing the configuration of an object position estimation device 6 according to the present Exemplary Embodiment 6. The object position estimation device 6 has functions equivalent to those of the object position estimation device 2 according to Exemplary Embodiment 2, except for the points described below.

[0154] As Figure 8 shown, the object position estimation device 6 according to the present Exemplary Embodiment 6 includes a first feature extraction unit 21, a second feature extraction unit 22, a first position likelihood estimation unit 23, and a second position likelihood estimation unit 24. The object position estimation device 6 further includes a training unit 41. The training unit 41 is an example of a training component.

[0155] In a modified example of the present Exemplary Embodiment 6, the object position estimation device 6 may include three or more feature extraction units and three or more position likelihood estimation units. For example, the object position estimation device 6 is provided with n (>2) feature extraction units and n position likelihood estimation units. In this case, the training data (i.e., teacher data) includes training images, object information, and n correct likelihood maps from the first correct likelihood map to the nth correct likelihood map. The n correct likelihood maps from the first correct likelihood map to the nth correct likelihood map may be referred to as correct values.

[0156] (Learning unit 41)

[0157] The training unit 41 trains each unit (except the training unit 41 itself) of the object position estimation device 6 by using pre-prepared training data (i.e., teacher data). The training data includes training images, object information, a first correct likelihood map, and a second correct likelihood map.

[0158] The first correct likelihood map indicates the probability of the position of an object of a first size in the training image and is determined based on the object region. The second correct likelihood map indicates the probability of the position of an object of a second size in the training image and is determined based on the object region. The method for generating the first correct likelihood map and the second correct likelihood map is not limited. For example, an operator can visually observe the object region in the training image displayed on the display device and manually generate the first correct likelihood map and the second correct likelihood map. The object position estimation device 6 may also include a training data generation unit 42 as shown in the object position estimation device 6a described later, and the training data generation unit 42 can generate the first correct likelihood map and the second correct likelihood map.

[0159] When the training data is generated by another device different from the object position estimation device 6, the object position estimation device 6 acquires the training data from that other device. For example, the training data is pre-stored in a storage device accessible from the object position estimation device 6. In this case, the object position estimation device 6 acquires the training data from the storage device. Alternatively, the object position estimation device 6 can acquire the training data generated by the training data generation unit 42 (a modified type described later).

[0160] The object position estimation device 6 does not learn the features of the shape of the object, but learns the positions of the objects in the training image while taking into account the overlap between the objects. As a result, the object position estimation device 6 can learn the overlap between the objects in the training image as it is.

[0161] The training unit 41 inputs the training image into the first feature extraction unit 21. The first feature extraction unit 21 generates a first feature map 80 from the training image. Then, the first position likelihood estimation unit 23 outputs a first likelihood map indicating the position of an object of a first size based on the first feature map 80. The first position likelihood estimation unit 23 outputs the first likelihood map to the training unit 41.

[0162] The first feature extraction unit 21 inputs the first feature map 80 into the second feature extraction unit 22. The second feature extraction unit 22 generates a second feature map 81 from the first feature map 80.

[0163] Alternatively, the second feature extraction unit 22 may generate a second feature map from the training image itself. In this case, the second feature extraction unit 22 acquires the training image instead of the first feature map 80. The second feature extraction unit 22 generates a second feature map 81 by performing more convolutional processing on the training image itself than the first feature extraction unit 21.

[0164] The second position likelihood estimation unit 24 outputs a second likelihood map indicating the positions of objects having the second size in the training image based on the second feature map 81. The second position likelihood estimation unit 24 outputs the second likelihood map to the training unit 41.

[0165] The training unit 41 calculates an error between each output (the first likelihood map, the second likelihood map) from the first position likelihood estimation unit 23 and the second position likelihood estimation unit 24 and the correct values (the first correct likelihood map, the second correct likelihood map) included in the training data as a first loss. For example, the training unit 41 calculates the mean square error between the first likelihood map / second likelihood map and the first correct likelihood map / second correct likelihood map. Then, the training unit 41 sets the calculated mean square error between the maps as the first loss. The training unit 41 trains each unit (except the training unit 41 itself) of the object position estimation device 6 to reduce the calculated first loss.

[0166] The term "training" used herein means updating the parameters of each unit of the object position estimation device 6. For example, the training unit 41 may perform a training process using a known technique such as backpropagation. Specifically, the training unit 41 calculates the first loss using a preset calculation formula (e.g., a loss function) of the first loss and trains each unit of the object position estimation device 6 to reduce the first loss. Alternatively, the training unit 41 acquires the calculation formula of the first loss stored in an accessible storage device, calculates the first loss, and trains each unit of the object position estimation device 6 to reduce the first loss.

[0167] In one example, the training unit 41 updates the parameters of each unit (except the training unit 41 itself) of the object position estimation device 6 based on the information (i.e., the first likelihood map / second likelihood map) fed back from the outputs of the first position likelihood estimation unit 23 / second position likelihood estimation unit 24 to the training unit 41.

[0168] After updating the parameters of each unit (except for the training unit 41) of the object position estimation device 6, each unit of the object position estimation device 6 uses another training data to estimate and output the first likelihood map / second likelihood map. The first likelihood map / second likelihood map is fed back from the output of the first position likelihood estimation unit 23 / second position likelihood estimation unit 24 to the training unit 41. The training unit 41 updates the parameters of each unit (except for the training unit 41) of the object position estimation device 6 again based on the feedback information (i.e., the first likelihood map / second likelihood map).

[0169] The training unit 41 can repeatedly perform the training of each unit of the object position estimation device 6 by the above method until the magnitude of the first loss becomes equal to or less than a predetermined threshold. However, the condition for the training unit 41 to end the training of each unit (except for the training unit 41) of the object position estimation device 6 is not limited. In this way, the training unit 41 repeatedly trains the parameters of each unit of the object position estimation device 6 to reduce the first loss. As a result, since the estimation of the first likelihood map and the estimation of the second likelihood map are trained simultaneously by the first feature extraction unit 21, the object position estimation device 6 can estimate the position of the object more accurately, and the training speed can be improved.

[0170] (Operation of the object position estimation device 6)

[0171] The operation of the object position estimation device 6 according to Example Embodiment 6 will be described with reference to Figure 9 FIG. Figure 9 FIG. is a flowchart showing the operation flow of the object position estimation device 6. Here, the case where the object position estimation device 6 performs training using a single training data will be described. When there are multiple pieces of training data, the object position estimation device 6 repeats Figure 9 the processing from step S20 to S23 shown in FIG., and performs the processing for each piece of training data.

[0172] As Figure 9 shown, first, the training unit 41 acquires the training data (S20). The training unit 41 inputs the training image included in the training data into the first feature extraction unit 21 (S21). The training unit 41 calculates the first loss indicating the error between the output of each position likelihood estimation unit and the correct value (S22), and performs the training (parameter update) of each unit of the object position estimation device 6 to reduce the calculated first loss (S23).

[0173] Thus, the operation of the object position estimation device 6 ends.

[0174] [Modified Type 1]

[0175] In Modification 1, in addition to the position and size of the object, the object information of the training data indicates the attributes of the object. The training unit 41 prepares, for each attribute of the object, a first correct likelihood map as the probability of the position of the object indicating the first size and a second correct likelihood map as the probability of the position of the object indicating the second size as training data. Then, the training unit 41 uses the training image, the first correct likelihood map as the probability of the position of the object indicating the first size for each attribute, and the second correct likelihood map as the probability of the position of the object indicating the second size for each attribute to perform training of each unit of the object position estimation device 6 by the above method ( Figure 9 )

[0176] According to the configuration of Modification 1, each unit of the object position estimation device 6 is trained using the first correct likelihood map and the second correct likelihood map for each attribute. As a result, the object position estimation device 6 can estimate the position of the object for each attribute of the object. For example, the object position estimation device 6 can estimate the position of an adult (an example of an attribute of the object) and can also estimate the position of a child (another example of the position of the object) separately.

[0177] [Modification 2]

[0178] When the total number of objects in the training image is small or the deviation in the arrangement of the objects is large, training may not be performed correctly. Specifically, in the first correct likelihood map or the second correct likelihood map as training data, there may be many coordinates with a likelihood of 0.

[0179] In the above training for minimizing the first loss, the training unit 41 of this Modification 2 trains each unit of the object position estimation device 6 so as to minimize the error in some coordinates without using the error in all coordinates in the first correct likelihood map / second correct likelihood map as training data and the first likelihood map / second likelihood map as the estimation result. Specifically, the training unit 41 of this Modification 2 selects some coordinates on the first correct likelihood map / second correct likelihood map as training data such that the number of coordinates with a likelihood of 0 and the number of other coordinates in the first correct likelihood map / second correct likelihood map become a predetermined ratio. Then, based on the coordinates on the selected first correct likelihood map / second correct likelihood map, the coordinates of the first likelihood map / second likelihood map as the estimation result are also selected. For example, the training unit 41 selects the same number of coordinates with a likelihood of 0 and other coordinates from the first correct likelihood map / second correct likelihood map, and also selects the coordinates of the first likelihood map / second likelihood map based on the selected coordinates on the first correct likelihood map / second correct likelihood map. The training unit 41 updates the parameters of each unit of the object position estimation device 6 so as to minimize the first error in the selected coordinates.

[0180] (Object position estimation device 6a)

[0181] Figure 10 is a block diagram showing the configuration of an object position estimation device 6a according to a modified example of Example Embodiment 6. The object position estimation device 6a according to this modified example includes a first feature extraction unit 21, a second feature extraction unit 22, a first position likelihood estimation unit 23, and a second position likelihood estimation unit 24. The object position estimation device 6a further includes a training unit 41 and a training data generation unit 42. The training data generation unit 42 is an example of a training data generation component. The object position estimation device 6a is different from the above-described object position estimation device 6 in that the object position estimation device 6a further includes a training data generation unit 42.

[0182] (Learning data generation unit 42)

[0183] The training data generation unit 42 generates training data (teacher data) for the training unit 41 to perform training.

[0184] Reference will be made to Figure 11 to describe the operation of the training data generation unit 42 according to this modified example. Figure 11 shows a processing flow executed by the training data generation unit to create a first correct likelihood map and a second correct likelihood map as training data.

[0185] The training data generation unit 42 acquires a training image. For example, the training image and object information are input to the object position estimation device 6a by an operator. Here, the training image includes an object of a first size / a second size (the head as the "target object" in Figure 11 as the target of position estimation performed by the object position estimation device 6a). The object region in the training image is specified by the object information associated with the training image.

[0186] The object region corresponds to the region occupied by the object in the training image. For example, the object region is a region in the training image surrounded by a rectangle or other two-dimensional shape circumscribing the object. For example, the object information specifies the coordinates of the upper left corner and the lower right corner of the object region (e.g., the circumscribing rectangle of the object) in the training image.

[0187] The training data generation unit 42 specifies the position and size of the object in the training image by using the object information associated with the training image. Then, according to the process described below, the training data generation unit 42 generates a first correct likelihood map and a second correct likelihood map.

[0188] As Figure 11As shown, the training data generation unit 42 first detects each of the objects having the first size and the objects having the second size based on the object information associated with the training image. The training data generation unit 42 specifies the positions of the objects having the first size / the objects having the second size in the training image.

[0189] Next, the training data generation unit 42 prepares an initial first correct likelihood map / an initial second correct likelihood map in which the likelihood of all coordinates is 0, and generates a normal distribution of likelihood centered on the center or centroid of the object region of the object having the first size / the object having the second size on the first correct likelihood map / the second correct likelihood map. When generating the normal distribution of likelihood, the training data generation unit 42 generates a normal distribution of likelihood for the object having the first size on the first correct likelihood map, and generates a normal distribution of likelihood for the object having the second size on the second correct likelihood map.

[0190] In addition, the training data generation unit 42 defines the spread of the normal distribution on the first correct likelihood map / the second correct likelihood map by parameters. For example, the parameters can be the parameters indicating the center (mean) and variance of the function of the normal distribution. In this case, the center of the function indicating the normal distribution can be the value indicating the position of the object (e.g., the center or centroid of the object region), and the variance of the function indicating the normal distribution can be the value related to the size of the object region. The form of the function indicating the normal distribution can be set such that the center value of the function indicating the normal distribution becomes 1.

[0191] As described above, the training data generation unit 42 generates a first correct likelihood map / a second correct likelihood map indicating the probability of the presence of an object having the first size / an object having the second size at each position of the training image. In the first correct likelihood map / the second correct likelihood map, the object region of the object having the first size / the object having the second size is related to the spread of the normal distribution of likelihood.

[0192] In the case where the normal distributions of multiple likelihoods overlap in a certain part of the first correct likelihood map and the second correct likelihood map, the training data generation unit 42 can set the maximum value of the likelihood at the same coordinates in this part as the likelihood at this coordinate. Alternatively, the training data generation unit 42 can set the average value of the likelihood at the coordinates of the part where multiple normal distributions overlap as the likelihood at this coordinate. However, the training data generation unit 42 can calculate the likelihood in the part where multiple normal distributions overlap on the first correct likelihood map and the second correct likelihood map by other methods.

[0193] The training data generation unit 42 calculates the total number of objects (the first number of objects) having a first size in the training image based on the object information. The training data generation unit 42 normalizes the likelihood of the first correct likelihood map so that the sum of the likelihoods in the first correct likelihood map is consistent with the first number of objects in the training image. In Figure 11 the normalized first correct likelihood map is omitted. Alternatively, the training data generation unit 42 may count the first number of objects by using the sum of the ratios of the object regions included in the training image.

[0194] The likelihood at each coordinate of the normalized first correct likelihood map represents the probability that an object having the first size exists at the position indicated by the coordinate. The sum of the likelihoods of the entire normalized first correct likelihood map is equal to the total number of objects having the first size included in the training image. That is, the sum of the likelihoods of the entire first correct likelihood map also has the meaning of the total number of objects existing in the first correct likelihood map.

[0195] In addition, the training data generation unit 42 makes the size of the normalized first correct likelihood map equal to the size of the first likelihood map that is the output of the first position likelihood estimation unit 23. In other words, the training data generation unit 42 transforms the first correct likelihood map so that each coordinate on the normalized first correct likelihood map is associated with each position in the training image on a one-to-one basis. In the above description, the case where the training data generation unit 42 performs normalization has been described as an example, but the normalization process is not necessary. That is, the training data generation unit 42 may not normalize the first correct likelihood map and the second correct likelihood map.

[0196] The training data generation unit 42 specifies an object having a second size from the training image using the object information. The training data generation unit 42 generates a normal distribution representing the positions of the objects having the specified second size. Then, similar to the process described for the first correct likelihood map, the training data generation unit 42 generates a second correct likelihood map and normalizes the second correct likelihood map. In Figure 11 the normalized second correct likelihood map is omitted.

[0197] In addition, the training data generation unit 42 matches the size of the normalized second correct likelihood map with the size of the second likelihood map. That is, the training data generation unit 42 transforms the second correct likelihood map such that each coordinate on the normalized second correct likelihood map is associated with each position in the training image on a one-to-one basis. The likelihood at each coordinate on the second correct likelihood map indicates the probability that an object having the second size exists at the associated position on the training image. In the above description, the case where the training data generation unit 42 performs normalization has been described as an example, but the normalization process is not essential. That is, the training data generation unit 42 may not normalize the first correct likelihood map and the second correct likelihood map.

[0198] The training data generation unit 42 associates the training image, the object information, and the correct values. The correct values include the first correct likelihood map and the second correct likelihood map.

[0199] (Effect of this exemplary embodiment)

[0200] According to the configuration of this exemplary embodiment, the first feature extraction unit 21 generates a first feature map 80 indicating the features of an object by performing a convolution process on the target image 70. The second feature extraction unit 22 further performs a convolution process on the first feature map 80 to generate a second feature map 81 indicating the object features. The first position likelihood estimation unit 23 uses the first feature map 80 to estimate a first likelihood map that indicates the probability that an object having the first size exists at each position of the target image 70. Using the second feature map 81, the second position likelihood estimation unit 24 estimates a second likelihood map that indicates the probability that an object having a second size greater than the first size exists at each position of the target image 70.

[0201] As described above, since the object position estimation device 6 (6a) uses the first feature map 80 and the second feature map 81 to estimate the positions of the objects having the first size / second size, even if these objects overlap each other in the target image 70, the positions of each object can be estimated robustly and with high accuracy.

[0202] The object position estimation device 6 (6a) uses the first correct likelihood map / second correct likelihood map to learn the positions of the objects having the first size / objects having the second size as the arrangement pattern of the objects including the overlap between the objects. In the first correct likelihood map / second correct likelihood map, the probability that an object having the first size / object having the second size exists at each coordinate of the training image is represented by the likelihood. As a result, even when the objects overlap each other in the target image 70, the object position estimation device 6 (6a) can robustly and accurately estimate the positions of the objects in the target image 70.

[0203] [Exemplary Embodiment 7]

[0204] Example embodiment 7 will be described in detail with reference to Figure 12 and Figure 13 .

[0205] (Object position estimation device 7)

[0206] Figure 12 FIG. is a block diagram showing the configuration of an object position estimation device 7 according to Example embodiment 7 of the present invention. As Figure 12 shown, the object position estimation device 7 includes a first feature extraction unit 21, a second feature extraction unit 22, a first position likelihood estimation unit 23, and a second position likelihood estimation unit 24. The object position estimation device 7 includes a training unit 41. In addition, the object position estimation device 7 further includes a first counting unit 25 and a second counting unit 26. For example, each unit of the object position estimation device 7 can operate independently or integrally in a neural network such as a convolutional neural network.

[0207] (Learning unit 41)

[0208] The training unit 41 trains each unit (except the training unit 41) included in the object position estimation device 7 by using pre-prepared training data (i.e., teacher data).

[0209] In Example embodiment 7 of the present invention, the training data includes training images and object information. The training images include objects for which the position likelihood is to be estimated. The training unit 41 uses the training images to learn the estimation of the likelihood of the position of the object and the total number of objects. The training data further includes the correct value of the first number of objects, the correct value of the second number of objects, the first correct likelihood map, and the second correct likelihood map. Hereinafter, the first correct likelihood map, the second correct likelihood map, the correct value of the first number of objects, and the correct value of the second number of objects may be collectively referred to as correct values. The method of generating the correct values is not limited.

[0210] For example, the operator designates the positions of objects of the first size / objects of the second size in the training images, and gives a normal distribution of likelihood centered on the positions of the objects of the first size / second size on the initial first correct likelihood map / second correct likelihood map where the likelihood of all coordinates is zero. The operator counts each of the objects of the first size and the objects of the second size that appear in the training images, and determines the total number of objects of the first size that appear in the training images as the correct value of the first number of objects, and determines the total number of objects of the second size that appear in the training images as the correct value of the second number of objects.

[0211] The likelihood in each coordinate of the first correct likelihood map indicates the probability that an object of the first size exists at the relevant position in the training image. The likelihood in each coordinate of the second correct likelihood map indicates the probability that an object of the second size exists at the relevant position in the training image.

[0212] The correct value of the first quantity of objects indicates the total number of objects of the first size included in the training image. The correct value of the second quantity of objects indicates the total number of objects of the second size included in the training image. In addition, the object position estimation device 7 may include a training data generation unit 42 shown in the object position estimation device 7a to be described later, and the training data generation unit 42 may generate each correct value.

[0213] The training unit 41 inputs the training image into the first feature extraction unit 21, and calculates the error between the first likelihood map / second likelihood map output from the first position likelihood estimation unit 23 and the second position likelihood estimation unit 24 and the correct value (first correct likelihood map / second correct likelihood map) included in the training data as the first loss. The training unit 41 calculates the error between the number of first objects / second objects output from the first counting unit 25 and the second counting unit 26 when the training image is input into the first feature extraction unit 21 and another correct value (the correct value of the first quantity of objects and the correct value of the second quantity of objects) included in the training data as the second loss.

[0214] The training unit 41 causes each unit of the object position estimation device 7 to learn so as to reduce at least one of the first loss and the second loss.

[0215] Specifically, the training unit 41 updates the parameters of each unit (except the training unit 41) of the object position estimation device 7 based on at least one of the first loss and the second loss. In one example, the training unit 41 causes each unit of the object position estimation device 7 to learn so that the first likelihood map output by the first position likelihood estimation unit 23 matches the first correct likelihood map. At the same time, the training unit 41 causes each unit of the object position estimation device 7 to learn so that the second likelihood map output by the second position likelihood estimation unit 24 matches the second correct likelihood map.

[0216] In addition, the training unit 41 causes each unit of the object position estimation device 7 to learn so that the first quantity of objects counted by the first counting unit 25 matches the correct value of the first quantity of objects. In addition, the training unit 41 causes each unit of the object position estimation device 7 to learn so that the second quantity of objects counted by the second counting unit 26 matches the correct value of the second quantity of objects.

[0217] There may be a situation where the deviation in the arrangement of objects in the training images is relatively large. In such a case, the training unit 41 can cause each unit of the object position estimation device 7 to learn so as to minimize the error in only some coordinates in the first likelihood map / second likelihood map. The example described here is shown in a modified type 2 of the object position estimation device 6.

[0218] (Object position estimation device 7a)

[0219] Figure 13 FIG. is a block diagram showing the configuration of an object position estimation device 7a according to a modified type of Example Embodiment 7 of the present example. The object position estimation device 7a according to this modified type includes a first feature extraction unit 21, a second feature extraction unit 22, a first position likelihood estimation unit 23, a second position likelihood estimation unit 24, a first counting unit 25, a second counting unit 26, and a training unit 41. The object position estimation device 7a further includes a training data generation unit 42. The object position estimation device 7a according to this modified type is different from the object position estimation device 7 in that the object position estimation device 7a further includes a training data generation unit 42.

[0220] As in the above Example Embodiment 6, the training data generation unit 42 generates training data (teacher data) for performing training related to the estimation of the position of an object having a first size / a position of an object having a second size in the target image 70. The training data generated by the training data generation unit 42 includes training images, object information, and correct values.

[0221] The training data generation unit 42 according to this modified type generates training data including the correct value of the first number of objects and the correct value of the second number of objects as the correct value. In this regard, the training data generation unit 42 of the object position estimation device 7a is different from the training data generation unit 42 of the object position estimation device 6a. The training data generation unit 42 of the object position estimation device 7a uses the total number of objects having a first size and the total number of objects having a second size to generate the correct value of the first number of objects and the correct value of the second number of objects, and the total number of objects having a first size and the total number of objects having a second size are obtained in the process of the training data generation unit 42 of the object position estimation device 6a according to the modified type of Example Embodiment 6. The total number of objects having a first size and the total number of objects having a second size are obtained through the counting process for normalizing the first correct likelihood map and the second correct likelihood map, as described for the training data generation unit 42 of the object position estimation device 6a according to the modified type of Example Embodiment 6.

[0222] (Effect of the present example embodiment)

[0223] According to the configuration of the present exemplary embodiment, the object position estimation device 7 according to the seventh exemplary embodiment of the present invention and the modified object position estimation device 7a according to the seventh exemplary embodiment are configured such that a plurality of parts are respectively and simultaneously connected to the subsequent stages of the first feature extraction unit 21 and the second feature extraction unit 22, and during training, the first feature extraction unit 21 and the second feature extraction unit 22 are affected by the plurality of parts to appropriately update the parameters. Further, the first feature extraction unit 21 and the second feature extraction unit 22 serve as a common part of the plurality of parts connected at the subsequent stage, and the first feature extraction unit 21 and the second feature extraction unit 22 are trained simultaneously. As a result, the accuracy of estimating the position of the object and the accuracy of counting the object in the object position estimation devices 7 and 7a can be improved, and the training speed can be improved.

[0224] [Hardware Configuration]

[0225] Figure 14 The hardware configuration of the object position estimation device 1 according to the first exemplary embodiment is shown. Each configuration of the object position estimation device 1 is implemented as a function of a computer 100 reading and executing an object position estimation program 101 (hereinafter simply referred to as the program 101). Referring to Figure 14 , an image acquisition device 90 is connected to the computer 100. A recording medium 102 storing the program 101 readable by the computer 100 is connected to the computer 100.

[0226] The recording medium 102 includes a magnetic disk, a semiconductor memory, etc. For example, the computer 100 reads the program 101 stored in the recording medium 102 at startup. The program 101 controls the operation of the computer 100 to cause the computer 100 to function as each unit in the object position estimation device 1 according to the first exemplary embodiment of the present invention described above.

[0227] Here, the configuration of the object position estimation device 1 according to the first exemplary embodiment implemented by the computer 100 and the program 101 is described. However, the object position estimation devices 2 to 7 (7a) according to the second to seventh exemplary embodiments can also be implemented by the computer 100 and the program 101.

[0228] [Supplementary Notes]

[0229] The exemplary embodiments of the present invention have been described above with reference to the accompanying drawings. However, these are examples of the present invention, and configurations in which the configurations of the exemplary embodiments are combined or various configurations other than the above can also be adopted. Some or all of the above exemplary embodiments can be described as the following supplementary notes, but are not limited to the following.

[0230] (Supplementary Note 1)

[0231] An object position estimation device, comprising:

[0232] A feature extraction component, including a first feature extraction component and a second feature extraction component, where the first feature extraction component is configured to generate a first feature map by performing convolution processing on a target image, and the second feature extraction component is configured to generate a second feature map by further performing convolution processing on the first feature map; and

[0233] A likelihood map estimation component, including a first position likelihood estimation component and a second position likelihood estimation component. The first position likelihood estimation component is configured to estimate a first likelihood map indicating the probability of the existence of an object with a first size at each position of the target image by using the first feature map, and the second position likelihood estimation component is configured to estimate a second likelihood map indicating the probability of the existence of an object with a second size greater than the first size at each position of the target image by using the second feature map.

[0234] (Supplementary Note 2)

[0235] The object position estimation device according to Supplementary Note 1, wherein

[0236] Each coordinate on the first likelihood map corresponds to a position on the target image, and the likelihood at each coordinate on the first likelihood map indicates the probability of the existence of an object with a first size at the corresponding position on the target image, or indicates the number of objects with a first size that additionally exist on the target image, and

[0237] Each coordinate on the second likelihood map corresponds to a position on the target image, and the likelihood at each coordinate on the second likelihood map indicates the probability of the existence of an object with a second size at the corresponding position on the target image, or indicates the number of objects with a second size that additionally exist on the target image.

[0238] (Supplementary Note 3)

[0239] The object position estimation device according to Supplementary Note 1 or 2, wherein

[0240] The first position likelihood estimation component is configured to: estimate the position of an object with a first size for each attribute of the object with a first size, and

[0241] The second position likelihood estimation component is configured to: estimate the position of an object with a second size for each attribute of the object with a second size.

[0242] (Supplementary Note 4)

[0243] The object position estimation device according to any one of Supplementary Notes 1 to 3 further includes:

[0244] A first counting component, configured to count the total number of objects with a first size in a target image based on a first feature map; and

[0245] A second counting component, configured to count the total number of objects with a second size in the target image based on a second feature map.

[0246] (Supplementary Note 5)

[0247] The object position estimation device according to any one of Supplementary Notes 1 to 4 further includes:

[0248] A first position specifying component, configured to specify the position of an object with a first size in the target image based on the coordinates of the local maximum indicating likelihood in a first likelihood map; and

[0249] A second position specifying component, configured to specify the position of an object with a second size in the target image based on the coordinates of the local maximum indicating likelihood in a second likelihood map.

[0250] (Supplementary Note 6)

[0251] The object position estimation device according to Supplementary Note 5, wherein

[0252] The first position specifying component is configured to:

[0253] Calculate the total number of objects with a first size in the target image according to the sum of the overall likelihoods of the first likelihood map, or the first counting component counts the total number of objects with a first size in the target image,

[0254] Extract the same number of coordinates as the total number of objects with a first size from among the coordinates of the local maximum indicating likelihood in the first likelihood map in descending order of the local maximum of the likelihood, and

[0255] Specify the position of the object with a first size in the target image based on the extracted coordinates of the local maximum indicating likelihood, and

[0256] The second position specifying component is configured to:

[0257] Calculate the total number of objects with a second size in the target image according to the sum of the overall likelihoods of the second likelihood map, or the second counting component counts the total number of objects with a first size in the target image,

[0258] Extract the same number of coordinates as the total number of objects with a second size from among the coordinates of the local maximum indicating likelihood in the second likelihood map in descending order of the local maximum of the likelihood, and

[0259] Specify the position of an object having a second size in the target image based on the coordinates of the local maximum of the extracted indication likelihood.

[0260] (Supplementary Note 7)

[0261] The object position estimation device according to any one of Supplementary Notes 1 to 6 further includes:

[0262] A training component configured to cause each unit of the object position estimation device to perform training in such a way that the error with respect to a previously obtained correct value in the first likelihood map and the second likelihood map output from the first position likelihood estimation component and the second position likelihood estimation component is reduced.

[0263] (Supplementary Note 8)

[0264] The object position estimation device according to Supplementary Note 7 further includes:

[0265] A training data generation component configured to generate training data to be used for training by the training component based on the training image and the object information, where

[0266] The training data includes the training image, the object information, and the correct value,

[0267] The correct value includes the first correct likelihood map and the second correct likelihood map, and

[0268] The first correct likelihood map indicates the position and the extension of the object region of the object having the first size in the training image, and the second correct likelihood map indicates the position and the extension of the object region of the object having the second size in the training image.

[0269] (Supplementary Note 9)

[0270] The object position estimation device according to Supplementary Note 8, wherein

[0271] The training component is configured to calculate a first loss indicating the error between the first likelihood map and the second likelihood map and the correct value by using the first correct likelihood map and the second correct likelihood map included in the training data as the correct value.

[0272] (Supplementary Note 10)

[0273] The object position estimation device according to any one of Supplementary Notes 1 to 9, wherein

[0274] The first size is any size within a first predetermined range from a first minimum size to a first maximum size,

[0275] The second size is any size within a second predetermined range from a second minimum size to a second maximum size, the first predetermined range and the second predetermined range do not overlap, and the second size is greater than the first size.

[0276] (Supplementary Note 11)

[0277] The object position estimation device according to any one of Supplementary Notes 1 to 10, wherein

[0278] The first size and the second size are proportional to the reciprocals of the data sizes of the first feature map and the second feature map.

[0279] (Supplementary Note 12)

[0280] An object position estimation method, comprising:

[0281] Generating a first feature map by performing a convolution process on a target image, and generating a second feature map by further performing a convolution process on the first feature map; and

[0282] Using the first feature map to estimate a first likelihood map indicating the probability of the presence of an object having a first size at each position of the target image, and using the second feature map to estimate a second likelihood map indicating the probability of the presence of an object having a second size greater than the first size at each position of the target image.

[0283] (Supplementary Note 13)

[0284] A non-transitory recording medium for causing a computer to execute:

[0285] Generating a first feature map by performing a convolution process on a target image, and generating a second feature map by further performing a convolution process on the first feature map; and

[0286] Using the first feature map to estimate a first likelihood map indicating the probability of the presence of an object having a first size at each position of the target image, and using the second feature map to estimate a second likelihood map indicating the probability of the presence of an object having a second size greater than the first size at each position of the target image.

[0287] [Industrial Applicability]

[0288] The present invention can be used in a video surveillance system for purposes such as detecting suspicious persons or objects from captured or recorded videos, or detecting suspicious behaviors or states. The present invention can be applied to applications in marketing, such as traffic flow analysis or behavior analysis. Additionally, the present invention can be applied to applications such as user interfaces for estimating the position of an object from a captured or recorded image and inputting the estimated position information in a two-dimensional or three-dimensional space. Furthermore, the present invention can also be applied to a video / video search device or a video search function that uses the estimated result of the position of an object and the position as a trigger key.

[0289] [List of reference numerals]

[0290] 1 Object position estimation device

[0291] 2(2a) Object position estimation device

[0292] 3 Object position estimation device

[0293] 4 Object position estimation device

[0294] 5 Object position estimation device

[0295] 6(6a) Object position estimation device

[0296] 7 Object position estimation device

[0297] 10 Feature extraction unit

[0298] 20 Likelihood map estimation unit

[0299] 21 First feature extraction unit

[0300] 22 Second feature extraction unit

[0301] 23 First position likelihood estimation unit

[0302] 24 Second position likelihood estimation unit

[0303] 25 First counting unit

[0304] 26 Second counting unit

[0305] 27 First position specifying unit

[0306] 28 Second position specifying unit

[0307] 29 First position specifying unit

[0308] 30 Second position specifying unit

[0309] 41 Training unit

[0310] 42 Training data generation unit

[0311] 80 First feature map

[0312] 81 Second feature map

[0313] 90 Image acquisition unit.

Claims

1. An object position estimation device, comprising: A feature extraction component, including a first feature extraction component and a second feature extraction component. The first feature extraction component is configured to generate a first feature map from the target image by performing a convolution process on the target image using a first filter. The second feature extraction component is configured to generate a second feature map from the first feature map by further performing a convolution process on the first feature map using a second filter; And A likelihood map estimation component, including: A first position likelihood estimation component, configured to estimate a first likelihood map by using the first feature map. The first likelihood map indicates the probability that an object with a first size exists at each position of the target image, and A second position likelihood estimation component, configured to estimate a second likelihood map by using the second feature map. The second likelihood map indicates the probability that an object with a second size greater than the first size exists at each position of the target image, where The first size is determined based on a first reference size of the object. The first size is any size within a first predetermined range from a first minimum size to a first maximum size; and the second size is determined based on a second reference size of the object. The second size is any size within a second predetermined range from a second minimum size to a second maximum size. The first predetermined range does not overlap with the second predetermined range, and the second size is greater than the first size.

2. The object position estimation device according to claim 1, wherein, Each coordinate on the first likelihood map corresponds to a position on the target image, and the likelihood at each coordinate on the first likelihood map indicates the probability that an object with the first size exists at the corresponding position on the target image, or indicates the number of objects with the first size that additionally exist on the target image, and Each coordinate on the second likelihood map corresponds to a position on the target image, and the likelihood at each coordinate on the second likelihood map indicates the probability that an object with the second size exists at the corresponding position on the target image, or indicates the number of objects with the second size that additionally exist on the target image.

3. The object position estimation device according to claim 1, wherein, The first position likelihood estimation component is configured to: estimate the position of an object with the first size for each attribute of the object with the first size, and The second position likelihood estimation component is configured to: estimate the position of an object with the second size for each attribute of the object with the second size.

4. The object position estimation device according to claim 1, further comprising: A first counting component, configured to count the total number of objects with the first size in the target image based on the first feature map; And A second counting component, configured to count the total number of objects with the second size in the target image based on the second feature map.

5. The object position estimation device according to any one of claims 1 to 4, further comprising: A first position specifying component, configured to specify the position of an object with the first size in the target image based on the coordinates indicating the local maximum of the likelihood in the first likelihood map; And A second position specifying component configured to specify a position of an object having the second size in the target image based on coordinates of local maxima indicating likelihood in the second likelihood map.

6. The object position estimation device according to claim 5, wherein, The first position specifying component is configured to: calculate a total number of objects having the first size in the target image based on a sum of all likelihoods of the first likelihood map, or count the total number of objects having the first size in the target image, extract, in descending order of the local maxima of the likelihood, the same number of coordinates as the total number of objects having the first size from among the coordinates of the local maxima indicating the likelihood in the first likelihood map, and specify a position of an object having the first size in the target image based on the extracted coordinates of the local maxima indicating the likelihood, and The second position specifying component is configured to: calculate a total number of objects having the second size in the target image based on a sum of all likelihoods of the second likelihood map, or count the total number of objects having the second size in the target image, extract, in descending order of the local maxima of the likelihood, the same number of coordinates as the total number of objects having the second size from among the coordinates of the local maxima indicating the likelihood in the second likelihood map, and specify a position of an object having the second size in the target image based on the extracted coordinates of the local maxima indicating the likelihood.

7. The object position estimation device according to any one of claims 1 to 4, further comprising: A training component configured to cause each unit of the object position estimation device to perform training in such a way that an error with respect to a previously obtained correct value in the first likelihood map and the second likelihood map output from the first position likelihood estimation component and the second position likelihood estimation component is reduced.

8. The object position estimation device according to claim 7, further comprising: A training data generation component configured to generate training data to be used for training by the training component based on a training image and object information, where the training data includes the training image, the object information, and the correct value, the correct value includes a first correct likelihood map and a second correct likelihood map, and the first correct likelihood map indicates a position and an extension of an object region of an object having the first size in the training image, and the second correct likelihood map indicates a position and an extension of an object region of an object having the second size in the training image.

9. The object position estimation device according to claim 8, wherein, The training component is configured to calculate a first loss indicating an error between the first likelihood map and the second likelihood map and the correct value by using the first correct likelihood map and the second correct likelihood map included in the training data as the correct value.

10. The object position estimation device according to any one of claims 1 to 4, wherein, The first size and the second size are proportional to reciprocals of data sizes of the first feature map and the second feature map.

11. An object position estimation method, comprising: Performing a convolution process on a target image by using a first filter to generate a first feature map from the target image, and further performing a convolution process on the first feature map by using a second filter to generate a second feature map from the first feature map; and Use the first feature map to estimate a first likelihood map indicating the probability of the presence of an object having a first size at each position of the target image, and use the second feature map to estimate a second likelihood map indicating the probability of the presence of an object having a second size greater than the first size at each position of the target image, where the first size is determined based on a first reference size of the object, where the first size is any size within a first predetermined range from a first minimum size to a first maximum size; and the second size is determined based on a second reference size of the object, where the second size is any size within a second predetermined range from a second minimum size to a second maximum size, the first predetermined range does not overlap with the second predetermined range, and the second size is greater than the first size.

12. A non-transitory recording medium for causing a computer to execute: generating a first feature map from the target image by performing convolution processing on the target image using a first filter, and generating a second feature map from the first feature map by further performing convolution processing on the first feature map using a second filter; and using the first feature map to estimate a first likelihood map indicating the probability that an object having a first size exists at each position of the target image, and using the second feature map to estimate a second likelihood map indicating the probability that an object having a second size greater than the first size exists at each position of the target image, wherein the first size is determined based on a first reference size of the object, wherein, the first size is any size within a first predetermined range from a first minimum size to a first maximum size; and the second size is determined based on a second reference size of the object, where the second size is any size within a second predetermined range from a second minimum size to a second maximum size, the first predetermined range does not overlap with the second predetermined range, and the second size is greater than the first size.

Citation Information

Patent Citations

  • Information processing device, method for discriminating subject and computer program

    JP2019139618A