Image processing device, image processing method and imaging apparatus

The proposed system addresses the challenge of tracking objects with changing postures or occlusions by calculating a loss based on feature amount distances and overlapping rates, enabling accurate identification and tracking of objects even when they appear similar.

JP2025083574AInactive Publication Date: 2025-05-30CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025043766
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing tracking methods using Deep Neural Networks struggle with correctly identifying objects when their posture changes or occlusion occurs, leading to potential tracking failures due to erroneous association with similar-looking objects.

Method used

A system that calculates a loss based on the distance between feature amounts of the target object and non-target objects, along with their overlapping rates, to learn a model for accurate feature extraction and tracking, even in the presence of similar-looking objects.

Benefits of technology

The system effectively identifies and tracks objects even when they appear similar, ensuring accurate tracking by learning feature amounts that distinguish between target and non-target objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025083574000001_ABST
    Figure 2025083574000001_ABST
Patent Text Reader

Abstract

To provide a technique for correctly identifying an object being a tracking object and another object and tracking the object even when there exists an object that has the similar appearance to the object being the tracking object.SOLUTION: An image processing device obtains a loss on the basis of a distance between images of a feature amount of a tracking object in each image, a distance between feature amounts of the tracking object and a non-tracking object in the image and an overlapping ratio of an image region of the tracking object to an image region of the non-tracking object in one or more images, performs learning of a model for extracting a feature amount of the object from the image on the basis of the loss, and performs tracking processing of tracking the tracking object in the image on the basis of the learned model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for tracking an object in an image.

Background Art

[0002] As techniques for tracking an object in an image, there are those that use luminance or color information, those that use template matching, those that use a Deep Neural Network, and the like. One method using a Deep Neural Network is a method called Tracking by Detection. In this method, an object in an image is detected, and tracking is performed by associating the object with past objects. In Non-Patent Document 1, this association is performed using feature amounts obtained from a Neural Network learned by Metric Learning. Metric Learning is a method of learning a transformation into a space in which data with a higher degree of similarity is distributed closer in space.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Consider a method of performing tracking by extracting feature amounts corresponding to respective objects using a neural network learned by metric learning and associating them with past feature amounts. In such a method, when the posture of the object to be tracked changes or occlusion occurs, there is a possibility that the object to be tracked may be erroneously associated with another object that looks similar, resulting in tracking failure.

[0006] The present invention provides a technique for correctly identifying an object to be tracked and other objects and performing tracking even when there are objects that look similar to the object to be tracked.

Means for Solving the Problems

[0007] One aspect of the present invention includes a calculation means for obtaining a loss based on the distance between the feature amounts of the object to be tracked in each image, the distance between the feature amounts of the object to be tracked and the non-target object in the image, and the overlapping rate between the image area of the object to be tracked and the image area of the non-target object in one or more images; a learning means for performing learning of a model for extracting the feature amounts of an object from an image based on the loss; and a tracking means for performing a tracking process of tracking the object to be tracked in the image based on the model learned by the learning means.

Effects of the Invention

[0008] According to the configuration of the present invention, even when there are objects that look similar to the object to be tracked, the object to be tracked and other objects can be correctly identified and tracking can be performed.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant explanations are omitted.

[0011] [First Embodiment] First, an image processing device as a learning device that performs learning of a learning model used to correctly identify a tracking target (object to be tracked) and other objects (non-object to be tracked) in continuously captured images and perform tracking of the object to be tracked will be described.

[0012] An example of the functional configuration of the learning device 100 according to this embodiment is shown in the block diagram of FIG. 1. The learning process of the learning model by the learning device 100 having the functional configuration example shown in FIG. 1 will be described according to the flowchart of FIG. 2.

[0013] In step S101, the acquisition unit 110 acquires a first image and correct data (first correct data) of the first image. The first correct data is created in advance (known) for the tracking target object in the first image, and includes region information indicating the image region of the tracking target object in the first image and a label unique to the tracking target object. The region information indicating the image region of the tracking target object is, for example, information including the center position (position of the tracking target object) of the rectangular region, the width and height of the rectangular region, when the image region is a rectangular region. Here, when the image region of the tracking target object is a rectangular region, the center position of the rectangular region is set as the position of the tracking target object, but any position of the rectangular region may be set as the position of the tracking target object.

[0014] The acquisition unit 110 also acquires a second image captured subsequent to the first image and correct data (second correct data) of the second image. Similar to the first correct data, the second correct data is created in advance (known) for the tracking target object in the second image, and includes region information indicating the image region of the tracking target object in the second image and a label unique to the tracking target object. The same label is assigned to the same tracking target object in the first image and the second image.

[0015] Each of the first image and the second image is, for example, an image of the first frame in a moving image and an image of the second frame subsequent to the first frame. Also, for example, each of the first image and the second image is a target still image in a plurality of still images captured periodically or irregularly and a still image captured after the target still image. Further, processing such as cutting out a part of the acquired image as necessary may be performed.

[0016] In step S102, the object detection unit 120 inputs the first image into an object detection CNN (Convolutional Neural Network), which is a learning model for object detection, and executes the arithmetic processing of the object detection CNN. As a result, the object detection unit 120 obtains, as the detection result of each object in the first image, region information indicating the image region of the object and a score of the object. The region information indicating the image region of the object is, for example, information including the center position of the rectangular region, the width and height of the rectangular region when the image region is a rectangular region. The score of the object is a score with a value range of 0 to 1, which means that the higher the value, the more object-like it is. Examples of techniques for detecting an object from an image include "Liu, SSD: Single Shot Multibox Detector: ECCV2016".

[0017] In step S103, the feature extraction unit 130 inputs the first image into a feature extraction CNN, which is a learning model for feature extraction, and executes the arithmetic processing of the feature extraction CNN. As a result, the feature extraction unit 130 obtains a feature map (first feature map) representing the features of each region when the first image is divided equally vertically and horizontally. Then, the feature extraction unit 130 specifies the corresponding position on the first feature map corresponding to the center position of the image region of the tracking target object indicated by the first correct data, and obtains the feature at the specified corresponding position as the "feature of the tracking target object". In addition, the feature extraction unit 130 specifies the corresponding position on the first feature map corresponding to the center position of the image region indicated by the region information obtained in step S102, and obtains the feature at the specified corresponding position as the "feature of the non-tracking target object".

[0018] Similarly, the feature extraction unit 130 inputs the second image into a CNN for feature extraction, which is a learning model for feature extraction, and executes the arithmetic processing of the CNN for feature extraction. As a result, the feature extraction unit 130 obtains a feature map (second feature map) representing the features of each region when the second image is divided into equal parts vertically and horizontally. Then, the feature extraction unit 130 identifies the corresponding position on the second feature map corresponding to the center position of the image region of the tracking target object indicated by the second correct data, and obtains the feature at the identified corresponding position as the "feature of the tracking target object".

[0019] Here, the process for obtaining the features of the tracking target object will be described with reference to FIG. 4. In FIG. 4, a first feature map, which is a map in which the features of each divided region when a first image having a size of 640 pixels in width and 480 pixels in height is divided into 24 divided regions horizontally and 16 divided regions vertically are registered, is shown superimposed. The first feature map is a map in which (24x16) features with the number of channels being C (an arbitrary natural number) are registered. Since the position of the upper left corner of the first image is set as the origin (0,0), the position (x, y) on the first image takes values in the range of 0≦x≦639, 0≦y≦439, and the position (p, q) (position in units of divided regions) on the first feature map takes values in the range of 0≦p≦23, 0≦q≦15.

[0020] Position 402 is the position (x, y) = (312, 276) of the tracking target object indicated by the first correct data. At this time, the feature extraction unit 130 calculates (312x24 / 640) = 11.7, sets 11 obtained by deleting the decimal part of 11.7 as p, calculates (276x16 / 480) = 9.2, and sets 9 obtained by deleting the decimal part of 9.2 as q. As a result, the feature extraction unit 130 obtains the position (11, 9) on the first feature map corresponding to the position (312, 276) of the tracking target object on the first image. Rectangle 403 indicates the position (11, 9) on the first feature map corresponding to the position (312, 276) of the tracking target object on the first image. Then, the feature extraction unit 130 obtains the feature at position 403 (the 1x1xC array element at position 403) as the feature of the tracking target object in the first image.

[0021] Also, the feature amount extraction unit 130 performs the same process on the second image to obtain the feature amount of the object to be tracked in the second image. That is, the feature amount extraction unit 130 obtains the position on the second feature amount map corresponding to the position of the object to be tracked indicated by the second correct data, and uses the feature amount (the 1x1xC array element at the obtained position) at the obtained position as the feature amount of the object to be tracked in the second image.

[0022] In addition, when specifying the corresponding position on the second feature amount map corresponding to the center position of the image area indicated by the area information obtained in step S102, it is also specified using the above method of obtaining (p, q) from (x, y).

[0023] Here, the relationship between the object detection CNN and the feature amount extraction CNN will be described with reference to FIG. 3. The object detection CNN and the feature amount extraction CNN may be implemented in any configuration. For example, as shown in FIG. 3(a), the object detection CNN and the feature amount extraction CNN may be provided separately, and each may be configured to operate independently on the input images (the first image and the second image). Also, as shown in FIG. 3(b), a common configuration may be provided as a shared CNN for the object detection CNN and the feature amount extraction CNN, and the configuration not included in the shared CNN in the object detection CNN may be reconfigured as the object detection CNN, and the configuration not included in the shared CNN in the feature amount extraction CNN may be reconfigured as the feature amount extraction CNN. In this case, the input image is input to the shared CNN, and the shared CNN inputs the processing result for the input image to the object detection CNN and the feature amount extraction CNN. The object detection CNN operates with the processing result output from the shared CNN as the input, and the feature amount extraction CNN operates with the processing result output from the shared CNN as the input.

[0024] Next, in step S104, the loss calculation unit 140 uses the feature amount obtained in step S103 to calculate the feature amount distance d, which is the Euclidean distance between the feature amount of the object to be tracked in the first image and the feature amount of the object to be tracked in the second image. 1Find it. Further, the loss calculation unit 140 uses the feature amount obtained in step S103 to calculate the Euclidean distance d between the feature amount of the object to be tracked in the first image and the feature amount of the non-target object in the first image. 2 Find it. The distance between feature amounts is the distance between each feature amount in the feature amount space to which each feature amount belongs. Since the method for obtaining the distance between feature amounts is well-known, the description thereof is omitted.

[0025] Next, in step S105, the loss calculation unit 140 obtains the parameters of the loss function. In the present embodiment, the Triplet loss given by the following (Equation 1) is used as the loss function.

[0026] loss = max(rd 1 -d 2 + m, 0) … (Equation 1) In this loss function, since loss occurs when d 1 is large and d 2 is small, it is possible to obtain feature amounts such that the distance between feature amounts for the same object is smaller than the distance between feature amounts for different objects.

[0027] In this (Equation 1), r and m are parameters of the loss function. r is the distance parameter and m is the margin parameter. When r is set large, learning is performed so that the distance between feature amounts for the same object becomes small, and when m is set large, learning is performed so that the distance between feature amounts for different objects becomes far. In the present embodiment, by dynamically setting these parameters using the correct data and the region of the object detected in step S102, feature amounts with higher object discrimination performance are learned.

[0028] Note that the loss function to be used is not limited to a specific function as long as distance learning can be performed. For example, Contrastive loss or Softmax loss may be used as the loss function.

[0029] Here, d 1 , d 2,Regarding the process for obtaining r and m, a specific example shown in FIG. 5 will be given for explanation. FIG. 5(a) shows the first image, and FIG. 5(b) shows the second image.

[0030] The first image in FIG. 5(a) includes the target object to be tracked 501 and non-target objects to be tracked 502 and 503. Rectangle 506 is the image area of the target object to be tracked 501 defined by the first correct data. Rectangle 507 is the image area of the non-target object to be tracked 502 detected in step S102. Rectangle 508 is the image area of the non-target object to be tracked 503 detected in step S102.

[0031] The second image in FIG. 5(b) includes the target object to be tracked 504 and the non-target object to be tracked 505. Rectangle 509 is the image area of the target object to be tracked 504 defined by the second correct data, and rectangle 510 is the image area of the non-target object to be tracked 505 detected in step S102.

[0032] In such a case, the feature extraction unit 130 calculates the feature distance d between the feature at the corresponding position on the first feature map corresponding to the "position of the target object to be tracked 501 in the first image" indicated by the first correct data and the feature at the corresponding position on the second feature map corresponding to the "position of the target object to be tracked 504 in the second image" indicated by the second correct data. 1 to obtain.

[0033] In addition, the feature extraction unit 130 selects one of the non-target objects to be tracked 502 and 503 (the non-target object to be tracked 502 in FIG. 5(a)) as the selected object, and calculates the feature distance d between the feature at the corresponding position on the first feature map corresponding to the position detected in step S102 for the selected object and the feature at the corresponding position on the first feature map corresponding to the "position of the target object to be tracked 501 in the first image" indicated by the first correct data. 2Find it. When a plurality of non-target objects to be tracked are included in the first image, the method for selecting a selected object from the plurality of non-target objects is not limited to a specific selection method. For example, the non-target object at the position closest to the position of the target object to be tracked (the closest one) is selected as the selected object.

[0034] In addition, the feature extraction unit 130 obtains the distance parameter r using the following (Equation 2) and obtains the margin parameter m using the following (Equation 3).

[0035] r = 1 … (Equation 2) m = m 0 + IoU … (Equation 3) Here, m 0 is an arbitrary real number. IoU (Intersection over Union) is the overlap ratio of the areas of two image regions. In the case of FIG. 5, the overlap ratio of the image region 506 of the target object to be tracked and the image region 507 of the non-target object 502 selected as the selected object is obtained as IoU.

[0036] In step S106, the loss calculation unit 140 calculates the loss loss by calculating (Equation 1) using d 1 , d 2 , r, and m obtained in the above process.

[0037] FIG. 6 is an image diagram showing the feature amount distance in a two-dimensional space when the loss is calculated by the above method according to the present embodiment. The feature amount 601 is the feature amount of the target object to be tracked in the first image, and the feature amount 602 is the feature amount of the target object to be tracked in the second image. d 1 is the feature amount distance between the feature amount 601 and the feature amount 602. The feature amounts 603 and 604 are the feature amounts of the non-target objects in the first image. d 2 is the feature amount distance between the feature amount 601 and the feature amount 604, and d 2 ’ is the feature amount distance between the feature amount 601 and the feature amount 603. Also, m, m 0 , and IoU are the margin parameters shown in (Equation 3).

[0038] In this method, when there is no overlap between the image region of the target object to be tracked and the image region of the non-target object to be tracked in the first image, the distance d between feature amounts 2 ’ is learned to be greater than the sum of the distance d between feature amounts 1 and the margin parameter m 0 . On the other hand, when there is overlap between the image region of the target object to be tracked and the image region of the non-target object to be tracked in the first image, the distance d between feature amounts 2 is learned to be greater than the sum of the distance d between feature amounts 1 and the margin parameter m. Therefore, for a non-target object with a high IoU with the image region of the target object to be tracked, that is, a non-target object that is close to the target object to be tracked and has a high risk of incorrect tracking switching, learning can be performed so that the distance between feature amounts becomes larger.

[0039] Note that the loss function is not limited to the above form as long as it changes parameters based on the IoU with the image region of the target object to be tracked. Also, when calculating the loss, a plurality of sets of the first image and the second image are prepared, the above processing is performed for each set to obtain the loss for each set, and the average of the losses obtained for each set may be used as the final loss.

[0040] In step S107, the learning unit 150 performs learning of the CNN for feature amount extraction by updating the parameters (such as weight values) in the CNN for feature amount extraction so as to minimize the loss loss obtained in step S106. Note that the learning unit 150 may simultaneously perform learning of the CNN for object detection in addition to learning of the CNN for feature amount extraction. Update of parameters such as weights can be considered to be performed using, for example, the error backpropagation method.

[0041] Next, an image processing apparatus as a tracking device that performs processing for tracking a target object in an image using the learned CNN described above will be described. A functional configuration example of the tracking device 200 according to the present embodiment is shown in the block diagram of FIG. 7. The tracking process of the target object by the tracking device 200 having the functional configuration example shown in FIG. 7 will be described according to the flowchart of FIG. 8.

[0042] In step S201, the image acquisition unit 210 acquires one captured image. The one captured image is the most recently captured image among the continuously captured images, and for example, it may be the captured image of the latest frame in a moving image, or the most recently captured still image in a still image captured periodically or irregularly.

[0043] In step S202, the object detection unit 220 inputs the captured image acquired in step S201 into the object detection CNN and executes the arithmetic processing of the object detection CNN. As a result, the object detection unit 220 acquires, as the detection result of the object, region information indicating the image region of the object and the score of the object for each object in the captured image acquired in step S201.

[0044] In step S203, the feature extraction unit 230 inputs the captured image acquired in step S201 into the feature extraction CNN and executes the arithmetic processing of the feature extraction CNN to obtain a feature map of the captured image. Then, for each object acquired in step S202, the feature extraction unit 230 specifies the corresponding position on the feature map corresponding to the center position of the image region indicated by the region information of the object, and acquires the feature at the specified corresponding position as "the feature of the object".

[0045] If the captured image acquired in step S201 is the first captured image (captured image at the initial time) since the start of the tracking process, the process proceeds to step S204. On the other hand, if the captured image acquired in step S201 is the second or later captured image (not the captured image at the initial time) since the start of the tracking process, the process proceeds to step S207.

[0046] In step S204, the feature quantity storage unit 240 assigns an ID (object ID) unique to the object to each feature quantity of the object acquired in step S203. For example, object IDs = 1, 2, 3,... are assigned in order from the object with the highest score.

[0047] And in step S205 (when proceeding from step S204 to step S205), the feature quantity storage unit 240 stores the feature quantities of the respective objects to which the object IDs were assigned in step S204 in the tracking device 200.

[0048] In step S207, the feature quantity comparison unit 250 selects one unselected feature quantity from among the feature quantities of the respective objects (feature quantities at the current time) acquired in step S203 as the selected feature quantity. Then, the feature quantity comparison unit 250 acquires n (n is a natural number) feature quantities (stored feature quantities) for each object in ascending order of proximity to the current time from the feature quantities stored in the tracking device 200. And the feature quantity comparison unit 250 obtains the feature quantity distance (the method for obtaining the feature quantity distance is the same as above) between each of the acquired stored feature quantities and the selected feature quantity. Note that the feature quantity comparison unit 250 may obtain one stored feature quantity (for example, the average value of the n stored feature quantities) for each object from the n stored feature quantities acquired for the object, and obtain the feature quantity distance between the obtained one stored feature quantity for each object and the selected feature quantity.

[0049] In step S208, the feature quantity comparison unit 250 specifies the minimum feature quantity distance among the feature quantity distances obtained in step S207. Then, the feature quantity comparison unit 250 assigns the object ID assigned to the object of the stored feature quantity corresponding to the specified minimum feature quantity distance to the object of the selected feature quantity. Also, when the "object ID assigned to the object of the stored feature quantity corresponding to the minimum feature quantity distance among the feature quantity distances obtained in step S207" has already been assigned to another object, the feature quantity comparison unit 250 issues a new object ID and assigns it to the object of the selected feature quantity. The object ID may be in descending order in order from the object with the highest score, but is not limited to this.

[0050] By performing the processes of step S207 and step S208 for the feature amounts of all the objects acquired in step S203, object IDs can be assigned to all the objects detected in step S202. Note that the above process for assigning object IDs to the feature amounts at the current time is just an example and is not intended to be limited to the above process. For example, the Hungarian algorithm may be used. Also, in step S207 and step S208, the following processes may be performed.

[0051] In step S207, the feature amount matching unit 250 selects one unselected feature amount from among the feature amounts (feature amounts at the current time) of each object acquired in step S203 as the selected feature amount. Then, the feature amount matching unit 250 obtains the feature amount distance (the method for obtaining the feature amount distance is the same as above) between the feature amount of the object stored in the tracking device 200 in the past (stored feature amount) and the selected feature amount.

[0052] In step S208, the feature amount matching unit 250 identifies the feature amount distances that are less than the threshold among the feature amount distances obtained in step S207. Note that when the feature amount matching unit 250 identifies a plurality of feature amount distances as the feature amount distances less than the threshold, the feature amount matching unit 250 identifies the minimum feature amount distance among the identified plurality of feature amount distances. Then, the feature amount matching unit 250 assigns the object ID assigned to the object with the most recent storage timing among the stored feature amounts corresponding to the identified feature amount distance to the object with the selected feature amount. Also, when there is no feature amount distance less than the threshold among the feature amount distances obtained in step S207, the feature amount matching unit 250 issues a new object ID and assigns it to the object with the selected feature amount.

[0053] That is, at the feature amount at the current time, if the feature amount distance from the target feature amount stored in the tracking device 200 in the past is less than the threshold, the object ID of the object with the target feature amount is assigned to the feature amount at the current time in order to associate the object with the feature amount at the current time and the object with the target feature amount.

[0054] And in step S205 (when proceeding from step S208 to step S205), the feature amount storage unit 240 stores the feature amounts of the respective objects to which the object IDs are assigned in step S208 in the tracking device 200.

[0055] Also, the feature amount comparison unit 250 outputs, as a tracking result, the object ID stored in the tracking device 200 in step S205 and the image region of the object to which the object ID is assigned. The tracking result may be such that the tracking device 200 or another device overlays a frame indicating the image region of the object included in the tracking result on the captured image and displays it together with the object ID of the object, or may transmit it to an external device. Note that the processing according to the flowchart of FIG. 8 above is performed for each captured image for which tracking processing is performed.

[0056] As described above, in the present embodiment, learning is performed such that the feature amount distance becomes larger for non-target objects that are close to the target object to be tracked in the captured image and have a high risk of incorrect tracking switching. Thereby, it is possible to obtain feature amounts with higher object identification performance, and even when there are objects with similar appearances, tracking can be performed with high accuracy.

[0057] <Modification Example 1> In this modification example, a case where correct data is assigned to each object (target object to be tracked and / or non-target object) in the image will be described. The correct data according to this modification example will be described with reference to FIG. 9. FIG. 9(a) shows a configuration example of correct data as a table in which the center positions (horizontal position coordinates, vertical position coordinates), widths, heights, and labels of the respective image regions 904, 905, and 906 of the objects 901, 902, and 903 in the image of FIG. 9(b) are registered. It is assumed that at least one object with the same label is present in the first image and the second image, respectively.

[0058] When calculating the loss in the loss calculation unit 140, an object with a certain label is regarded as a target object to be tracked, and an object with a different label is regarded as a non-target object to be tracked. Alternatively, the above calculation may be performed for objects with different labels, and the final loss may be determined by taking the average of them.

[0059] <Modification Example 2> In this modification example, the method of obtaining the feature amounts of the target object to be tracked and the non-target object to be tracked by the feature amount extraction unit 130 will be described taking FIG. 10 as an example. In the image 1001, the area 1002 is the image area of the target object to be tracked indicated by the correct data, and the areas 1003 and 1004 are the image areas of the non-target objects to be tracked detected by the object detection unit 120. The image 1005 is an image within the area 1002, the image 1006 is an image within the area 1003, and the image 1007 is an image within the area 1004.

[0060] The feature amount extraction unit 130 converts the images 1005 to 1007 into images of a specified image size (here, 32 pixels x 32 pixels) (such as resizing). That is, the feature amount extraction unit 130 converts the image 1005 into an image 1008 of 32 pixels x 32 pixels, converts the image 1006 into an image 1009 of 32 pixels x 32 pixels, and converts the image 1007 into an image 1010 of 32 pixels x 32 pixels.

[0061] Then, the feature amount extraction unit 130 inputs the image 1008 into the CNN for feature amount extraction and executes the arithmetic processing of the CNN for feature amount extraction, thereby obtaining the feature amount of the object corresponding to the image 1008. Also, the feature amount extraction unit 130 inputs the image 1009 into the CNN for feature amount extraction and executes the arithmetic processing of the CNN for feature amount extraction, thereby obtaining the feature amount of the object corresponding to the image 1009. Also, the feature amount extraction unit 130 inputs the image 1010 into the CNN for feature amount extraction and executes the arithmetic processing of the CNN for feature amount extraction, thereby obtaining the feature amount of the object corresponding to the image 1010. The feature amount of the object can be obtained as a vector of an arbitrary number of dimensions.

[0062] By performing such processing on the first image and the second image, the feature amounts of the objects in the respective images can be obtained. Therefore, hereinafter, the same processing as in the first embodiment may be performed.

[0063] <Modification Example 3> In this modification example, the above distance parameter r and margin parameter m are obtained according to (Equation 4) and (Equation 5), respectively.

[0064] r = 1 + IoU 1 ·IoU 2 (Equation 4) m = m 0 (Equation 5) Here, m 0 is an arbitrary positive real number. For the calculation method of IoU 1 , IoU 2 , the calculation method of IoU will be described with reference to FIG. 11. FIG. 11(a) shows the first image, and (b) shows the second image.

[0065] As shown in FIG. 11(a), the first image includes a target object 1101 to be tracked and a non-target object 1102 (when the first image includes a plurality of non-target objects, the non-target object closest to the target object 1101). IoU 1 is the IoU between the image region 1105 of the target object 1101 to be tracked and the image region 1106 of the non-target object 1102.

[0066] Also, as shown in FIG. 11(b), the second image includes a target object 1103 to be tracked and a non-target object 1104 (when the second image includes a plurality of non-target objects, the non-target object closest to the target object 1103). IoU 2 is the IoU between the image region 1107 of the target object 1103 to be tracked and the image region 1108 of the non-target object 1104.

[0067] That is, when there is a non-target object to be tracked that has an overlapping portion with the image area of the target object to be tracked in both the first image and the second image, r takes a value greater than 1. By calculating the loss in this way, when there is a non-target object to be tracked approaching the target object to be tracked, learning can be performed so that the distance between feature amounts becomes smaller. As a result, even when the target object to be tracked is blocked by another object or blocks another object, it is possible to obtain a feature amount that robustly identifies the target object to be tracked.

[0068] Note that the loss function is not limited to the above form as long as it changes the parameter when there is a non-target object to be tracked that has an overlapping portion with the image area of the target object to be tracked in both the first image and the second image.

[0069] <Modification Example 4> In this modification example, when calculating the loss, a plurality of sets of the first image and the second image are prepared, the above processing is performed for each set to obtain the loss for each set, and the sum of the losses obtained for each set weighted according to the score of the non-target object to be tracked corresponding to the set is taken as the final loss. In this modification example, the loss loss is obtained according to the following (Equation 6).

[0070]

Equation

[0071] Here, N is the total number of the above sets, and loss i is the loss obtained for the i-th set. Also, p i is the score of the non-target object to be tracked selected when obtaining d 2 in Equation (1) for the i-th set. When calculating the loss in this way, the contribution to the loss of a non-target object to be tracked with a high score, that is, an object with a high risk of an incorrect tracking switch occurring, becomes large. As a result, it is possible to focus on important cases for the tracking performance and perform efficient learning. Note that the weighting of the loss is not limited to the above form as long as it is changed by the score.

[0072] <Modified Example 5> In this modified example, Contrastive loss is used as the loss function. Contrastive loss is represented by the following (Equation 7).

[0073]

Equation

[0074] Here, d is the distance between feature quantities corresponding to two objects, and y is 1 when the two objects are the same object and 0 when they are different objects. In this loss function, learning is performed such that d is small when the two objects are the same object and d is large when they are different objects. Also, r and m are the distance parameter and the margin parameter, respectively. Similar to the first embodiment, when r is set large, learning is performed such that the distance between feature quantities of the same objects becomes small, and when m is made large, learning is performed such that the distance between feature quantities of different objects becomes far. When setting the parameters, for example, (Equation 2) and (Equation 3) can be used in the same manner as in the first embodiment.

[0075] <Modified Example 6> In the first embodiment, CNN was used as the learning model. However, the model applicable to the learning model is not limited to CNN, and other models (for example, machine learning models) may be used.

[0076] In the first embodiment and its modified examples, an example of a method for obtaining the loss based on the distance between images of feature quantities of the object to be tracked in each image, the distance between the feature quantities of the object to be tracked and the non-object to be tracked in the image, and the overlapping rate between the image region of the object to be tracked and the image region of the non-object to be tracked in one or more images was described. However, the method for obtaining the loss using these pieces of information is not limited to the above method.

[0077] In the first embodiment, the learning device 100 and the tracking device 200 were described as separate image processing devices, but they may be incorporated into one image processing device. That is, one image processing device may perform both the processes described as being performed by the learning device 100 and the processes described as being performed by the tracking device 200. In such a case, for example, an imaging device can be configured that includes an imaging unit that captures a moving image or captures a still image periodically or irregularly, and an image processing device capable of executing both the processes described as being performed by the learning device 100 and the processes described as being performed by the tracking device 200. In such an imaging device, the operation of the learning device 100 (learning of the object detection CNN and the feature extraction CNN) can be executed using the captured image captured by the device itself. Then, after the operation, the imaging device can execute the tracking process of the object in the captured image captured by the device itself using the learned object detection CNN and feature extraction CNN.

[0078] Also, in the first embodiment, the Euclidean distance was used as the distance between features, but the distance between features may be obtained by any calculation and is not limited to a specific calculation method (any type of distance may be adopted as the distance between features).

[0079] [Second Embodiment] Each functional unit shown in FIGS. 1 and 7 may be implemented in hardware or in software (computer program). In the latter case, a computer device capable of executing such a computer program is applicable to the above-described learning device 100 and tracking device 200. A hardware configuration example of a computer device applicable to the learning device 100 and the tracking device 200 will be described with reference to the block diagram of FIG. 12.

[0080] The CPU 1201 executes various processes using the computer programs and data stored in the RAM 1202 and the ROM 1203. As a result, the CPU 1201 controls the operation of the entire computer device and also executes or controls the various processes described as being performed by the learning device 100 and the tracking device 200.

[0081] The RAM 1202 has areas for storing computer programs and data loaded from the ROM 1203 and the external storage device 1206. Also, the RAM 1202 has an area for storing data received from the outside via the I / F 1207. Further, the RAM 1202 has a work area used when the CPU 1201 executes various processes. Thus, the RAM 1202 can appropriately provide various areas.

[0082] The ROM 1203 stores computer device setting data, computer programs and data related to the startup of the computer device, computer programs and data related to the basic operations of the computer device, and the like.

[0083] The operation unit 1204 is a user interface such as a keyboard, a mouse, or a touch panel, and various instructions can be input to the CPU 1201 by the user's operation.

[0084] The display unit 1205 has a liquid crystal screen or a touch panel screen, and can display the processing result by the CPU 1201 as an image, characters, or the like. For example, the display unit 1205 can display the result of the above tracking process. Further, the display unit 1205 may be a projection device such as a projector that projects images and characters.

[0085] The external storage device 1206 is a large-capacity information storage device such as a hard disk drive. The external storage device 1206 stores an OS (operating system), computer programs, and data for causing the CPU 1201 to execute or control various processes described as being performed by the learning device 100 and the tracking device 200. The data stored in the external storage device 1206 includes the data described as being stored in the learning device 100 and the tracking device 200, the data treated as known data in the above description, data related to CNN, and the like. The computer programs and data stored in the external storage device 1206 are appropriately loaded into the RAM 1202 according to the control by the CPU 1201 and become the processing targets by the CPU 1201.

[0086] The I / F 1207 is a communication interface for performing data communication with an external device. For example, the I / F 1207 can be connected to an imaging device capable of imaging moving images and still images, a server device holding moving images and still images, or a network to which these devices are connected. In this case, the moving images and still images imaged by the imaging device and the moving images and still images supplied from the server device are supplied to the computer device via the I / F 1207, and the supplied moving images and still images are stored in the RAM 1202 and the external storage device 1206.

[0087] The CPU 1201, the RAM 1202, the ROM 1203, the operation unit 1204, the display unit 1205, the external storage device 1206, and the I / F 1207 are all connected to the system bus 1208.

[0088] Note that the hardware configuration shown in FIG. 12 is an example of the hardware configuration of a computer device applicable to the learning device 100 and the tracking device 200, and can be appropriately modified / changed.

[0089] Also, the numerical values, processing timings, processing orders, processing entities, transmission destinations / sources / storage locations of data (information), etc. used in each of the above-described embodiments and each modification example are given as examples for the purpose of specific explanation, and are not intended to be limited to such examples.

[0090] Further, some or all of each of the above-described embodiments and each modification example may be used in appropriate combination. Also, some or all of each of the above-described embodiments and each modification example may be selectively used.

[0091] (Other Embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and causing one or more processors in the computer of the system or device to read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0092] The invention is not limited to the above-described embodiments, and various changes and modifications are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention.

Explanation of Reference Numerals

[0093] 100: Learning device 110: Acquisition unit 120: Object detection unit 130: Feature amount extraction unit 140: Loss calculation unit 150: Learning unit

Claims

1. a calculation means for calculating a loss based on a distance between images of feature amounts of a tracked object in each image, a distance between feature amounts of a tracked object and a non-tracked object in each image, and an overlap rate between an image area of ​​the tracked object and an image area of ​​a non-tracked object in one or more images; A learning means for learning a model for extracting a feature amount of an object from an image based on the loss; A tracking means for performing a tracking process for tracking a tracking target object in an image based on the model learned by the learning means; An image processing device comprising:

2. The calculation means is The feature amounts of the tracked object and the non-tracked object in the first image are obtained from a feature amount map that is a result of the calculation of the model to which the first image is input, and the feature amount of the tracked object in a second image subsequent to the first image is obtained from the feature amount map that is a result of the calculation of the model to which the second image is input.

2. The image processing device according to claim 1,

3. The calculation means is An image region of an object in a first image is resized, and a result of an operation of the model with the resized image region input is acquired as a feature amount of the object in the first image; an image region of an object in a second image subsequent to the first image is resized, and a result of an operation of the model with the resized image region input is acquired as a feature amount of the object in the second image.

2. The image processing device according to claim 1,

4. 4. The image processing device according to claim 1, wherein the calculation means calculates a loss based on a distance between images of feature amounts of the tracked object in each image, a distance between feature amounts of the tracked object and non-tracked object in the image, and an overlap rate between an image area of ​​the tracked object and an image area of ​​the non-tracked object in the image.

5. 4. The image processing device according to claim 1, wherein the calculation means calculates a loss based on a distance between images of feature amounts of the tracked object in each image, a distance between the feature amounts of the tracked object and the non-tracked object in the image, an overlap rate between an image area of ​​the tracked object and an image area of ​​the non-tracked object in a first image, and an overlap rate between an image area of ​​the tracked object and an image area of ​​the non-tracked object in a second image subsequent to the first image.

6. the calculation means calculates the loss for each set of images, and calculates a final loss by summing the losses calculated for each set of images weighted according to a score indicating an object likelihood obtained in the detection of the non-tracked object; The learning means learns the model based on the final loss.

6. The image processing apparatus according to claim 1,

7. The calculation means calculates the loss for each set of images, and calculates an average of the losses calculated for each set of images as a final loss; The learning means learns the model based on the final loss.

6. The image processing apparatus according to claim 1,

8. 4. The image processing apparatus according to claim 2, wherein the non-tracking target object in the first image is the non-tracking target object that is closest to the tracking target object in the first image.

9. An imaging unit that captures an image; The image processing device according to any one of claims 1 to 8, An imaging device comprising:

10. An image processing method performed by an image processing device, comprising: a calculation step in which a calculation means of the image processing device calculates a loss based on a distance between images of feature amounts of the tracked object in each image, a distance between feature amounts of the tracked object and non-tracked object in each image, and an overlap rate between an image area of ​​the tracked object and an image area of ​​a non-tracked object in one or more images; a learning step in which a learning means of the image processing device learns a model for extracting a feature amount of an object from an image based on the loss; a tracking step in which a tracking means of the image processing device performs a tracking process for tracking a tracking target object in an image based on the model learned in the learning step; An image processing method comprising:

11. A computer program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Object tracking method, object tracking device, and program

    JP2018026108A