Information processing apparatus and information processing method

JP2024008592A5Active Publication Date: 2025-07-22CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022110586
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-07-08
Publication Date
2025-07-22
Estimated Expiration
2042-07-08

AI Technical Summary

Technical Problem

Existing tracking technologies using Deep Neural Networks (DNNs) suffer from erroneous tracking when similar objects are close to the target due to independent similarity calculations for feature parts, leading to incorrect predictions.

Method used

A method involving multiple calculation means to generate an inference map by transforming and combining feature amounts using Convolutional Neural Networks (CNNs) to accurately predict the position of a tracking target, suppressing erroneous tracking by learning from both target and similar object features.

Benefits of technology

The method effectively suppresses erroneous tracking by accurately distinguishing between tracking targets and similar objects, enhancing the precision of target localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a technology for suppressing erroneous tracking of a tracking target.SOLUTION: An information processing apparatus obtains a first feature of an image of a tracking target, obtains a second feature of an image of a search area, uses the first feature and the second feature to obtain an inference tensor representing the likelihood that the tracking target exists at each position in the image of the search area. Then, using the inference tensor, and obtains an inference map representing the position of the tracking target in the image of the search area.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a technology for tracking a target in an image. [Background technology]

[0002] There are technologies for tracking specific subjects within an image, such as those that use brightness or color information or template matching, but in recent years, technology that uses Deep Neural Networks (DNNs) has been attracting attention as a highly accurate tracking technology.

[0003] The technology described in Non-Patent Document 1 is one of the technologies for tracking a specific subject in an image. An image showing the target to be tracked and an image serving as a search region are input to a Convolutional Neural Network (CNN) with the same weights, and the position of the target to be tracked in the image of the search region is identified by calculating the cross-correlation between the results obtained from the CNN. While such a tracking method can accurately predict the position of the target to be tracked, if an object similar to the target to be tracked is present in the image, the cross-correlation value with the similar object becomes high, which tends to cause erroneous tracking in which the similar object is mistakenly tracked as the target to be tracked.

[0004] The technology described in Patent Document 1 attempts to suppress erroneous tracking of similar objects by predicting the positions of the tracking target and similar objects when the similar object is present near the tracking target. The technology described in Non-Patent Document 1 extracts the features of the tracking target and the search image, respectively, and calculates the similarity for each channel to detect the tracking target. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] JP 2013-219531 A [Non-patent literature]

[0006] [Non-Patent Document 1] “Fully-Convolutional Siamese Networks for Object Tracking”,arXiv 2016 Summary of the Invention [Problem to be solved by the invention]

[0007] However, in the method shown in Non-Patent Document 1, similarity is calculated independently for each part of the feature amount, so erroneous tracking occurs when a similar object having some features similar to the tracking target, such as contour or color, is nearby. The present invention provides a technology for suppressing erroneous tracking of a tracking target. [Means for solving the problem]

[0008] One aspect of the present invention is characterized in that it comprises a first calculation means for calculating a first feature of an image of a tracking target, a second calculation means for calculating a second feature of an image of a search area, a third calculation means for calculating an inference tensor representing the likelihood that the tracking target is present at each position of the image of the search area using the first feature and the second feature, and a fourth calculation means for calculating an inference map representing the position of the tracking target in the image of the search area using the inference tensor. Effect of the Invention

[0009] According to the present invention, erroneous tracking of a tracking target can be suppressed. [Brief description of the drawings]

[0010] [Figure 1] FIG. 1 is a block diagram showing an example of a hardware configuration of an information processing device. [Diagram 2] FIG. 2 is a block diagram showing an example of the functional configuration of an information processing device. [Diagram 3] 13 is a flowchart of an inference process. [Figure 4] FIG. 11 is a diagram showing details of the process in step S303. [Diagram 5] FIG. 4 is a diagram for explaining an inference map 412. [Figure 6] Flowchart of the CNN learning process. [Figure 7] Flowchart of the CNN learning process. [Figure 8] A diagram showing how CNN weight parameters are generated. [Figure 9] FIG. 13 is a diagram showing an operation when convolution processing is performed multiple times. [Figure 10] 6A and 6B are diagrams showing a generation process of a tracking target template feature amount; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.

[0012] [First embodiment] First, an example of the hardware configuration of an information processing device according to this embodiment will be described with reference to the block diagram of Fig. 1. Note that the hardware configuration shown in Fig. 1 is merely an example and can be changed / modified as appropriate.

[0013] The CPU 101 executes various processes using computer programs and data stored in the ROM 102 and the RAM 103. As a result, the CPU 101 controls the operation of the entire information processing device, and executes or controls each process that will be described as being performed by the information processing device.

[0014] The ROM 102 stores setting data for the information processing device, computer programs and data relating to the startup of the information processing device, computer programs and data relating to the basic operation of the information processing device, and the like.

[0015] The RAM 103 has an area for storing computer programs and data loaded from the ROM 102 or the storage unit 104, and an area for storing information received from an external device via the I / F 107. The RAM 103 also has a work area used when the CPU 101 executes various processes. In this way, the RAM 103 can provide various areas as needed.

[0016] The storage unit 104 is a large-capacity information storage device such as a hard disk drive device. The storage unit 104 stores an OS (operating system), computer programs and data for causing the CPU 101 to execute or control each process described as a process performed by the information processing device, etc. The computer programs and data stored in the storage unit 104 are loaded into the RAM 103 as appropriate under the control of the CPU 101, and become targets for processing by the CPU 101.

[0017] The storage unit 104 can be implemented by a medium (recording medium) and an external storage drive for realizing access to the medium. Examples of such media include flexible disks (FDs), CD-ROMs, DVDs, USB memories, MOs, and flash memories. The storage unit 104 may also be a server device that can be accessed by the information processing device via a network.

[0018] Display unit 105 is a device having a liquid crystal screen, a touch panel screen, an organic EL display, etc., and displays the processing results by CPU 101 as images, characters, etc. Note that display unit 105 may be a projection device such as a projector that projects images and characters.

[0019] The operation unit 106 is a user interface such as a keyboard, a mouse, a touch panel, etc., and can input various instructions to the CPU 101 by operating the operation unit 106. The operation unit 106 may be an external device capable of communicating with the information processing device. The operation unit 106 may also be a pen tablet, or the display unit 105 and the operation unit 106 may be combined to form a tablet. In this way, the type of device that implements the display function and the operation input function is not limited to a specific form.

[0020] The I / F 107 is a communication interface for performing data communication with an external device. The CPU 101, the ROM 102, the RAM 103, the storage unit 104, the display unit 105, the operation unit 106, and the I / F 107 are all connected to a system bus .

[0021] Next, a functional configuration example of an information processing device is shown in the block diagram of FIG. 2. In this embodiment, all of the storage units other than the storage unit 208 among the functional units shown in FIG. 2 are implemented by computer programs. In the following, the functional units (excluding the storage unit 208) of FIG. 2 may be described as the subject of processing, but in reality, the functions of the functional units are realized by the CPU 101 executing the computer program corresponding to the functional units. Note that some or all of the storage units other than the storage unit 208 among the functional units shown in FIG. 2 may be implemented by hardware. The storage unit 208 can be implemented by a memory device such as the RAM 103, the ROM 102, or the storage unit 104.

[0022] Next, a process (inference process) performed by the information processing device to output information indicating the position of a tracking target in an image as an inference map will be described with reference to the flowchart in Fig. 3. Note that the process according to the flowchart shown in Fig. 3 may be performed by the information processing device alone, or may be performed by a plurality of devices including the information processing device operating in cooperation with each other.

[0023] In step S301, the input unit 203 acquires an image including various objects (people, cars, animals, buildings, trees, etc.). The method of acquiring the image is not limited to a specific method, and the image may be acquired from the storage unit 208, or may be acquired from an external device such as an imaging device or a server device connected via the I / F 107.

[0024] In step S302, the input unit 203 acquires an image within the image area of ​​the tracking target as the tracking target image from the image acquired in step S301. The method of acquiring the tracking target image is not limited to a specific acquisition method. For example, the image acquired in step S301 may be displayed on the display unit 105, and an image within the image area of ​​the tracking target designated by the user by operating the operation unit 106 may be acquired as the tracking target image.

[0025] In step S303, the generation unit 204 performs CNN calculations to which the tracking target image is input, thereby generating a tracking target template feature amount representing the feature amount of the tracking target image as an output of the CNN. Then, the transformation unit 205 transforms the tracking target template feature amount in accordance with the number of dimensions of the weight parameters of the CNN used by the detection unit 206.

[0026] In step S304, the input unit 201 acquires an input image and search area information indicating a range (image area) in which the tracking target is searched in the input image. The method of acquiring the input image is not limited to a specific acquisition method, and may be acquired from the storage unit 208, or may be acquired from an external device such as an imaging device or a server device connected via the I / F 107. The method of acquiring the search area information is not limited to a specific acquisition method, and may be acquired from the storage unit 208, or may be acquired from an external device such as a server device connected via the I / F 107. Information that defines a region specified by a user operating the operation unit 106 may be acquired as the search area information. The extraction unit 202 then acquires an image within the image area indicated by the search area information from the input image as a search image, and performs a CNN operation to which the search image is input, thereby generating a feature amount (search area feature amount) of the search image as an output of the CNN.

[0027] In step S305, the detection unit 206 inputs the search area feature to the CNN in which the tracking target template feature transformed in step S303 is set as a weight parameter, and performs calculations on the CNN to generate an inference tensor representing the likelihood that the tracking target exists at each position of the search image as an output of the CNN. Here, the weight parameter is a parameter representing a weight value between layers in the CNN. The detection unit 206 then performs calculations on the CNN to which the inference tensor is input, and generates an inference map indicating the position of the tracking target in the search image as an output of the CNN.

[0028] The details of the process in step S303 above will be described with reference to FIG. 4. The generating unit 204 inputs the tracking target image 401 to the CNN 402 and performs the calculation of the CNN 402 to generate the tracking target template feature 403 as the output of the CNN 402. The CNN 402 is a CNN that has been trained in advance so as to obtain a tracking target template feature that makes it easy to distinguish between a tracking target and a non-tracking target. In this embodiment, the tracking target template feature 405 obtained by transforming the tracking target template feature 403 to match the number of dimensions of the weighting parameter of the CNN 409 is used as the weighting parameter of the CNN 409. Therefore, the CNN 402 is a CNN with a weight independent of the CNN 407. The dimension of the weighting parameter of the CNN 409 is four-dimensional, and for example, in the case of 3x3xCxY (C and Y are natural numbers), the dimension of the tracking target template feature 403 is set to 3x3xCY.

[0029] The transformation unit 205 transforms the 3x3xCY tracked template feature 403 into a 3x3xCxY tracked template feature 405 (transformation 404). The transformation method may be dimensional block division or shuffle division, and the same transformation method is used during inference and learning.

[0030] Next, the details of the processes in the above steps S304 and S305 will be described with reference to Figs. 4 and 5. The extraction unit 202 acquires an image within the image area indicated by the search area information from the input image as a search image 406. The extraction unit 202 then performs calculations on the CNN 407 to which the search image 406 is input, thereby generating search area features 408, which are features of the search image 406, as an output of the CNN 407.

[0031] The detection unit 206 sets the tracking target template feature 405 as a weight parameter of the CNN 409. The detection unit 206 inputs the search area feature 408 to the CNN 409 in which the tracking target template feature 405 is set as a weight parameter and performs a calculation (convolution calculation) of the CNN 409 to generate an inference tensor 410 representing the likelihood that the tracking target exists at each position of the search image 406. Here, while Non-Patent Document 1 calculates the cross-correlation by depthwise convolution of each feature of the tracking target image and the search area image, in this embodiment, the CNN 409 is full convolution. When the dimension of the search area feature 408 is WxHxC and the dimension of the weight parameter of the CNN 409 is 3x3xCxY (W, H, C, and Y are natural numbers), the dimension of the inference tensor 410 is WxHxY. At this time, the tracking target template feature amount 405 is trained so that the value of the inference tensor 410 corresponding to a position where the likelihood that the tracking target exists in the search image 406 is high increases. Then, the detection unit 206 performs an operation (convolution operation) of the CNN 411 to which the inference tensor 410 is input, thereby generating an inference map 412 indicating the position of the tracking target in the search image 406.

[0032] The inference map 412 will be described with reference to Fig. 5. For each position (each rectangle) in the search image 501, the "likelihood that the tracking target 502 is located in that rectangle" is obtained, and in Fig. 5, a rectangle 503 near the center of the tracking target 502 shows a high likelihood value. If the likelihood is equal to or greater than a threshold, it can be estimated that the tracking target 502 is located at a position corresponding to the rectangle 503.

[0033] Next, the process (CNN learning process) performed by the information processing device to learn CNN402, CNN407, and CNN411 that appeared in Fig. 4 will be described with reference to the flowchart in Fig. 6. Note that the process according to the flowchart shown in Fig. 6 may be performed by the information processing device alone, or may be performed by a plurality of devices including the information processing device operating in cooperation with each other.

[0034] In step S601, the input unit 203 acquires an image of an object to be learned (an object that can be a tracking target) as a tracking target image. As mentioned in the above description of step S302, the method of acquiring the tracking target image is not limited to a specific acquisition method.

[0035] In step S602, the generation unit 204 performs calculations on the CNN 402 to which the tracking target image acquired in step S601 is input, thereby generating tracking target template features representing features of the tracking target image as an output of the CNN 402. Then, the transformation unit 205 transforms the tracking target template features in accordance with the number of dimensions of the weight parameters of the CNN used by the detection unit 206.

[0036] In step S603, the input unit 201 acquires an input image and search area information indicating a range (image area) for searching for an object to be learned (an object that can be a tracking target) in the input image. As mentioned in the above description of step S304, the method of acquiring the input image and the search area information is not limited to a specific acquisition method. The extraction unit 202 acquires an image within the image area indicated by the search area information from the input image as a search image.

[0037] In step S604, the extraction unit 202 performs calculations on the CNN 407 to which the search image acquired in step S603 is input, thereby generating search region features, which are features of the search image, as an output of the CNN 407.

[0038] In the flowchart of FIG. 6, the processes of steps S603 and S604 and the processes of steps S601 and S602 are executed in parallel, but this is not limiting, and these processes may be executed sequentially.

[0039] In step S605, the detection unit 206 sets the tracking target template feature amount transformed in step S602 as a weight parameter of the CNN 409. The detection unit 206 inputs the search area feature amount generated in step S604 to the CNN 409 in which the tracking target template feature amount is set as a weight parameter, and performs calculations of the CNN 409 to generate an inference tensor. This inference tensor represents the likelihood that the learning target exists at each position of the search image. The detection unit 206 then performs calculations of the CNN 411 to which the inference tensor 410 is input, thereby generating an inference map indicating the position of the learning target in the search image.

[0040] In the CNN learning process, the goal is to update the weight parameters of each CNN so that the values ​​corresponding to the position of the learning target in the inference map are high and the values ​​at positions other than the learning target are low.

[0041] In step S606, the learning unit 207 calculates the difference between a map (teacher data) in which the position of the learning target in the search image is assigned with "1" and the positions other than the position of the learning target is assigned with "0" and the inference map generated in step S605 as a loss. As a loss function for calculating the loss, Cross Entropy Loss, Smooth L1 Loss, etc. can be used.

[0042] In step S607, the learning unit 207 updates the weight parameters of the CNN 402, CNN 407, and CNN 411 based on the loss calculated in step S606. At this time, in the transformation 404, a transformation inverse to the transformation in the above inference process is performed, and the error is propagated to the CNN 402. The weight parameters of the CNN 409 are set to the tracking target template feature amount 405, and do not need to be updated. The weight parameters are updated based on the back propagation method using Momentum SGD or the like.

[0043] In this embodiment, the output of the loss function for one image has been described for simplicity, but when multiple images are targeted, the loss is calculated for the scores estimated for the multiple images.Then, the weight parameters between layers of the CNN are updated so that the losses for the multiple images are all smaller than a predetermined threshold.

[0044] In step S608, the learning unit 207 stores the weight parameters (weight parameters of CNN402, CNN407, and CNN411) updated in step S607 in the storage unit 208. Then, when CNN402 is used in the above inference process, CNN402 to which the "weight parameters of CNN402" stored in the storage unit 208 are set is used. Similarly, when CNN407 is used in the above inference process, CNN407 to which the "weight parameters of CNN407" stored in the storage unit 208 are set is used. Similarly, when CNN411 is used in the above inference process, CNN411 to which the "weight parameters of CNN411" stored in the storage unit 208 are set is used.

[0045] In step S609, the learning unit 207 determines whether or not a learning end condition has been satisfied. The learning end condition is not limited to a specific end condition. For example, the learning end condition may be "the loss calculated in step S606 is equal to or less than a threshold", "the rate of change in the loss calculated in step S606 is equal to or less than a threshold", "the number of times the processes in steps S601 to S608 are repeated is equal to or greater than a threshold", etc.

[0046] Various modified examples of this embodiment will be described below. In each description of the modified example, the differences from the first embodiment will be described, and unless otherwise specified, the modified example will be considered to be the same as the first embodiment.

[0047] <Variation 1> In the first embodiment, the input unit 203 acquires an image of an object to be learned (an object that can be a tracking target) as a tracking target image (positive example) in the CNN learning process. In this modification, the input unit 203 acquires an image of a similar object that is similar to the learning target object as a tracking non-target image (negative example) in addition to the tracking target image in the CNN learning process.

[0048] The learning process of the CNN according to this modification will be described with reference to Fig. 7. The input unit 203 acquires a non-tracking target image 701 in addition to the tracking target image 401. As with the tracking target image 401, the method of acquiring the non-tracking target image 701 is not limited to a specific acquisition method.

[0049] Then, the generation unit 204 performs calculations on the CNN 402 to which a concatenated image formed by concatenating the tracking target image 401 and the tracking non-target image 701 is input, thereby generating the tracking target template feature amount 403 corresponding to the concatenated image as an output of the CNN 402.

[0050] Note that the non-tracking target image 701 is not limited to an image of a similar object similar to the object to be learned, but may be an image of an object different from the object to be learned, such as an image showing only the background. In this way, the features of similar objects and the background can be learned as negative examples, and erroneous tracking can be suppressed.

[0051] <Variation 2> In the first embodiment, the transformation unit 205 transforms the tracking target template feature amount in accordance with the number of dimensions of the weight parameters of the CNN used by the detection unit 206, and sets the transformed result as the weight parameters of the CNN. In this modification, the transformation unit 205 generates the tracking target template feature amount combined with a feature amount prepared in advance as the weight parameters of the CNN. A method of generating the weight parameters of the CNN according to this modification will be described with reference to FIG. 8.

[0052] The generating unit 204 inputs the tracking target image 401 to the CNN 402 and performs the calculation of the CNN 402 to generate a tracking target template feature 803 as an output of the CNN 402. The transforming unit 205 generates a tracking target template feature 805 by concatenating (combining) the tracking target template feature 803 with a feature 806 previously stored in the storage unit 208 (transformation 804). Instead of "concatenation (combination)", "reshape addition" may be applied. Here, the feature 806 is a weight previously learned so that the tracking target can be detected. In the example of FIG. 8, some dimensions of the feature 806 are replaced with the tracking target template feature 803. For example, when the dimension of the tracking target template feature 803 is 3x3x16 and the dimension of the feature 806 is 3x3x16x16, the transformation unit 205 replaces some dimensions of the feature 806 with the tracking target template feature 803 to generate a tracking target template feature 805 of 3x3x16x16. The storage unit 208 stores the feature 806 for each object category such as a person, a car, an animal, etc. The transformation unit 205 acquires the feature 806 corresponding to the target object category from the storage unit 208 and links (combines) it with the tracking target template feature 803. In this way, by combining the weight calculated from the tracking target image 401 with the weight sufficiently learned in advance, the tracking accuracy can be improved.

[0053] <Modification 3> In the second modification, the CNN402 and the CNN407 are separate CNNs having different weight parameters, but the same CNN may be applied to the CNN402 and the CNN407. By applying the same CNN to the CNN402 and the CNN407, for example, when the search area image 406 is used as the tracking target image 401, the tracking target template feature 803 and the search area feature 408 become the same. Therefore, it is not necessary to obtain both the tracking target template feature 803 and the search area feature 408 (it is sufficient to obtain one of them), and the overall processing can be speeded up. Furthermore, since the CNN409 includes the depthwise convolution of Non-Patent Document 1, it is possible to achieve a tracking accuracy equal to or higher than that of Non-Patent Document 1.

[0054] <Modification 4> In the first embodiment, the calculation (convolution process) by the CNN in which the weight parameters based on the tracking target template feature amount calculated from the tracking target image 401 are set is performed only once, but it may be performed multiple times. The operation in the case where such convolution process is performed multiple times will be described with reference to FIG.

[0055] The generation unit 204 inputs the tracking target image 401 to the CNN 902 and performs calculations of the CNN 902 to generate a tracking target template feature 403 and a tracking target template feature 903 as an output of the CNN 902. Here, the dimension of the search area feature 408 is WxHxCh, the dimension of the tracking target template feature 403 is 3x3x(ChxY), the dimension of the tracking target template feature 405 is 3x3xChxY, and the dimension of the tracking target template feature 903 is 3x3x(YxOUT).

[0056] The transformation unit 205 transforms (transforms 904) the tracked target template feature 903 in accordance with the number of dimensions of the weight parameters of the CNN 411, and generates a tracked target template feature 905 with dimensions 3x3xYxOUT.

[0057] The detection unit 206 inputs an inference tensor 410 with dimensions WxHxY to a CNN 911 in which the tracked target template feature amount 905 is set as a weight parameter, and performs calculations on the CNN 911 to generate an inference map 412 with dimensions WxHxOUT.

[0058] In the learning process, the CNN 911 does not update the weight parameters, similar to the CNN 409. In this way, by generating multiple weights and performing multiple stages of convolution processing, complex and nonlinear processing according to the tracking target image 401 becomes possible, and the detection accuracy is improved.

[0059] <Variation 5> In the first embodiment, the generation unit 204 generates the tracking target template feature amount only from the image acquired by the input unit 203, but the tracking target template feature amount may be generated using the tracking target template feature amount generated in the past in addition to the image. The generation process of the tracking target template feature amount according to this modification will be described with reference to FIG.

[0060] The generation unit 204 inputs the tracking target image 401 to the CNN 1002, performs calculations on the CNN 1002 to generate a tracking target template feature 1003 (dimension: 3x3x16), and inputs the generated tracking target template feature 1003 to the CNN 1006. The generation unit 204 also reads out a previously determined tracking target template feature 1005 (dimension: 3x3x(16xY)) from the storage unit 208 and inputs it to the CNN 1006. The generation unit 204 then performs calculations on the CNN 1006 to which the tracking target template feature 1003 and the tracking target template feature 1005 have been input, to generate a tracking target template feature 1007 (dimension: 3x3x(16xY)).

[0061] The transformation unit 205 transforms the tracked target template feature 1007 to match the number of dimensions of the weight parameters of the CNN 409, thereby generating the tracked target template feature 405 (dimension: 3x3x16xY) (transformation 404).

[0062] Note that instead of inputting the previously obtained tracking target template feature amount 1005 to the CNN 1006 as is, a moving average of the previously obtained tracking target template feature amount 1005 may be input to the CNN 1006 .

[0063] Also, instead of the CNN 1006, a GRU (Gated Recurrent Unit) or an LSTM (Long Short Term Memory) may be used.

[0064] Also, instead of inputting the tracking target template feature 1005 to the CNN 1006, the tracking target template feature 1003 may be integrated with the calculation result of the CNN 1006 to which it has been input to generate the tracking target template feature 1007. The integration method may be a method of calculating a moving average or a method of concatenating and convoluting. In this way, even if the visibility of the tracking target changes or it is blocked, it is possible to suppress a decrease in detection accuracy.

[0065] In this way, according to the first embodiment and its modified examples, even if a similar object similar to a tracking target is close to the tracking target, erroneous tracking of the similar object can be suppressed.

[0066] [Second embodiment] The information processing device described in the first embodiment and its modified examples may be a device such as a PC (personal computer), a tablet terminal device, or a smartphone. In this case, the information processing device may be configured with one such device or multiple devices. In the latter case, each device does not need to have the same configuration, and in this case, the information processing device may be a collection of devices each having a role, such as one or more devices that execute the process according to the above flowchart, a device that functions as a storage, etc.

[0067] In the first embodiment and its modified examples, the inference process and the CNN learning process are performed by the same device, but each process may be performed by a separate device. In this case, the device that performs the inference process uses the CNN generated by the device that performed the CNN learning process.

[0068] Furthermore, the numerical values, processing timing, processing order, processing subject, data (information) structure (including number of dimensions), acquisition method, transmission destination, transmission source, storage location, etc. used in each of the above embodiments and each of the modified examples are given as examples to provide a concrete explanation, and are not intended to be limited to these examples.

[0069] In addition, any part or all of the embodiments and modifications described above may be used in appropriate combination. In addition, any part or all of the embodiments and modifications described above may be used selectively.

[0070] (Other embodiments) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0071] The invention of this specification includes the following system, information processing device, information processing method, and computer program.

[0072] (Item 1) A first calculation means for calculating a first feature amount of an image of a tracking target; A second calculation means for calculating a second feature amount of the image in the search area; a third calculation means for calculating an inference tensor representing a likelihood that the tracking target exists at each position of the image in the search area using the first feature amount and the second feature amount; a fourth calculation means for calculating an inference map representing a position of the tracked target in an image of the search area using the inference tensor; An information processing device comprising:

[0073] (Item 2) moreover, 2. The information processing device according to item 1, further comprising a transformation means for transforming the first feature amount into a feature amount with a number of dimensions equal to the number of dimensions of the third calculation means.

[0074] (Item 3) moreover, 2. The information processing device according to item 1, further comprising: a transformation unit that transforms the first feature amount by combining the first feature amount with a feature amount that is stored in advance.

[0075] (Item 4) moreover, 2. The information processing device according to item 1, further comprising: a transformation means for transforming a feature amount obtained based on the first feature amount and a first feature amount previously obtained by the first calculation means.

[0076] (Item 5) the first calculation means, the second calculation means, the third calculation means, and the fourth calculation means are a Convolutional Neural Network (CNN); The third calculation means is a CNN that sets the first feature amount transformed by the transformation means as a weight parameter, and obtains the inference tensor using the second feature amount as an input. 5. The information processing device according to any one of items 2 to 4.

[0077] (Item 6) 6. The information processing device according to item 5, characterized in that the first calculation means and the second calculation means are different CNNs.

[0078] (Item 7) 6. The information processing device according to item 5, wherein the first calculation means and the second calculation means are the same CNN.

[0079] (Item 8) The information processing device described in any one of items 1 to 7, wherein the fourth calculation means calculates an inference map representing a position of the tracking target in the image of the search area by using a third feature of the image of the tracking target calculated by the first calculation means and the inference tensor.

[0080] (Item 9) moreover, The information processing device according to any one of items 1 to 8, further comprising a learning means for calculating an inference tensor representing the likelihood that a training object exists at each position of an image of a search area using a first feature of the image of the training object and a second feature of the image of the search area, calculating an inference map representing the position of the training object in the image of the search area using the inference tensor, and performing training of the first calculation means, the second calculation means, the third calculation means, and the fourth calculation means using the inference map.

[0081] (Item 10) moreover, The information processing device according to any one of items 1 to 8, further comprising a learning means for calculating an inference tensor representing the likelihood that a training object exists at each position of an image in a search area using a first feature of a concatenated image of the image of the training object and an image of an object other than the training object and a second feature of the image of the search area, calculating an inference map representing the position of the training object in the image of the search area using the inference tensor, and performing training of the first calculation means, the second calculation means, the third calculation means, and the fourth calculation means using the inference map.

[0082] (Item 11) An information processing method performed by an information processing device, A first calculation step in which a first calculation means of the information processing device calculates a first feature amount of an image of a tracking target; a second calculation step in which a second calculation means of the information processing device calculates a second feature amount of the image in the search area; a third calculation step in which a third calculation means of the information processing device calculates an inference tensor representing a likelihood that the tracking target exists at each position of the image in the search area using the first feature amount and the second feature amount; a fourth calculation step in which a fourth calculation means of the information processing device uses the inference tensor to obtain an inference map representing a position of the tracking target in an image of the search area; An information processing method comprising:

[0083] (Item 12) A computer program for causing a computer to function as each of the means of the information processing device according to any one of items 1 to 10.

[0084] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0085] 201: Input unit 202: Extraction unit 203: Input unit 204: Generation unit 205: Transformation unit 206: Detection unit 207: Learning unit 208: Storage unit

Claims

1. A first calculation means for calculating a first feature amount of an image of a tracking target; A second calculation means for calculating a second feature amount of the image in the search area; a third calculation means for calculating an inference tensor representing a likelihood that the tracking target exists at each position of the image in the search area using the first feature amount and the second feature amount; a fourth calculation means for calculating an inference map representing a position of the tracked target in an image of the search area using the inference tensor; An information processing device comprising:

2. moreover, 2 . The information processing apparatus according to claim 1 , further comprising: a transformation unit that transforms the first feature amount into a feature amount with a number of dimensions equal to that of the third calculation unit.

3. moreover, 2. The information processing apparatus according to claim 1, further comprising a transformation unit that transforms the first feature amount by combining the first feature amount with a feature amount that is stored in advance.

4. moreover, 2 . The information processing apparatus according to claim 1 , further comprising: a transformation unit that transforms a feature amount obtained based on the first feature amount and a first feature amount previously obtained by the first calculation unit.

5. the first calculation means, the second calculation means, the third calculation means, and the fourth calculation means are a Convolutional Neural Network (CNN); The third calculation means is a CNN that sets the first feature amount transformed by the transformation means as a weight parameter, and obtains the inference tensor using the second feature amount as an input.

3. The information processing apparatus according to claim 2.

6. The information processing apparatus according to claim 5 , wherein the first calculation means and the second calculation means are different CNNs.

7. The information processing apparatus according to claim 5 , wherein the first calculation means and the second calculation means are the same CNN.

8. The information processing device according to claim 1, characterized in that the fourth calculation means calculates an inference map representing the position of the tracking target in the image of the search area using a third feature of the image of the tracking target calculated by the first calculation means and the inference tensor.

9. moreover, The information processing device according to claim 1, further comprising a learning means for calculating an inference tensor representing the likelihood that a learning object exists at each position of an image of a search area using a first feature of the image of the learning object and a second feature of the image of the search area, calculating an inference map representing the position of the learning object in the image of the search area using the inference tensor, and using the inference map to train the first calculation means, the second calculation means, the third calculation means, and the fourth calculation means.

10. moreover, The information processing device according to claim 1, further comprising a learning means for calculating an inference tensor representing the likelihood that a learning object is present at each position of an image in a search area using a first feature of a concatenated image of the image of the learning object and an image of an object other than the learning object and a second feature of the image of the search area, calculating an inference map representing the position of the learning object in the image of the search area using the inference tensor, and performing training of the first calculation means, the second calculation means, the third calculation means, and the fourth calculation means using the inference map.

11. An information processing method performed by an information processing device, A first calculation step in which a first calculation means of the information processing device calculates a first feature amount of an image of a tracking target; a second calculation step in which a second calculation means of the information processing device calculates a second feature amount of the image in the search area; a third calculation step in which a third calculation means of the information processing device calculates an inference tensor representing a likelihood that the tracking target exists at each position of the image of the search area using the first feature amount and the second feature amount; a fourth calculation step in which a fourth calculation means of the information processing device uses the inference tensor to obtain an inference map representing a position of the tracking target in an image of the search area; An information processing method comprising:

12. A computer program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 10.