Information processing device and control method thereof

By alternately combining and grouping feature amounts in the spatial direction within an information processing apparatus, the computational inefficiencies in token mixing are addressed, resulting in improved efficiency for image recognition and tracking tasks.

JP2025080589APending Publication Date: 2025-05-26CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023193846
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-26

AI Technical Summary

Technical Problem

The processing amount in token mixing, such as Multi-head Self Attention (MSA), increases significantly due to the square of the size in the spatial direction, leading to computational inefficiencies in image recognition tasks.

Method used

An information processing apparatus that alternately combines and groups template and search feature amounts in the spatial direction, allowing for token mixing within reduced spatial groups, thereby minimizing computational load.

Benefits of technology

This approach reduces the processing amount in token mixing, enhancing computational efficiency and enabling more effective image recognition and tracking tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025080589000001_ABST
    Figure 2025080589000001_ABST
Patent Text Reader

Abstract

To reduce an amount of processing in token mixing.SOLUTION: An information processing device comprises: first generation means for generating a first feature from a first image; second generation means for generating a second feature from a second image; combination means for combining the first feature and the second feature; division means for dividing the combined features into a plurality of groups; and mixing means for mixing the features included in each of the plurality of groups for each group. At least one group included in the plurality of groups includes both a first feature and a second feature.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to processing using a neural network.

Background Art

[0002] In recent years, image recognition technologies such as image classification, object detection, and object tracking have seen a dramatic improvement in accuracy due to the emergence of deep neural networks (DNNs). Although there are various DNN structures, in image recognition, convolutional neural networks (CNNs) that perform convolutional operations over multiple layers have mainly been used. On the other hand, in Non-Patent Document 1, a Vision Transformer (ViT) that applies a Transformer, which is used in natural language processing, to image recognition has been proposed. A Transformer is a structure that represents the relationship between words using Attention in natural language processing. However, ViT has problems such as a large number of parameters and a large amount of computation.

[0003] In Non-Patent Document 2, a technique has been proposed that divides features into several rectangular windows and performs MSA (Multi-head Self Attention) for each window to reduce the amount of computation and the number of parameters. The method of Non-Patent Document 2 is called SwinTransformer. Also, in Non-Patent Document 3, a technique has been proposed that uses SwinTransformer to find a tracking target from a search range.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

[0005] However, in the method disclosed in Non-Patent Document 3, after simply combining the template features to be tracked and the search features to be searched in the spatial direction, token mixing such as MSA is performed. Since the processing amount of MSA is proportional to the square of the size in the spatial direction, there is a problem that the processing amount in token mixing increases significantly.

[0006] The present invention has been made in view of such problems, and an object thereof is to provide a technique for reducing the processing amount in token mixing. [Means for Solving the Problems]

[0007] In order to solve the above problems, the information processing apparatus according to the present invention has the following configuration. That is, the information processing apparatus a first generation means for generating a first feature amount from a first image, a second generation means for generating a second feature amount from a second image different from the first image, a combining means for combining the first feature amount and the second feature amount, a dividing means for dividing the combined feature amount into a plurality of groups, a mixing means for mixing the feature amounts included in each group for each of the plurality of groups, and At least one of the plurality of groups includes both the first feature amount and the second feature amount.

Advantages of the Invention

[0008] According to the present invention, a technique for reducing the processing amount in a token mix can be provided.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Embodiments for Carrying Out the Invention

[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential for the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are given the same reference numerals, and redundant explanations are omitted.

[0011] (First Embodiment) As a first embodiment of an information processing apparatus according to the present invention, an information processing apparatus that executes a tracking task using a neural network (NN) for moving image data will be described below as an example.

[0012] <Hardware Configuration of Information Processing Apparatus> FIG. 1 is a diagram showing the hardware configuration of an information processing apparatus 100. The information processing apparatus 100 executes image processing including a tracking task described later. Note that in FIG. 1, it is shown as a single information processing apparatus, but each function may be realized by being distributed among a plurality of information processing apparatuses. When it is configured by a plurality of information processing apparatuses, they are connected by a local area network (LAN) or the like so as to be able to communicate with each other.

[0013] An input device 109, an output device 110, a network 111 such as the Internet, and a camera 112 are connected to the information processing apparatus 100. Note that the connection method is not limited to that shown in FIG. 1. For example, each may be connected separately by wire or via wireless communication. Further, the information processing apparatus 100 and the input device 109 or the output device 110 may be independent devices or integrated devices (for example, a touch panel display).

[0014] The input device 109 is a device for performing user input to the information processing apparatus 100. The input device may be, for example, a pointing device or a keyboard. The output device 110 is a device such as a monitor capable of displaying images and characters in order to display data held by the information processing apparatus 100, data supplied by user input, and execution results of programs. The camera 112 is an imaging device capable of acquiring an imaging image. The camera 112 may acquire continuous imaging images having a predetermined interval Δt, for example, for input to an image acquisition unit 201 described later. Further, the camera 112 is not limited to a monocular camera and may be a compound eye camera.

[0015] The CPU 101 is a central processing unit that controls the entire information processing apparatus 100. The CPU 101 realizes various functions and operations described later in the information processing apparatus 100 by executing various software (computer programs) stored in an external storage device 104, for example. The ROM 102 is a read-only memory that stores programs and parameters that do not require modification. The RAM 103 is a random access memory that temporarily stores programs and data supplied from external devices and the like.

[0016] The external storage device 104 is an external storage device that can be read by the information processing device 100, and stores programs, data, etc. in the long term. The external storage device 104 may be, for example, a hard disk and a memory card fixedly installed in the information processing device 100. Also, for example, the external storage device 104 may be an optical disk such as a flexible disk (FD) or a compact disk (CD) detachable from the information processing device 100, a magnetic or optical card, an IC card, and a memory card, etc.

[0017] The input device interface 105 is an interface with the input device 109. The output device interface 106 is an interface with the output device 110. The communication interface 107 is an interface for connecting to a network 111 such as the Internet or a camera 112, etc. The system bus 108 is a bus that communicably connects each unit within the information processing device 100.

[0018] Note that in FIG. 1, the camera 112 is shown in a form directly connected to the information processing device 100 via the communication interface 107, but it may also be in a form connected to the information processing device 100 via a network 111 or the like. Also, the camera 112 does not have to be one unit, and a plurality of cameras may be connected.

[0019] As described above, the programs that realize various functions and operations are stored in the external storage device 104. When executing a program, the CPU 101 reads the program into the RAM 103. Then, by the CPU 101 executing the program, various functions and operations are realized. Note that although it is assumed that various programs and setting data sets are stored in the external storage device 104, it may be configured to be stored in an external server (not shown). In that case, the information processing device 100 acquires programs and setting data sets from the external server via, for example, the network 111.

[0020] <Functional Configuration of Information Processing Device> FIG. 2 is a diagram showing the functional configuration of the information processing apparatus 100. As described above, in the present embodiment, the information processing apparatus 100 executes a tracking task of detecting a specific tracking target from an image using an NN. In the present embodiment, it is assumed that the tracking task is executed according to the method of Document 1 below. (Document 1) Zhang et al., "Ocean: Object-aware Anchor-free Tracking", ECCV2020, arXiv:2006.10721, 2020

[0021] The information processing apparatus 100 includes an image acquisition unit 201, a tracking target designation unit 202, a feature quantity generation unit 203, a matching unit 204, an identification unit 205, and a post-processing unit 206. These functional units are communicably connected to a storage unit 207. In FIG. 2, the storage unit 207 is shown as existing outside the information processing apparatus 100, but the storage unit 207 may be included in the information processing apparatus 100.

[0022] The image acquisition unit 201 acquires an image (for example, a frame image included in moving image data) obtained by imaging a subject with an imaging device. The subject is an object such as a person or an animal. In the following description, an example of tracking a person will be described.

[0023] The tracking target designation unit 202 determines a tracking target in the image according to an instruction via the input device 109 by the user. As a specific method for determining the tracking target, there is a method of determining the tracking target by touching a subject displayed on the output device 110. Note that, in addition to being specified by the input device 109, the tracking target may be automatically detected and determined as a main subject or the like in the image. As a method for automatically detecting the main subject in the image, for example, the method disclosed in Japanese Patent No. 6556033 can be cited. Further, it may be determined based on both the designation by the input unit 105 and the detection result by the detection process (for example, selecting and designating one from a plurality of detection results). As a technique for detecting an object from an image, the method disclosed in Document 2 below can be cited. (Reference 2) Liu et al., "SSD: Single Shot Multibox Detector", ECCV2016, arXiv:1512.02325, 2015

[0024] The feature generation unit 203 generates features from the image obtained by the image acquisition unit 201. The features are generated using a CNN or the like. Although details will be described later, template features and search image features are generated. The matching unit 204 performs a matching process on the features generated by the feature generation unit 203. The matching process is a process of mixing (mixing) features at the token level and is also called token mixing. Here, the template features and the search image features are token-mixed. Details will be described later with reference to FIG. 4.

[0025] The identification unit 205 infers the features (feature maps) necessary to specify the position and size of the subject in the image based on the token mixing result by the matching unit 204.

[0026] The post-processing unit 206 forms and outputs a bounding box (BB) indicating the position and size of the person who is the subject based on the feature map inferred by the identification unit 205. Task execution such as the tracking task is performed by the identification unit 205 and the post-processing unit 206 and the like.

[0027] <Operation of the information processing device> FIG. 3 is a flowchart showing the flow of the process (tracking task) executed by the information processing device 100. Note that the information processing device 100 does not necessarily perform all the steps described in this flowchart.

[0028] In S301, the image acquisition unit 201 acquires an image in which the object to be tracked exists. FIG. 8(a) is a diagram for explaining a template image. The image 801 is the original image acquired by the image acquisition unit 201. The image 801 includes the object to be tracked 803, and the BB 804 indicates the position and size of the object to be tracked 803. The BB of the object to be tracked is obtained by the object to be tracked designation unit 202.

[0029] In S302, based on the BB 804 of the object to be tracked 803 obtained by the object to be tracked designation unit 202, the image acquisition unit 201 cuts out and resizes the peripheral image 802 of the object to be tracked within the image 801 to generate a template image. For example, it may be cut out at a constant multiple of the size of the object to be tracked centered on the position of the object to be tracked 803.

[0030] In S303, the feature quantity generation unit 203 extracts features from the template image generated in S302 to generate a template feature quantity. The feature quantity can be generated using, for example, a CNN.

[0031] FIG. 9 is a flowchart showing the processing by the CNN. Further, FIG. 10 is a diagram showing the processing of each layer of the CNN. The CNN is composed of one or more layers, and here, as an example, a CNN composed of three layers 901 to 903 is shown. Each layer is composed of a convolution operation (Convolution) and a non-linear operation (such as ReLU). There may be a plurality of these elements, or there may be no non-linear operation such as ReLU. Also, Average Pooling and Max Pooling may be combined. However, the generation of the feature quantity in the specific example in the feature quantity generation unit 203 is not limited to a specific method.

[0032] In S304, the image acquisition unit 201 acquires a search image that is the object to be searched for the object to be tracked. FIG. 8(b) is a diagram for explaining the search image. The image 805 is the original image acquired by the image acquisition unit 201. The image 805 includes the object to be tracked 807.

[0033] In S305, the image acquisition unit 201 cuts out and resizes the image 806 around the tracking target in the image 805 to obtain a search image. For example, it may be cut out at a constant multiple of the size of the tracking target centered on the position of the tracking target in the previous frame.

[0034] In S306, the feature quantity generation unit 203 extracts features from the search image generated in S305 to generate a search image feature quantity. The search image feature quantity can be generated in the same manner as the template feature quantity (S303).

[0035] In S307, the matching unit 204 matches the template feature quantity obtained in S303 and the search image feature quantity obtained in S306. As described above, matching means token mixing. FIG. 4 is a detailed flowchart of the matching process (S307).

[0036] In S401, the matching unit 204 alternately combines the template feature quantity and the search image feature quantity in the spatial direction.

[0037] FIG. 5 is a diagram for explaining the combination and group division of feature quantities. In FIG. 5, the vertical direction and the horizontal direction represent the "spatial direction", and the depth direction represents the "channel direction". FIG. 5(a) shows the template feature quantity 501 and the search image feature quantity 502 respectively. When these are alternately combined in the spatial direction (for example, the vertical direction in FIG. 5), the feature quantity shown in FIG. 5(b) (alternately combined in the spatial direction) is obtained. Note that the combination is not limited to the vertical direction, and the combination in the horizontal direction may also be used.

[0038] In S402, the matching unit 204 divides the feature quantity alternately combined in the spatial direction into predetermined groups. FIG. 5(c) is a diagram showing an example of the feature quantity divided into groups. Here, it is divided into two rectangles in the vertical direction to generate the feature quantity 504 belonging to group 1 and the feature quantity 505 belonging to group 2 respectively.

[0039] Note that the group division is not limited to the division shown in Fig. 5(c). Any group division is acceptable as long as both the template feature amount and the search image feature amount are included in one or more of the divided groups. Fig. 7 is a diagram for explaining another example of the group division of feature amounts. For example, it may be divided into three groups such as feature amounts 701 to 703.

[0040] In S403, the matching unit 204 performs token mixing for each of the divided groups. As the token mixing method, there are an MLP in which a plurality of fully-connected (FC) layers as shown in Fig. 11 are connected in series, and an MSA.

[0041] Note that the token mixing by the MSA is expressed by the following mathematical formula (1). Q, K, and V are obtained through FC for the features divided into groups.

[0042]

Equation

[0043] In S404, the matching unit 204 restores the group division of the features token-mixed for each group in S403 (returns to the spatial position before division).

[0044] In S405, the matching unit 204 performs token mixing in the channel direction (for example, the depth direction in Fig. 5). The token mixing in the channel direction can be realized by a CNN as shown in Fig. 9.

[0045] In S406, the matching unit 204 releases the combination of the template feature amount and the search image feature amount combined in the spatial direction (returns to the spatial position before combination).

[0046] Through the processing of the above token mix (S401 to S406), the feature amounts are grouped, so that it becomes possible to reduce the processing amount (computation amount). In particular, both the template feature amount and the search image feature amount will be included in the same group. Note that the processing of the token mix may be repeated multiple times.

[0047] In S308, the identification unit 205 generates feature amounts (likelihood map, BB map) necessary for identifying the position and size of the person who is the tracking target from the search image feature amount output in S307. Note that the identification unit 205 is composed of a CNN as shown in FIG. 9.

[0048] FIG. 12 is a diagram for explaining an input image and the outputs (likelihood map, BB map) of the CNN. The CNN is trained to generate feature amounts as shown in the likelihood map 1203 and the BB map 1211 for the input image 1201 (FIG. 12(a)). The likelihood map 1203 is a map representing the existence likelihood of the tracking target at each position, and the BB map 1211 is a map representing the BB of the tracking target at each position.

[0049] In S309, the post-processing unit 206 specifies the BB of the person who is the tracking target based on the feature amounts obtained in S308. The method for specifying the BB of an object from the output of the neural network is described in detail in the following Literature 3 and the like. (Literature 3) Tian et al., "FCOS: Fully Convolutional One-Stage Object Detection", arXiv:1904.01355, 2019

[0050] Figure 12(b) shows the likelihood map 1203 inferred by the identification unit 205 for the input image 1201. In the likelihood map 1203, for each grid region, a large value is output for the region where the tracking target exists, and a small value is output for other regions. In Figure 12(b), a large value is output for the grid region 1204 (for example, compared with a threshold value), suggesting that a person (tracking target) exists at this position. On the other hand, a small value is output for the grid region 1205, suggesting that no person exists at this position.

[0051] Figure 12(c) shows the BB map 1211 inferred by the identification unit 205 for the input image 1201. In the BB map 1211, for each grid region, a value indicating the distance from the center of the grid region to the upper, lower, left, and right ends of the person is output. For example, in the grid region 1206 corresponding to the grid region 1204 of the likelihood map 1203, the distance 1207 to the upper end of the person, the distance 1208 to the right end, the distance 1209 to the lower end, and the distance 1210 to the left end are output. Thus, the BB corresponding to the person 1202 can be formed. The learning method of the tracking task is detailed in the above-mentioned reference 1.

[0052] As described above, according to the first embodiment, prior to token mixing the template feature amount and the search image feature amount, the template feature amount and the search image feature amount are alternately combined in the spatial direction and grouped in the spatial direction. By this process, it is possible to reduce the size in the spatial direction in each group, so that it is possible to reduce the processing amount of token mixing. Also, it is possible to perform token mixing in each group with both the template feature amount and the search image feature amount included.

[0053] (Modification Example 1-1) In Modification 1-1, another example of the matching process (S307) will be described. FIG. 13 is a flowchart showing the matching process in Modification 1-1. The processes of S401 to S406 are the same as those in the first embodiment (FIG. 4), but the difference is that S1302 is performed prior to S402 (that is, after the combination in S401).

[0054] In S1302, the matching unit 204 rearranges the feature amounts. FIG. 6 is a diagram for explaining the rearrangement of the feature amounts. The feature amount 503 shown in FIG. 6(a) is the same as that in FIG. 5(b), which is obtained by alternately combining the template feature amount 501 and the search image feature amount 502 in the spatial direction in S401.

[0055] The feature amount 602 shown in FIG. 6(b) is obtained by cyclically shifting (rearranging) only the template feature amount (the part shown in white) in the vertical direction with respect to the feature amount 502. In the subsequent S402, group division is performed on the feature amount 602. Note that the rearrangement of the feature amounts is not limited to the cyclic shift in the vertical direction. It may be rearranged in the horizontal direction or irregularly.

[0056] Also, the feature amount to be rearranged may be the search image feature amount, or both the template feature amount and the search image feature amount may be rearranged. That is, it is sufficient if the result of the group division (S402) is different from the result when the feature amounts are not rearranged.

[0057] Also, the rearrangement (S1302) may be configured to be performed prior to the combination in the spatial direction (S401) (that is, before the combination). After the group division, it is sufficient if both the template feature amount and the search image feature amount are included in at least one or more groups.

[0058] Furthermore, when the processes of S401 to S405 are repeated over a plurality of layers, the method of rearranging the feature amounts may be changed for each layer. Thereby, it becomes possible to calculate various combinations of the template feature amount and the search image feature amount.

[0059] When the feature amounts are not rearranged, matching will contribute only to feature amounts that are spatially close to each other. The more efficiently all the feature amounts are mixed, the easier it becomes to recognize various patterns, and as a result, the recognition accuracy will be improved. That is, by rearranging the feature amounts, it becomes possible to perform matching even between spatially separated positions, and as a result, it becomes possible to improve the tracking accuracy.

[0060] (Modification Example 1-2) In Modification Example 1-2, still another example of the matching process (S307) will be described. FIG. 14 is a flowchart showing the matching process in Modification Example 1-2. Note that the processes of S401 to S406 are the same as those in the first embodiment (FIG. 4), but the difference is that S1401 is performed prior to S401 and S1402 can be performed prior to S405.

[0061] In S1401, the matching unit 204 reduces the spatial resolution of the template feature amount and the search image feature amount by downsampling. FIG. 15 is a diagram for explaining downsampling.

[0062] The feature amount 1501 shown in FIG. 15(a) is the feature amount before downsampling. Here, the feature amounts arranged in the spatial direction are rearranged in the channel direction. In this case, as shown in FIG. 15(b), while reducing the resolution in the spatial direction to 1 / 2 in both the vertical and horizontal directions, the resolution in the channel direction is increased by 4 times. The resulting feature amounts after increasing by 4 times in the channel direction become feature amounts 1502 to 1505. The processes of S401 to S404 are performed on the downsampled feature amounts.

[0063] In S1402, the matching unit 204 upsamples to the original resolution. Note that for upsampling, the reverse operation of downsampling may be performed. That is, the features of the feature amounts 1502 to 1505 may be rearranged again like the feature amount 1501.

[0064] As described above, the processing of token mixing in the spatial direction (S403) is proportional to the square of the spatial resolution. Therefore, it is possible to reduce the computational amount by performing the processing of S402 to S404 after reducing the spatial resolution by downsampling.

[0065] (Second Embodiment) As a second embodiment of the information processing apparatus according to the present invention, in this embodiment, an information processing apparatus that executes a disparity estimation task using a neural network (NN) for binocular images obtained by a stereo camera will be described below as an example. Note that the hardware configuration of the information processing apparatus is the same as that of the first embodiment (FIG. 1), and thus the description thereof will be omitted.

[0066] <Functional Configuration of Information Processing Apparatus> FIG. 16 is a diagram showing the functional configuration of the information processing apparatus 100 in the second embodiment. In this embodiment, it is assumed that a disparity estimation task is executed according to the method of the following Document 4. (Document 4) Liang et al., "Learning for Disparity Estimation through Feature Constancy", arXiv:1712.01039, 2017

[0067] The information processing apparatus 100 includes an image acquisition unit 1601, a feature quantity generation unit 203, a matching unit 204, and a refinement unit 1602. These functional units are communicably connected to a storage unit 207. The image acquisition unit 1601 acquires a right-eye image and a left-eye image of the stereo camera, respectively. The refinement unit 1602 refines the feature quantity obtained from the matching unit 204 and generates a disparity image.

[0068] <Operation of Information Processing Apparatus> FIG. 17 is a flowchart showing the flow of processing (disparity estimation task) executed by the information processing apparatus 100. Note that the information processing apparatus 100 does not necessarily perform all the steps described in this flowchart.

[0069] In S1701, the image acquisition unit 1601 acquires the right-eye image obtained from the stereo camera. In S303, the feature quantity generation unit 203 generates the feature quantity of the right-eye image. In S1702, the image acquisition unit 1601 acquires the left-eye image obtained from the stereo camera. In S306, the feature quantity generation unit 203 generates the feature quantity of the left-eye image. Note that the method for generating the feature quantity is the same as that in the first embodiment.

[0070] In S307, the matching unit 204 matches the right-eye image feature quantity and the left-eye image feature quantity. Similar to the first embodiment, matching means token mixing.

[0071] In S1703, the refinement unit 1602 refines the feature quantity obtained from S307 to obtain a disparity image. The refinement process can be performed by combining Convolution and upsampling as shown in FIG. 10. Note that the learning method for disparity estimation by the stereo camera is detailed in the above-mentioned Document 4.

[0072] As described above, according to the second embodiment, prior to token mixing the right-eye image feature quantity and the left-eye image feature quantity, the right-eye image feature quantity and the left-eye image feature quantity are alternately combined in the spatial direction and grouped in the spatial direction. By this process, it becomes possible to reduce the size in the spatial direction in each group, and thus it becomes possible to reduce the processing amount of token mixing. Also, it becomes possible to perform token mixing in each group with both the right-eye image feature quantity and the left-eye image feature quantity included.

[0073] <Third Embodiment> As understood from the first and second embodiments, the present invention can be applied to any process for comparing two images. For example, it can also be applied to a face authentication task by an authentication system. In that case, as the two images, the first face image obtained at the time of authentication processing and the second face image registered in advance are used.

[0074] The feature quantity generation unit 203 generates the feature quantities of each of the two face images. Then, the matching unit 204 matches the feature quantity of the first face image with the feature quantity of the second face image.

[0075] After the matching, Global Average Pooling and the MLP shown in FIG. 11 may be performed to output whether the first face image and the second face image are of the same person.

[0076] The disclosure of this specification includes the following information processing apparatus, control method, and program. (Item 1) A first generation means for generating a first feature quantity from a first image, A second generation means for generating a second feature quantity from a second image different from the first image, A combining means for combining the first feature quantity and the second feature quantity, A dividing means for dividing the combined feature quantity into a plurality of groups, For each of the plurality of groups, a mixing means for mixing the feature quantities included in each group for each group, Comprising At least one of the plurality of groups includes both the first feature quantity and the second feature quantity An information processing apparatus characterized by this. (Item 2) The combining means combines the first feature quantity and the second feature quantity alternately in the spatial direction The information processing apparatus according to Item 1, characterized by this. (Item 3) The combining means combines the first feature quantity and the second feature quantity irregularly in the spatial direction The information processing apparatus according to Item 1, characterized by this. (Item 4) The mixing means A first means for mixing the feature quantities included in each of the plurality of groups for each group in the spatial direction, A second means for returning the position of the feature amount in the spatial direction obtained by the first means to the position in the spatial direction before the division by the division means; A third means for mixing the feature amounts obtained by the second means in the channel direction; A fourth means for returning the position of the feature amount in the spatial direction obtained by the third means to the position in the spatial direction before the combination by the combination means; and having The information processing apparatus according to any one of Items 1 to 3, characterized in that. (Item 5) The information processing apparatus further includes a shift means for performing a cyclic shift in the spatial direction with respect to at least one of the first feature amount and the second feature amount before or after the combination by the combination means. The information processing apparatus according to any one of Items 1 to 4, characterized in that. (Item 6) The mixing means includes a fully connected layer or an MSA (Multi-head Self Attention). The information processing apparatus according to any one of Items 1 to 5, characterized in that. (Item 7) The information processing apparatus further includes a task execution means for executing a predetermined task using a neural network (NN) based on the feature amount obtained by the mixing means. The information processing apparatus according to any one of Items 1 to 6, characterized in that. (Item 8) The first image is a first frame image included in moving image data, The second image is a second frame image included in the moving image data, The predetermined task is a tracking task The information processing apparatus according to Item 7, characterized in that. (Item 9) The first image is a right-eye image obtained from a stereo camera, The second image is a left-eye image obtained from the stereo camera, The predetermined task is a disparity estimation task The information processing apparatus according to Item 7, characterized in that. (Item 10) The first image is a first face image obtained during authentication processing by the authentication system, the second image is a second face image pre-registered in the authentication system, and the predetermined task is a face authentication task The information processing apparatus according to item 7, characterized in that (Item 11) A control method for an information processing apparatus, comprising: a first generation step of generating a first feature amount from a first image; a second generation step of generating a second feature amount from a second image different from the first image; a combining step of combining the first feature amount and the second feature amount; a dividing step of dividing the combined feature amount into a plurality of groups; for each of the plurality of groups, a mixing step of mixing the feature amounts included in each group for each group; including at least one of the plurality of groups includes both the first feature amount and the second feature amount A control method characterized by the above. (Item 12) A program for causing a computer to execute the control method according to item 11.

[0077] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. Further, it can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0078] The invention is not limited to the above embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Therefore, the claims are attached to disclose the scope of the invention.

Explanation of Symbols

[0079] 100 Information processing apparatus; 201 Image acquisition unit; 202 Tracking target designation unit; 203 Feature quantity generation unit; 204 Matching unit; 205 Identification unit; 206 Post-processing unit; 207 Storage unit

Claims

1. A first generation means for generating a first feature amount from a first image, A second generation means for generating a second feature amount from a second image different from the first image, Combining means for combining the first feature amount and the second feature amount, Dividing means for dividing the combined feature amount into a plurality of groups, Mixing means for mixing the feature amounts included in each of the plurality of groups for each group, Comprising, At least one of the plurality of groups includes both the first feature amount and the second feature amount An information processing apparatus characterized by this.

2. The combining means alternately combines the first feature amount and the second feature amount in the spatial direction The information processing apparatus according to claim 1, characterized by this.

3. The combining means irregularly combines the first feature amount and the second feature amount in the spatial direction The information processing apparatus according to claim 1, characterized by this.

4. The mixing means, First means for mixing the feature amounts included in each of the plurality of groups for each group in the spatial direction, Second means for returning the spatial position of the feature amount obtained by the first means to the spatial position before division by the dividing means, Third means for mixing the feature amount obtained by the second means in the channel direction, Fourth means for returning the spatial position of the feature amount obtained by the third means to the spatial position before combination by the combining means, Having The information processing apparatus according to claim 1, characterized by this.

5. Further comprising shift means for cyclically shifting in the spatial direction with respect to at least one of the first feature amount and the second feature amount before or after combination by the combining means The information processing apparatus according to claim 1, characterized by this.

6. The mixing means includes a fully connected layer or MSA (Multi-head Self Attention) The information processing apparatus according to claim 1, characterized by this.

7. Further comprising task execution means for executing a predetermined task using a neural network (NN) based on the feature amount obtained by the mixing means The information processing apparatus according to claim 1, characterized by this.

8. The first image is a first frame image included in moving image data, The second image is a second frame image included in the moving image data, The predetermined task is a tracking task The information processing apparatus according to claim 7, characterized in that...

9. The first image is a right-eye image obtained from a stereo camera, The second image is a left-eye image obtained from the stereo camera, The predetermined task is a disparity estimation task The information processing apparatus according to claim 7, characterized in that...

10. The first image is a first face image obtained during authentication processing by an authentication system, The second image is a second face image registered in advance in the authentication system, The predetermined task is a face authentication task The information processing apparatus according to claim 7, characterized in that...

11. A control method for an information processing apparatus, comprising: A first generation step of generating a first feature amount from a first image; A second generation step of generating a second feature amount from a second image different from the first image; A combination step of combining the first feature amount and the second feature amount; A division step of dividing the combined feature amount into a plurality of groups; For each of the plurality of groups, a mixing step of mixing the feature amounts included in each group for each group; Including At least one of the plurality of groups includes both the first feature amount and the second feature amount A control method characterized by the above.

12. A program for causing a computer to execute the control method according to claim 11.