Video processing device, video processing method, and program

The video processing apparatus enhances feature transformation by incorporating identification element extraction and attention mechanisms, addressing the limitations of existing neural network-based video recognition systems to improve task-specific accuracy.

JP7859530B2Active Publication Date: 2026-05-15NEC CORP
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC CORP
Filing Date
2022-12-28
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing video recognition using neural networks lacks optimal feature transformation and fails to explicitly encode identification elements, limiting the effectiveness of downstream recognition tasks.

Method used

A video processing apparatus and method that includes video acquisition, feature map conversion, identification element extraction, and a discriminant element reference non-local layer to generate optimal feature maps by focusing on identification elements, using attention mechanisms to enhance feature transformation.

Benefits of technology

Enables optimal feature transformation considering identification elements, improving the accuracy and effectiveness of video recognition tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007859530000003
    Figure 0007859530000003
  • Figure 0007859530000004
    Figure 0007859530000004
  • Figure 0007859530000005
    Figure 0007859530000005
Patent Text Reader

Abstract

In this video processing device, a video acquisition means acquires a video. A plurality of sequentially arranged conversion means converts a feature map of an input video and outputs a converted feature map. An identification element extraction means detects an identification element from the video and outputs an identification element feature quantity. An identification element reference conversion means is disposed between the plurality of conversion means and generates and outputs a feature map on the basis of an input feature map and the identification element feature quantity.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to feature extraction from video. [Background technology]

[0002] In video recognition technology using neural networks (NN), a method of introducing non-local layers into the intermediate layers (Non-local Neural Networks) has been proposed to improve accuracy (see Non-Patent Document 1). Furthermore, Patent Document 1 proposes a feature conversion device and an image recognition device that improve upon the technology related to the above-mentioned method of introducing non-local layers into the intermediate layers. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] International Publication No. WO2021-176566 [Non-patent literature]

[0004] [Non-Patent Document 1] Non-Local Neural Networks, Xiaolong Wang, Ross Girshick, Abhinav Gupta, Kaiming He Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7794-7803 [Overview of the project] [Problems that the invention aims to solve]

[0005] In video recognition using a neural network (NN), it is necessary to perform feature transformation, extract identification elements included in the video, and perform recognition based on the extracted identification elements. Thus, in video recognition using an NN, since the identification elements are only used in downstream recognition tasks, it has not been possible to perform optimal feature transformation based on the identification elements. Also, in the method of Non-Patent Document 1, since the non-local layer does not explicitly encode the identification elements, optimal feature transformation has not been achieved.

[0006] One object of the present disclosure is to enable optimal feature transformation in consideration of identification elements corresponding to a target task.

Means for Solving the Problem

[0007] In one aspect of the present disclosure, a video processing apparatus includes video acquisition means for acquiring a video, conversion means for converting a feature map of the video and outputting a converted feature map, identification element extraction means for detecting an identification element from the video and From the aforementioned identification element outputting an identification element feature amount Extract Features A means for focusing on a discriminant element to generate a discriminant element response based on the converted feature map and the discriminant element feature quantities, A means for converting the aforementioned discriminant element response into a first feature map with the same dimension as the converted feature map, A video feature focusing means generates a video feature response focused on video features by obtaining, for each point in the transformed feature map, a video feature response based on the relationship between each point in the transformed feature map as output, based on the transformed feature map; A response value conversion means that converts the aforementioned video feature response into a second feature map with the same dimension as the converted feature map, Adding means for generating a third feature map by adding the first feature map and the second feature map to the converted feature map, and includes. [[ID=?]] [[ID=?]]<? [[ID=?]]

[0008] [[ID=?]] In another aspect of the present disclosure, a video processing method is executed by a computer and includes acquiring a video, converting a feature map of the video and outputting a converted feature map, detecting an identification element from the video, outputting an identification element feature amount, From the aforementioned identification element and based on the converted feature map and the identification element feature amount, , generate an identification element response that focuses on the aforementioned identification element, ​​​ The aforementioned discriminant element response is transformed into a first feature map having the same dimension as the transformed feature map. Based on the transformed feature map, a video feature response is obtained as output for each point in the transformed feature map, based on its relationship with each other in the transformed feature map, thereby generating the video feature response that focuses on video features. The aforementioned video feature response is transformed into a second feature map having the same dimension as the transformed feature map. A third feature map is generated by adding the first and second feature maps to the converted feature map. .

[0009] In yet another aspect of this disclosure, the program is Obtain the video, The feature map of the aforementioned video is transformed and the transformed feature map is output. Identifying elements are detected from the aforementioned video, From the aforementioned identification element Output the discriminant feature quantities, Based on the transformed feature map and the discriminant element features, , generate an identification element response that focuses on the aforementioned identification element, The aforementioned discriminant element response is transformed into a first feature map having the same dimension as the transformed feature map. Based on the transformed feature map, a video feature response is obtained as output for each point in the transformed feature map, based on its relationship with each other in the transformed feature map, thereby generating the video feature response that focuses on video features. The aforementioned video feature response is transformed into a second feature map having the same dimension as the transformed feature map. A third feature map is generated by adding the first and second feature maps to the converted feature map. Have the computer perform the process. [Effects of the Invention]

[0010] According to this disclosure, it becomes possible to perform optimal feature transformation by taking into account the identification elements corresponding to the target task. [Brief explanation of the drawing]

[0011] [Figure 1] This is a block diagram showing the schematic configuration of a video processing device according to the first embodiment. [Figure 2] This is a block diagram showing the hardware configuration of a feature conversion device. [Figure 3] This is a block diagram showing the functional configuration of a feature conversion device. [Figure 4] This block diagram shows the functional configuration of the nonlocal layer of the identification element reference. [Figure 5] This is a block diagram showing the configuration of the identification element reference nonlocal layer according to modified example 2. [Figure 6] This is a block diagram showing another configuration of the identification element reference nonlocal layer according to Modification 2. [Figure 7] This block diagram shows the functional configuration of the feature conversion device in modified example 3. [Figure 8] This is a block diagram showing the functional configuration of the video processing apparatus according to the second embodiment. [Figure 9] This is a flowchart of the processing performed by the video processing device of the second embodiment. [Modes for carrying out the invention]

[0012] Preferred embodiments of this disclosure will be described below with reference to the drawings. <First Embodiment> [System Configuration] Figure 1 shows the schematic configuration of the video processing device 1. The video processing device 1 consists of a feature conversion device 100 and a prediction device 5. The feature conversion device 100 generates a feature map from the input video and outputs it to the prediction device 5. Based on the feature map input from the feature conversion device 100, the prediction device 5 recognizes objects contained in the video and outputs the result.

[0013] The feature conversion device 100 of this embodiment is characterized by generating a feature map using identification elements. Identification elements are elements necessary for identifying an object and are elements that contain information necessary for performing tasks performed downstream of the feature conversion device 100, which in this embodiment is the prediction performed by the prediction device 5. For example, in the case of recognizing human behavior, the identification elements would be the target person included in the video or objects related to the target person's behavior. The feature conversion device 100 detects identification elements from the input video and extracts the feature quantities of the identification elements. Then, the feature conversion device 100 generates a feature map using the feature quantities of the identification elements and outputs it to the prediction device 5.

[0014] [Hardware configuration] Figure 2 is a block diagram showing the hardware configuration of the feature conversion device 100 in the first embodiment. As shown in the figure, the feature conversion device 100 includes an interface (I / F) 11, a processor 12, a memory 13, a recording medium 14, and a database (DB) 15.

[0015] I / F11 performs data input and output with external devices. Specifically, when video used by the feature conversion device 100 is input via external communication, I / F11 receives the input data. I / F11 is also used when the feature conversion device 100 outputs the feature map it has generated to an external device.

[0016] The processor 12 is a computer such as a CPU (Central Processing Unit) and controls the entire feature conversion device 100 by executing a pre-prepared program. The processor 12 may also be a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array).

[0017] Memory 13 consists of ROM (Read Only Memory), RAM (Random Access Memory), and other components. Memory 13 is also used as working memory while the processor 12 is executing various processes.

[0018] The recording medium 14 is a non-volatile, non-temporary recording medium such as a disk-shaped recording medium or semiconductor memory, and is configured to be detachable from the feature conversion device 100. The recording medium 14 stores various programs that the processor 12 executes. When the feature conversion device 100 performs various processes, the programs stored on the recording medium 14 are loaded into the memory 13 and executed by the processor 12.

[0019] DB15 stores data input and output via I / F11 as needed.

[0020] [Functional Configuration] FIG. 3 is a block diagram showing a functional configuration of the feature conversion device 100 according to the first embodiment. Functionally, the feature conversion device 100 includes conversion layers 21a to 21n, an identification element extractor 22, and an identification element reference non-local layer 23. In the following description, when the individual conversion layers 21a to 21n are not distinguished, they may be simply referred to as "conversion layer 21".

[0021] A video is input to the feature conversion device 100 from an external device through I / F 11. The video is input to the conversion layer 21a and the identification element extractor 22. Note that the video is a tensor of C i ×T i ×H i ×W i . "C" i indicates the number of channels. When the input video has three RGB channels, "C" i is 3. Also, "T" i indicates time, "H" i indicates the height (vertical pixel count) of the image, and "W" i indicates the width (horizontal pixel count) of the image.

[0022] The conversion layer 21 includes, for example, a convolutional layer, a residual block layer, a pooling layer, etc., and converts the input feature map to output a converted feature map. Specifically, the conversion layers 21a to 21n are arranged continuously. The conversion layer 21a performs processing such as convolution on the input video and outputs the generated feature map to the subsequent conversion layer 21b. The conversion layer 21b performs processing such as convolution on the input feature map and outputs the generated feature map to the subsequent conversion layer 21c. Then, the feature maps generated through the processing in each conversion layer are output to the prediction device 5.

[0023] The identification element extractor 22 detects identification elements from the input video and extracts their feature amounts (hereinafter also referred to as "identification element feature amounts"). One type of identification element is set in the identification element extractor 22 in advance. When the identification element extractor 22 detects K identification elements corresponding to that one type from the input video, it is a vector value of C D dimensions for the identification element, CD Extract ×K). C D The dimension is a predetermined number of dimensions. For example, if a cart is set as the identification element and the input video contains three carts, the identification element extractor 22 extracts the identification element features (C) of each cart. D The (x3) is extracted. The discriminant element extractor 22 then outputs the extracted discriminant element features to the discriminant element reference nonlocal layer 23.

[0024] For example, if the identification element is a person, the identification element extractor 22 outputs the person's coordinates, size, joint point coordinates, and features of how the person appears as identification element features. If the identification element is an object, the identification element extractor 22 extracts the object's coordinates, size, type, posture information, and features of how the object appears as identification element features. Furthermore, if the identification element is an interaction between a person and an object, the identification element extractor 22 extracts the coordinates of the interacting elements (i.e., the person and the object), the types of those elements, and the type of interaction as identification element features.

[0025] The discriminant element reference nonlocal layer 23 is positioned at any location between the transformation layers 21a to 21n. The discriminant element reference nonlocal layer 23 generates a feature map based on the feature map input from the preceding transformation layer 21 and the discriminant element features input from the discriminant element extractor 22, and outputs the generated feature map to the subsequent transformation layer 21.

[0026] Figure 4 shows the configuration of the discriminant element reference nonlocal layer 23. The discriminant element reference nonlocal layer 23 includes a video feature attention unit 30, a resolution conversion unit 32, a response value conversion unit 37, a discriminant element attention unit 40, a discriminant element response value conversion unit 47, and an addition unit 48. The video feature attention unit 30 is a type of attention device equipped with an attention mechanism. The discriminant element reference nonlocal layer 23 enables optimal feature map conversion by adding the discriminant element attention unit 40, which is an attention mechanism for discriminant elements, to the video feature attention unit 30. An attention device is a device that takes query elements, key elements, and value elements as inputs and outputs response features. More specifically, an attention device is a device equipped with a mechanism that controls the values ​​to be taken in (generated from value elements) based on the similarity between each query generated from the query element and the key generated from the key element.

[0027] The discriminant element reference nonlocal layer 23 receives a feature map from the transformation layer 21. The feature map is a tensor of type C×T×H×W. The discriminant element reference nonlocal layer 23 also receives K discriminant element features (C) from the discriminant element extractor 22. D ×K) is entered.

[0028] First, the video feature attention unit 30 will be explained. Based on the feature maps input from the conversion layer 21 and the resolution conversion unit 32, the video feature attention unit 30 generates response features (hereinafter referred to as "video feature response") that focus on the video features indicated by the feature maps, and outputs them to the response value conversion unit 37. Specifically, the video feature attention unit 30 includes a query generation unit 31, a video feature key generation unit 33, a video feature value generation unit 34, and a weighted sum calculation unit 36.

[0029] The query generation unit 31 receives the feature map from the transformation layer 21. The query generation unit 31 generates a query using the input feature map (C×T×H×W) as query elements. Specifically, the query generation unit 31 uses the C-dimensional vector of each point in the feature map as C q Dimensional queries (C qA dimensional transformation is performed to convert to ×THW). This dimensional transformation can be performed, for example, by a linear transformation, or by a combination of a linear transformation, an activation function, and sheath normalization. C q The dimension is a predetermined number of dimensions. The query generation unit 31 transforms the query so that its dimension matches the number of dimensions of the key, which will be described later. Since the feature map has THW points, the number of queries is THW. The query generation unit 31 outputs the generated queries to the video feature weight generation unit 35 and the discrimination element weight generation unit 45.

[0030] The resolution conversion unit 32 receives the feature map from the conversion layer 21. The resolution conversion unit 32 converts the resolution of the input feature map to a predetermined value. The resolution conversion unit 32 outputs the feature map with the converted resolution (C×T'×H'×W') to the video feature key generation unit 33 and the video feature value generation unit 34.

[0031] The video feature key generation unit 33 generates keys using the feature map (C×T'×H'×W') input from the resolution conversion unit 32 as key elements. Specifically, the video feature key generation unit 33 uses the C-dimensional vector of each point in the feature map as C q Dimensional key (C q The dimension is converted to ×T'H'W'). This dimension transformation can be performed, for example, by a linear transformation, or by a combination of a linear transformation, an activation function, and layer normalization. Since the feature map has T'H'W' points, the number of keys will be T'H'W'. The video feature key generation unit 33 outputs the generated keys to the video feature weight generation unit 35.

[0032] The video feature value generation unit 34 generates a value using the feature map (C×T'×H'×W') input from the resolution conversion unit 32 as value elements. Specifically, the video feature value generation unit 34 uses the C-dimensional vector of each point in the feature map as C v Dimensional Value (C v This transforms to ×T'H'W'). This dimensional transformation can be performed, for example, by a linear transformation, or by a combination of a linear transformation, an activation function, and sheath normalization. C vThe dimension is a predetermined number of dimensions. Since the feature map has T'H'W' points, the number of values ​​is T'H'W'. The video feature value generation unit 34 outputs the generated values ​​to the weighted sum calculation unit 36.

[0033] Note that the resolution conversion unit 32 may be omitted. In this case, the video feature key generation unit 33 and the video feature value generation unit 34 only need to generate keys and values ​​from the feature map (C×T×H×W) input from the conversion layer 21.

[0034] The video feature weight generation unit 35 receives queries from the query generation unit 31 and keys from the video feature key generation unit 33. The video feature weight generation unit 35 generates weights for each value based on the similarity between the queries and their corresponding keys. Since there are THW queries and T'H'W' keys, THW × T'H'W' weights are generated.

[0035] Specifically, the video feature weight generation unit 35 generates weights using the dot product of the query and the key. In one example, C is created by arranging THW queries. q Let the ×THW matrix be "Q", and then arrange T'H'W' keys in C. q Let the ×T'H'W' matrix be "K", and the weight matrix of the output THW×T'H'W' be "A". Then the weight matrix A is given by the following equation.

[0036]

number

[0037]

number

[0038] The weighted sum calculation unit 36 ​​receives values ​​from the video feature value generation unit 34 and weights from the video feature weight generation unit 35. Based on the values ​​and weights, the weighted sum calculation unit 36 ​​generates a video feature response. For example, the weighted sum calculation unit 36 ​​calculates the values ​​(C v By taking the matrix product of (×T'H'W') and the weights (THW×T'H'W'), the weighted sum of values ​​is calculated, and the video feature response (C v The ×THW) is generated. The weighted sum calculation unit 36 ​​outputs the generated video feature response to the response value conversion unit 37.

[0039] The response value conversion unit 37 converts the video feature response input from the weighted sum calculation unit 36 ​​to the same dimension as the feature map input from the conversion layer 21. Specifically, the response value conversion unit 37 calculates the C for each query. v Dimensional video feature response (C v The ×THW) is converted to the same C dimension as the feature map input from the conversion layer 21 and aligned into a C×T×H×W type tensor. The response value conversion unit 37 then outputs the aligned C×T×H×W type tensor to the addition unit 48 as the converted video feature response (hereinafter also referred to as the "converted video feature response").

[0040] Next, the identification element observation unit 40 will be described. The identification element observation unit 40 includes a query generation unit 31, an identification element key generation unit 43, an identification element value generation unit 44, an identification element weight generation unit 45, and an identification element weighted sum calculation unit 46.

[0041] The identification element key generation unit 43 receives K identification element features from the identification element extractor 22. The identification element key generation unit 43 generates an identification element key using the input identification element features as key elements. Specifically, the identification element key generation unit 43 generates a C D Dimensional discriminant feature(C) D ×K) is the same C as the query generated by the query generation unit 31. q Dimensional identification element key (C qThe dimensional transformation is converted to ×K). This dimensional transformation can be performed, for example, by a linear transformation, or by a combination of a linear transformation, an activation function, and sheath normalization. The discrimination element key generation unit 43 outputs the generated discrimination element key to the discrimination element weight generation unit 45.

[0042] The identification element value generation unit 44 receives K identification element features from the identification element extractor 22. The identification element value generation unit 44 generates an identification element value using the input identification element features as value elements. Specifically, the identification element value generation unit 44 generates a C D Dimensional discriminant feature(C) D ×K) to C Dv Dimensional identification element value (C Dv Convert to ×K). This dimensional transformation can be performed, for example, by a linear transformation, or by a combination of a linear transformation, an activation function, and sheath normalization. C Dv The dimensions are predetermined to a set number of dimensions. The identification element value generation unit 44 outputs the generated identification element values ​​to the identification element weighted sum calculation unit 46.

[0043] The identification element weight generation unit 45 receives a query from the query generation unit 31 and an identification element key from the identification element key generation unit 43. The identification element weight generation unit 45 generates the query (C q ×THW) and the identification element key (C q Based on the similarity with ×K, the identification element weights (THW×K) are generated. The identification element weight generation unit 45 outputs the generated identification element weights to the identification element weight sum calculation unit 46. For example, the identification element weight generation unit 45 can calculate the identification element weights using the inner product of the query and the identification element key, similar to the video feature weight generation unit 35 described above.

[0044] The identification element weighted sum calculation unit 46 receives the identification element value from the identification element value generation unit 44 and the identification element weight from the identification element weight generation unit 45. The identification element weighted sum calculation unit 46 calculates the identification element value (C Dv By taking the matrix product of (×K) and the discriminant weights (THW×K), a weighted sum of discriminant values ​​is calculated for each query, and the discriminant response (CDv It generates ×THW). Specifically, the identification element weighted sum calculation unit 46 calculates a weighted sum of the K identification element values ​​according to the identification element weights for each query, and C Dv The dimensional discriminant element response is output. The discriminant element weighted sum calculation unit 46 outputs the discriminant element response to the discriminant element response value conversion unit 47.

[0045] The discriminant element response value conversion unit 47 converts the discriminant element response input from the discriminant element weighted sum calculation unit 46 to the same dimension as the feature map input from the conversion layer 21. Specifically, the discriminant element response value conversion unit 47 calculates the C for each query. Dv THW-dimensional discriminant response (C Dv The ×THW) is converted to the same C dimension as the feature map input from the conversion layer 21 and aligned into a C×T×H×W type tensor. The discriminant element response value conversion unit 47 then outputs the aligned C×T×H×W type tensor to the addition unit 48 as the converted discriminant element response (hereinafter also referred to as the "converted discriminant element response").

[0046] The summing unit 48 receives the feature map from the conversion layer 21, the converted video feature response from the response value conversion unit 37, and the converted discriminant element response from the discriminant element response value conversion unit 47. The input feature map, converted video feature response, and converted discriminant element response are all C×T×H×W type tensors, and the summing unit 48 adds the converted video feature response and the converted discriminant element response to the feature map. The summing unit 48 then outputs the feature map obtained by the addition to the subsequent conversion layer 21.

[0047] Returning to Figure 3, the feature map output from the identification element reference nonlocal layer 23 is input to the subsequent transformation layer 21. Then, after processing in each transformation layer in sequence, the finally generated feature map is output to the prediction device 5.

[0048] As described above, in this embodiment, the nonlocal layer 23 that references the identification element performs a transformation of the feature map using a fixation mechanism for the identification element, thereby enabling the output of an optimal feature map for downstream tasks that use the identification element.

[0049] In the above configuration, the conversion layer 21 is an example of a video acquisition means and a conversion means, the identification element extractor 22 is an example of a video acquisition means and an identification element extraction means, and the identification element reference nonlocal layer 23 is an example of an identification element reference conversion means. Furthermore, the converted identification element response is an example of a first feature map, and the converted video feature response is an example of a second feature map.

[0050] [Differentiation] Next, a modified version of the first embodiment will be described. The following modifications can be combined as appropriate and applied to the first embodiment.

[0051] (Variation 1) In the first embodiment described above, the video feature attention unit 30 and the identification element attention unit 40 share a query generation unit 31. Alternatively, the feature conversion device 100 may provide query generation units for both the video feature attention unit 30 and the identification element attention unit 40. Specifically, in addition to the query generation unit 31 in Figure 4, an additional query generation unit can be provided that generates queries using a feature map as input, and the output of the additional query generation unit can be input to the identification element weight generation unit 45.

[0052] (Modification 2) In the feature transformation device 100 described above, the identification element extractor 22 extracts one type of identification element, such as a person. However, the feature transformation device 100 can also perform feature transformation using multiple types of identification elements. Multiple types of identification elements are identification elements with different dimensions or numbers of identification element features. Examples of using multiple types of identification elements include using person features and object features together, or using person and object features together with interaction features between people and objects. In this case, the identification element reference nonlocal layer 23 is provided with identification element gaze units 40 in parallel for each type of identification element. Figures 5 and 6 show examples of the configuration of the identification element reference nonlocal layer 23 when there are multiple identification element gaze units 40.

[0053] Figure 5 shows a first example of the identification element nonlocal layer according to Modification 2. In the first example, the identification element reference nonlocal layer 23a comprises a video feature gaze unit 30, a plurality of identification element gaze units 40a to 40n, and a combination transformation unit 49. The plurality of identification element gaze units 40a to 40n each correspond to different types of identification elements. For example, when using person features, object features, and person-object interaction features as multiple types of identification elements, the identification element gaze unit 40a is configured to have a gaze mechanism for person features, the identification element gaze unit 40b is configured to have a gaze mechanism for object features, and the identification element gaze unit 40c is configured to have a gaze mechanism for person-object interaction features.

[0054] The combined transformation unit 49 receives the video feature response generated by the video feature gaze unit 30 and the discrimination element responses generated by the discrimination element gaze units 40a to 40n, respectively. The combined transformation unit 49 combines the video feature response and each discrimination element response and converts them into a tensor with the same shape as the feature map input from the transformation layer 21. The combined transformation unit 49 outputs the combined and transformed result (hereinafter also called the "transformed response") to the adder unit 48. The adder unit 48 adds the transformed response to the input feature map and outputs the added feature map to the subsequent transformation layer 21.

[0055] Figure 6 shows a second example of the identification element nonlocal layer according to Modification 2. In the second example, the identification element reference nonlocal layer 23b comprises a video feature gaze unit 30, a response value conversion unit 37, a plurality of identification element gaze units 40a to 40n, and a plurality of identification element response value conversion units 47a to 47n. That is, in the example of Figure 6, identification element response value conversion units 47a to 47n are provided in correspondence with each of the identification element gaze units 40a to 40n. Each of the identification element response value conversion units 47a to 47n converts the identification element response input from the identification element gaze units 40a to 40n into a tensor with the same shape as the feature map input from the conversion layer 21. Then, the identification element response value conversion units 47a to 47n output the converted identification element response to the summation unit 48. Furthermore, the response value conversion unit 37 converts the video feature response input from the video feature gaze unit 30 into a tensor with the same shape as the feature map input from the conversion layer 21. The response value conversion unit 37 outputs the converted video feature response to the adder 48. The adder 48 adds the converted video feature response and each converted discriminant element response to the input feature map and outputs the added feature map to the subsequent conversion layer 21.

[0056] Furthermore, the first example of the non-local identification element layer shown in Figure 5 and the second example of the non-local identification element layer shown in Figure 6 may be combined. That is, the outputs of some of the identification element viewing units 40 may be input to the adder 48 via a coupling conversion unit, while the outputs of the other identification element viewing units 40 may be input to the adder 48 via individually provided identification element response value conversion units 47.

[0057] (Variation 3) In the first embodiment described above, the discriminant element extractor 22 extracts discriminant element features based on a video input from an external device. Alternatively, the discriminant element extractor 22 may extract discriminant element features using a video input from an external device and a feature map output by any transformation layer located before the discriminant element reference nonlocal layer 23. Figure 7 shows the configuration of the feature transformation device 100x according to Modification 3. In the example of Figure 7, the discriminant element extractor 22 extracts discriminant element features based on a video input from an external device and a feature map input from the transformation layer 21c. For example, the discriminant element extractor 22 detects discriminant elements such as objects from the video and extracts feature quantities from the feature map corresponding to those discriminant elements. That is, the discriminant element extractor 22 detects the position of the target discriminant element in the image based on the input video and extracts the feature quantities of the part of the input feature map corresponding to that position as discriminant element features. The discriminant element extractor 22 then outputs the extracted feature quantities as discriminant element features to the discriminant element reference nonlocal layer 23.

[0058] <Second Embodiment> Figure 8 is a block diagram showing the functional configuration of a video processing device according to the first embodiment. The video processing device 200 includes a video acquisition means 201, a plurality of conversion means 202, an identification element extraction means 203, and an identification element reference conversion means 204.

[0059] Figure 9 is a flowchart of the processing performed by the video processing device of the second embodiment. The video acquisition means 201 acquires a video (step S201). Multiple sequentially arranged conversion means 202 convert the feature map of the input video and output the converted feature map (step S202). The identification element extraction means 203 detects identification elements from the video and outputs the identification element feature quantities (step S203). The identification element reference conversion means 204 is positioned between the multiple conversion means and generates and outputs a feature map based on the input feature map and the identification element feature quantities (step S204).

[0060] According to the video processing device 200 of the second embodiment, it is possible to perform optimal feature transformation by taking into account the identification elements corresponding to the target task.

[0061] Some or all of the above embodiments may also be described as follows, but are not limited to the following:

[0062] (Note 1) Methods for obtaining videos, Multiple transformation means are arranged in a continuous manner to transform the feature map of the input video and output the transformed feature map, A means for extracting identifying elements from the aforementioned video and outputting identifying element features, An identification element reference conversion means is positioned between the plurality of conversion means and generates and outputs a feature map based on the input feature map and the identification element feature quantities. A video processing device equipped with the following features.

[0063] (Note 2) The aforementioned identification element reference conversion means is A means for focusing on the identification element, which generates a first feature map that focuses on the identification element based on the aforementioned identification element features, A video feature-focusing means generates a second feature map that focuses on the input feature map, A video processing device as described in Appendix 1, comprising:

[0064] (Note 3) The video processing apparatus according to Appendix 2, wherein the identification element focusing means generates a first query based on the input feature map, generates a first key and a first value based on the identification element feature quantities, and generates the first feature map using the first query, the first key and the first value.

[0065] (Note 4) The video feature-focusing means generates a second query, a second key, and a second value based on the input feature map, and generates the second feature map using the second query, the second key, and the second value, as described in Appendix 3 of the video processing apparatus.

[0066] (Note 5) The aforementioned identification element reference conversion means comprises a plurality of identification element gaze means provided for each type of identification element, Each of the plurality of identification element focusing means generates the first feature map based on the individual identification element features, as described in Appendix 2.

[0067] (Note 6) The video processing apparatus described in Appendix 1 outputs an identification element feature quantity based on the video and a feature map output by a conversion layer located before the identification element reference conversion means.

[0068] (Note 7) The video processing apparatus described in Appendix 1, wherein the plurality of conversion means and the identification element reference conversion means are configured by a neural network.

[0069] (Note 8) Executed by a computer, Obtain the video, The feature map of the input video is converted and the converted feature map is output. The video above is used to detect identifying elements and output the features of those identifying elements. A video processing method that generates and outputs a feature map based on the converted feature map and the identification element feature quantities.

[0070] (Note 9) Obtain the video, The feature map of the input video is converted and the converted feature map is output. The video above is used to detect identifying elements and output the features of those identifying elements. A recording medium that stores a program that causes a computer to execute a process to generate and output a feature map based on the converted feature map and the discriminant element feature quantities.

[0071] Although the present disclosure has been described above with reference to embodiments and examples, the present disclosure is not limited to the above embodiments and examples. Various modifications to the structure and details of the present disclosure can be understood by those skilled in the art within the scope of the present disclosure. [Explanation of Symbols]

[0072] 1. Video Processing Device 5 Prediction device 21 Conversion Layer 22. Identifier Extractor 23, 23a, 23b Identification element reference nonlocal layer 30. Video Features: Focus Area 40 Identification element viewing area 100, 100x Feature Conversion Device

Claims

1. Methods for obtaining videos, A conversion means that converts the feature map of the aforementioned video and outputs the converted feature map, A means for extracting identification elements from the aforementioned video, and for extracting and outputting identification element features from the aforementioned identification elements, A means for focusing on a discriminant element to generate a discriminant element response based on the converted feature map and the discriminant element feature quantities, A means for converting the aforementioned identification element response into a first feature map with the same dimension as the converted feature map, A video feature focusing means generates a video feature response that focuses on video features by obtaining, for each point in the transformed feature map, a video feature response based on the relationship between each point in the transformed feature map as an output, based on the transformed feature map; A response value conversion means that converts the aforementioned video feature response into a second feature map with the same dimension as the converted feature map, Adding means for generating a third feature map by adding the first feature map and the second feature map to the converted feature map, A video processing device equipped with the following features.

2. The video processing apparatus according to claim 1, wherein the identification element focusing means generates a first query based on the transformed feature map, generates a first key and a first value based on the identification element feature quantities, and generates the identification element response using the first query, the first key and the first value.

3. The video processing apparatus according to claim 2, wherein the video feature looking means generates a second query, a second key, and a second value based on the converted feature map, and generates the video feature response using the second query, the second key, and the second value.

4. The aforementioned video processing device includes a plurality of identification element viewing means provided for each type of identification element, The video processing apparatus according to claim 1, wherein each of the plurality of identification element gaze means generates the identification element response based on the individual identification element feature quantities.

5. The video processing apparatus according to claim 1, wherein the distinguishing element feature extraction means outputs distinguishing element features based on the video and the converted feature map.

6. The video processing device is The system further comprises a second conversion means for converting the feature map of the aforementioned video and outputting a feature map, The conversion means converts the feature map output by the second conversion means and outputs the converted feature map. The aforementioned identification element feature extraction means outputs identification element features based on the video and the feature map output by the second conversion means. The video processing apparatus according to claim 1.

7. The video processing apparatus according to claim 1, wherein the conversion means is configured by a neural network.

8. The identification element feature extraction means detects at least one of a person, an object, and an interaction between a person and an object included in the video as the identification element, The aforementioned identification element feature quantity includes at least one of the following features of the detected identification element: coordinates, size, type, orientation information, and appearance. The video processing apparatus according to claim 1.

9. Executed by a computer, Obtain the video, The feature map of the aforementioned video is transformed and the transformed feature map is output. The system detects the identification elements from the aforementioned video and outputs the identification element features from the aforementioned identification elements. Based on the transformed feature map and the identification element features, an identification element response focused on the identification element is generated. The aforementioned discriminant element response is converted into a first feature map having the same dimension as the converted feature map. Based on the transformed feature map, a video feature response is obtained as output for each point in the transformed feature map, based on its relationship with each other in the transformed feature map, thereby generating the video feature response that focuses on video features. The aforementioned video feature response is transformed into a second feature map having the same dimension as the transformed feature map. A video processing method that generates a third feature map by adding the first feature map and the second feature map to the converted feature map.

10. Obtain the video, The feature map of the aforementioned video is transformed and the transformed feature map is output. The system detects the identification elements from the aforementioned video and outputs the identification element features from the aforementioned identification elements. Based on the transformed feature map and the identification element features, an identification element response focused on the identification element is generated. The aforementioned discriminant element response is converted into a first feature map having the same dimension as the converted feature map. Based on the transformed feature map, a video feature response is obtained as output for each point in the transformed feature map, based on its relationship with each other in the transformed feature map, thereby generating the video feature response that focuses on video features. The aforementioned video feature response is transformed into a second feature map having the same dimension as the transformed feature map. A program that causes a computer to perform a process of adding the first feature map and the second feature map to the converted feature map to generate a third feature map.