Video processing device, video processing method, and program
Patent Information
- Application Number
- JP2025504910
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-05
AI Technical Summary
In video recognition using neural networks, existing methods fail to perform optimal feature conversion based on identification elements, as the non-local layer does not explicitly encode these elements, limiting the accuracy of downstream recognition tasks.
A video processing device is designed with a feature conversion mechanism that includes identification element extraction and an identification element reference non-local layer, which generates a transformed feature map and identification element feature amount, enabling optimal feature conversion by utilizing the features of identification elements.
This approach allows for improved feature conversion and recognition by explicitly encoding identification elements, enhancing the accuracy of video recognition tasks.
Abstract
Description
Video processing device, video processing method, and recording medium
[0001] The present disclosure relates to feature extraction from video.
[0002] In video recognition technology using neural networks (NNs), a method of introducing a non-local layer into an intermediate layer (non-local neural networks) has been proposed to improve accuracy (see Non-Patent Document 1). Patent Document 1 also proposes a feature conversion device and an image recognition device that improve on the technology related to the method of introducing a non-local layer into the intermediate layer.
[0003] International Publication No. WO2021-176566
[0004] Non-Local Neural Networks, Xiaolong Wang, Ross Girshick, Abhinav Gupta, Kaiming He Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7794-7803
[0005] In video recognition using neural networks, it is necessary to perform feature transformation, extract discriminant elements contained in the video, and perform recognition based on the extracted discriminant elements. In this way, in video recognition using neural networks, the discriminant elements are only used in downstream recognition tasks, so optimal feature transformation based on the discriminant elements cannot be achieved. Furthermore, in the method of Non-Patent Document 1, the non-local layer does not explicitly encode the discriminant elements, so optimal feature transformation cannot be achieved.
[0006] One object of the present disclosure is to provide a moving image processing device that is capable of performing optimal feature conversion by using features of discriminant elements.
[0007] In one aspect of the present disclosure, a video processing device comprises: a video acquisition means for acquiring a video; a plurality of conversion means, which are arranged successively, for converting a feature map of the input video and outputting a converted feature map; a discrimination element extraction means for detecting discrimination elements from the video and outputting discrimination element features; and a discrimination element reference conversion means, which is arranged between the plurality of conversion means, for generating and outputting a converted feature map and a converted discrimination element features based on the input feature map and the discrimination element features.
[0008] In another aspect of the present disclosure, a video processing method is executed by a computer, acquires a video, converts a feature map of the input video and outputs the converted feature map, detects discriminant elements from the video and outputs discriminant element features, and generates and outputs a feature map and discriminant element features based on the converted feature map and the discriminant element features.
[0009] In yet another aspect of the present disclosure, a recording medium records a program that causes a computer to execute the following processes: acquire a video; convert a feature map of the input video and output the converted feature map; detect discriminant elements from the video and output discriminant element features; and generate and output a feature map and discriminant element features based on the converted feature map and the discriminant element features.
[0010] According to the present disclosure, by using the features of the discriminant element, it is possible to perform optimal feature transformation.
[0011] 1 is a block diagram showing a schematic configuration of a moving image processing device according to a first embodiment; FIG. 2 is a block diagram showing a hardware configuration of a feature conversion device; FIG. 3 is a block diagram showing a functional configuration of a feature conversion device; FIG. 4 is a block diagram showing a functional configuration of an identification element reference non-local layer; FIG. 5 is a block diagram showing a configuration of an identification element reference non-local layer according to a first modified example; FIG. 6 is a block diagram showing another configuration of an identification element reference non-local layer according to a first modified example; FIG. 7 is a block diagram showing a functional configuration of a feature conversion device according to a second modified example; and FIG. 8 is a flowchart of processing by the moving image processing device according to the second embodiment.
[0012] Preferred embodiments of the present disclosure will now be described with reference to the drawings. First Embodiment [System Configuration] Fig. 1 shows a schematic configuration of a video processing device 1. The video processing device 1 comprises a feature conversion device 100 and a prediction device 5. The feature conversion device 100 generates a feature map from an input video and outputs it to the prediction device 5. The prediction device 5 recognizes objects included in the video based on the feature map input from the feature conversion device 100 and outputs the results.
[0013] The feature transformation device 100 of this embodiment is characterized by generating a feature map using discriminant elements. Discriminant elements are elements necessary for identifying an object and have information necessary for executing a task performed downstream of the feature transformation device 100, which in this embodiment is the prediction performed by the prediction device 5. For example, in the case of human behavior recognition, discriminant elements include a target person included in a video and objects related to the target person's behavior. The feature transformation device 100 detects discriminant elements from the input video and extracts feature quantities of the discriminant elements. The feature transformation device 100 then generates a feature map using the feature quantities of the discriminant elements and outputs the feature map to the prediction device 5.
[0014] 2 is a block diagram showing the hardware configuration of the feature transformation device 100 according to the first embodiment. As shown in the figure, the feature transformation device 100 includes an interface (I / F) 11, a processor 12, a memory 13, a recording medium 14, and a database (DB) 15.
[0015] The I / F 11 inputs and outputs data to and from an external device. Specifically, when a video used by the feature transformation device 100 is input from an external device via communication, the I / F 11 receives the input data. The I / F 11 is also used when the feature transformation device 100 outputs a feature map generated by the feature transformation device 100 to an external device.
[0016] The processor 12 is a computer such as a CPU (Central Processing Unit), and executes a pre-prepared program to control the entire feature transformation device 100. The processor 12 may also be a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array).
[0017] The memory 13 is configured by a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The memory 13 is also used as a working memory while the processor 12 is executing various processes.
[0018] The recording medium 14 is a non-volatile, non-transitory recording medium such as a disk-shaped recording medium or a semiconductor memory, and is configured to be detachable from the feature conversion device 100. The recording medium 14 records various programs to be executed by the processor 12. When the feature conversion device 100 executes various processes, the programs recorded on the recording medium 14 are loaded into the memory 13 and executed by the processor 12.
[0019] The DB 15 stores data input and output via the I / F 11 as necessary.
[0020] [Functional Configuration] Figure 3 is a block diagram showing the functional configuration of the feature conversion device 100 according to the first embodiment. Functionally, the feature conversion device 100 includes conversion layers 21a to 21n, an identification element extractor 22, and identification element reference non-local layers 23a and 23b. In the following description, when the individual conversion layers 21a to 21n are not distinguished, they may be simply referred to as "conversion layer 21." Furthermore, when the individual identification element reference non-local layers 23a and 23b are not distinguished, they may be simply referred to as "identification element reference non-local layer 23." For convenience of illustration, Figure 3 shows two identification element reference non-local layers, the identification element reference non-local layers 23a and 23b. However, it is assumed that multiple identification element reference non-local layers are also provided downstream of the identification element reference non-local layer 23b.
[0021] A moving image is input to the feature conversion device 100 from an external device through the I / F 11. The moving image is input to the conversion layer 21a and the discriminant element extractor 22. The moving image is input to the C i ×T i ×H i ×W i is a tensor of "C i " indicates the number of channels. If the input video has three channels of RGB, "C i " is 3. Also, "T i " indicates time, and "H i " indicates the height of the image (number of vertical pixels), and "W i " indicates the width of the image (number of horizontal pixels).
[0022] The transformation layer 21 includes, for example, a convolution layer, a residual block layer, and a pooling layer, and transforms an input feature map and outputs the transformed feature map. Specifically, the transformation layers 21a to 21n are arranged consecutively, and the transformation layer 21a performs processing such as convolution on the input video and outputs the generated feature map to the subsequent transformation layer 21b. The transformation layer 21b performs processing such as convolution on the input feature map and outputs the generated feature map to the subsequent transformation layer 21c. The feature maps generated through processing in each transformation layer are then output to the prediction device 5.
[0023] The discriminant element extractor 22 detects discriminant elements from the input video and extracts their feature quantities (hereinafter also referred to as "discriminant element feature quantities"). One type of discriminant element is set in advance in the discriminant element extractor 22. When the discriminant element extractor 22 detects K discriminant elements corresponding to that one type from the input video, it performs C D The discrimination element feature (C D × K). D For example, if a cart is set as the discrimination element and three carts are included in the input video, the discrimination element extractor 22 extracts the discrimination element feature quantity (C D Then, the discrimination element extractor 22 outputs the extracted discrimination element feature amount to the discrimination element reference non-local layer 23a.
[0024] For example, if the identification element is a person, the identification element extractor 22 outputs the person's coordinates, size, joint point coordinates, and features of how the person looks as identification element features. If the identification element is an object, the identification element extractor 22 extracts the object's coordinates, size, type, posture information, and features of how the object looks as identification element features. Furthermore, if the identification element is an interaction between a person and an object, the identification element extractor 22 extracts the coordinates of the interacting elements (i.e., the person and the object), the types of those elements, the type of interaction, and the like as identification element features.
[0025] The discrimination element reference non-local layer 23a is disposed at any position between the transformation layers 21a to 21n. The discrimination element reference non-local layer 23a generates a feature map and a discrimination element feature based on the feature map input from the preceding transformation layer 21 and the discrimination element feature input from the discrimination element extractor 22. The discrimination element reference non-local layer 23a outputs the generated feature map and the generated discrimination element feature to the subsequent discrimination element reference non-local layer 23b.
[0026] The discrimination element reference non-local layer 23b is disposed after the discrimination element reference non-local layer 23a. The discrimination element reference non-local layer 23b generates a feature map and discrimination element features based on the feature map and discrimination element features input from the discrimination element reference non-local layer 23a in the preceding stage. The discrimination element reference non-local layer 23b outputs the generated feature map to the subsequent conversion layer 21. The discrimination element reference non-local layer 23b also outputs the generated discrimination element features to the subsequent discrimination element reference non-local layer 23.
[0027] 3, the discrimination element reference non-local layers 23a and 23b are arranged continuously, but a plurality of transformation layers 21 may be arranged between the discrimination element reference non-local layer 23a and the discrimination element reference non-local layer 23b. In this case, the discrimination element reference non-local layer 23a outputs the generated feature map to the subsequent transformation layer 21, and outputs the generated discrimination element feature amount to the discrimination element reference non-local layer 23b.
[0028] 4 shows the configuration of the discrimination element reference non-local layer 23. The discrimination element reference non-local layer 23 includes a video feature query generation unit 31, a resolution conversion unit 32, a video feature key generation unit 33, a video feature value generation unit 34, connection units 35a, 35b, and 35c, a discrimination element query generation unit 41, a discrimination element key generation unit 43, a discrimination element value generation unit 44, a gaze weight generation unit 45, a weighted sum calculation unit 46, a response value conversion unit 47, a first adder 48, and a second adder 49.
[0029] The discrimination element reference non-local layer 23 includes a type of attention device equipped with four attention mechanisms. The attention device is a device that receives a query element, a key element, and a value element as input and outputs a response feature. More specifically, the attention device is a device equipped with a mechanism that controls the value (generated from the value element) to be captured for each query generated from the query element based on the similarity with the key generated from the key element.
[0030] The identification element reference non-local layer 23 in Figure 4 has four gaze mechanisms: self-gaze in which the query element is an animation and the key element and value element are animations; cross-gaze in which the query element is an animation and the key element and value element are identification elements; cross-gaze in which the query element is an identification element and the key element and value element are animations; and self-gaze in which the query element is an identification element and the key element and value element are identification elements.
[0031] The feature map is input to the discriminant element reference non-local layer 23 from the transformation layer 21. The feature map is a C×T×H×W type tensor. The discriminant element reference non-local layer 23 also receives K discriminant element feature amounts (C D ×K) is input.
[0032] The video feature query generation unit 31 receives the feature map from the conversion layer 21. The video feature query generation unit 31 generates a query using the input feature map (C×T×H×W) as a query element. Specifically, the video feature query generation unit 31 converts the C-dimensional vector of each point on the feature map into C q Dimensional query (C qThis dimension transformation can be performed by, for example, a linear transformation or a combination of a linear transformation, an activation function, and layer normalization. q The number of dimensions is a predetermined number. The video feature query generation unit 31 converts the query so that the number of dimensions of the query matches the number of dimensions of the key described below. Since there are T HW points in the feature map, the number of queries is T HW. The video feature query generation unit 31 outputs the generated video feature query to the linking unit 35a.
[0033] The feature map is input to the resolution conversion unit 32 from the conversion layer 21. The resolution conversion unit 32 converts the resolution of the input feature map to a predetermined value. The resolution conversion unit 32 outputs the feature map (C×T'×H'×W') with the converted resolution to the video feature key generation unit 33 and the video feature value generation unit 34.
[0034] The moving image feature key generating unit 33 generates a key using the feature map (C×T′×H′×W′) input from the resolution converting unit 32 as a key element. Specifically, the moving image feature key generating unit 33 converts a C-dimensional vector of each point of the feature map into C q Dimensional Key (C q ×T'H'W'). This dimensional transformation can be performed by, for example, a linear transformation or a combination of a linear transformation, an activation function, and layer normalization. Note that since there are T'H'W' points in the feature map, the number of keys is T'H'W'. The video feature key generation unit 33 outputs the generated video feature keys to the connection unit 35b.
[0035] The video feature value generation unit 34 generates a value using the feature map (C×T′×H′×W′) input from the resolution conversion unit 32 as a value element. Specifically, the video feature value generation unit 34 converts a C-dimensional vector of each point on the feature map into a C v Dimensional Value (C v × T'H'W'). This dimension transformation can be performed by, for example, a linear transformation or a combination of a linear transformation, an activation function, and layer normalization. C vThe number of dimensions is a predetermined number. Since there are T'H'W' points in the feature map, the number of values is T'H'W'. The video feature value generation unit 34 outputs the generated video feature values to the connection unit 35c.
[0036] The resolution conversion unit 32 may be omitted. In this case, the video feature key generation unit 33 and the video feature value generation unit 34 may generate keys and values from the feature map (C×T×H×W) input from the conversion layer 21.
[0037] The discriminant element query generation unit 41 receives K discriminant element feature amounts from the discriminant element extractor 22 or the preceding discriminant element reference non-local layer 23. The discriminant element query generation unit 41 calculates the input discriminant element feature amounts (C D ×K) as a query element. D Dimensional discriminant feature (C D ×K) is the same as the query generated by the video feature query generation unit 31. q Dimensional query (C q ×K). This dimensional transformation can be performed by, for example, a linear transformation or a combination of a linear transformation, an activation function, and layer normalization. The discriminant element query generation unit 41 outputs the generated discriminant element query to the connection unit 35a.
[0038] The discrimination element key generation unit 43 receives K discrimination element feature amounts from the discrimination element extractor 22 or the preceding discrimination element reference non-local layer 23. The discrimination element key generation unit 43 generates a discrimination element key using the input discrimination element feature amounts as key elements. Specifically, the discrimination element key generation unit 43 generates a discrimination element key using C D Dimensional discriminant feature (C D ×K) is converted into the same C as the key generated by the moving image feature key generating unit 33. q Dimension identification element key (C q ×K). This dimension transformation can be performed by, for example, a linear transformation or a combination of a linear transformation, an activation function, and layer normalization. The identification element key generation unit 43 outputs the generated identification element key to the connection unit 35b.
[0039] The discrimination element value generation unit 44 receives K discrimination element feature amounts from the discrimination element extractor 22 or the preceding discrimination element reference non-local layer 23. The discrimination element value generation unit 44 generates a discrimination element value using the input discrimination element feature amounts as value elements. Specifically, the discrimination element value generation unit 44 generates a discrimination element value using C D Dimensional discriminant feature (C D ×K) is converted into the same value C as the value generated by the video feature value generating unit 34. v Dimensional discrimination factor value (C v ×K). This dimension transformation can be performed by, for example, a linear transformation or a combination of a linear transformation, an activation function, and layer normalization. The discrimination element value generation unit 44 outputs the generated discrimination element value to the connection unit 35c.
[0040] The linking unit 35a receives the video feature query (C q ×THW) is input, and the discrimination element query generation unit 41 outputs a discrimination element query (C q ×K) is input to the connection unit 35a. The connection unit 35a connects the T H W video feature queries to C q ×THW matrix and C q ×K matrix. The concatenation unit 35a concatenates the concatenated query (C q ×(THW+K)) to the gaze weight generating unit 45.
[0041] The linking section 35b receives the moving image feature key (C q ×T'H'W') is input, and the identification element key generating unit 43 generates an identification element key (C q ×K) is input to the connection unit 35b. The connection unit 35b connects T'H'W' movie feature queries to C q ×T'H'W' matrix and C q ×K matrix. The concatenation unit 35b concatenates the concatenated key (C q ×(T′H′W′+K)) to the gaze weight generating unit 45.
[0042] The linking section 35c receives the moving image feature value (C v×T'H'W') is input, and the discrimination element value generating unit 44 outputs the discrimination element value (C v ×K) is input to the connection unit 35c. The connection unit 35c connects T'H'W' movie feature queries to C v ×T'H'W' matrix and C v ×K matrix. The concatenation unit 35c concatenates the concatenated value (C v ×(T′H′W′+K)) to the weighted sum calculation unit 46 .
[0043] The gaze weight generating unit 45 receives the query (C q × (THW+K)) is input, and the linked key (C q ×(T'H'W'+K)) is input. The gaze weight generation unit 45 generates a weight for each value based on the similarity between the query and the corresponding key.
[0044] Specifically, the attention weight generation unit 45 generates weights using the inner product of the query and the key. q × (THW+K) matrix is "Q", and C is a matrix of T'H'W'+K keys. q When the matrix (THW+K)×(TH'H'W'+K) is represented as "K" and the weight matrix of the output (THW+K)×(TH'H'W'+K) is represented as "A", the weight matrix A is given by the following equation.
[0045] In another example, the attention weight generation unit 45 generates weights by calculating the inner product of the query and the key and normalizing it using the SoftMax function. In this case, the weight matrix A is given by the following formula:
[0046] Then, the gaze weight generating unit 45 outputs the generated weight to the weighted sum calculating unit 46 .
[0047] The weighted sum calculation unit 46 receives the linked value (C v ×(T'H'W'+K)) is input, and weights are input from the gaze weight generation unit 45. The weighted sum calculation unit 46 generates a response value based on the values and weights. For example, the weighted sum calculation unit 46 calculates the value (C vThe weighted sum of values is calculated by taking the matrix product of (THW+K) × (TH'H'W'+K) and the weight (THW+K) × (TH'H'W'+K), and the response value (C v ×(THW+K)) is generated. The weighted sum calculation unit 46 outputs the generated response value to the response value conversion unit 47.
[0048] The response value conversion unit 47 decomposes the response value input from the weighted sum calculation unit 46 into a response value for the feature map and a response value for the discrimination element feature amount. v × (THW+K) is the response value for the feature map, C v ×THW matrix and the response value for the discrimination element feature, C v ×K matrix.
[0049] Then, the response value converter 47 calculates C v The response value conversion unit 47 converts the C × THW matrix into a C-dimensional vector (C × THW) that is the same as the feature map input from the conversion layer 21. The response value conversion unit 47 outputs the converted response value (hereinafter also referred to as the "feature map residual" or the "first feature map") to the first adder 48. The response value conversion unit 47 also converts the C v ×K matrix is converted into the same C D dimensional vector (C D ×K) to the second adder 49. Then, the response value converter 47 outputs the converted response value (hereinafter also referred to as a “discriminant component feature residual” or a “second feature map”).
[0050] The first adder 48 receives the feature map from the transformation layer 21 and the feature map residual from the response value transformation unit 47. The first adder 48 calculates the sum of the input feature map and the feature map residual. The first adder 48 outputs the calculation result (hereinafter also referred to as the "transformed feature map") to a transformation layer or a discrimination element reference non-local layer arranged after the discrimination element reference non-local layer 23.
[0051] The second adder 49 receives the discrimination element feature from the discrimination element extractor 22 or the preceding discrimination element reference non-local layer 23, and receives the discrimination element feature residual from the response value conversion unit 47. The second adder 49 calculates the sum of the discrimination element feature and the discrimination element feature residual. The second adder 49 outputs the calculation result (hereinafter also referred to as the "transformed discrimination element feature") to a discrimination element reference non-local layer arranged subsequent to the discrimination element reference non-local layer 23.
[0052] 3 , the transformed feature map output from the discriminant element reference non-local layer 23 is input to the subsequent transformation layer 21 or discriminant element reference non-local layer. The transformed discriminant element feature values output from the discriminant element reference non-local layer 23 are also input to the subsequent discriminant element reference non-local layer. The feature map generated through processing in each layer is then output to the prediction device 5.
[0053] As described above, in this embodiment, the discrimination element reference non-local layer 23 generates a feature map using a discrimination element, thereby enabling an optimal feature map to be output for a downstream task. Furthermore, in this embodiment, each discrimination element reference non-local layer 23 generates transformed discrimination element features, and the transformed discrimination element features are used as discrimination element features in the subsequent discrimination element reference non-local layer 23, enabling the subsequent discrimination element reference non-local layer 23 to generate a feature map using the optimal discrimination element features.
[0054] In the above configuration, the conversion layer 21 is an example of a video acquisition means and a conversion means, the identification element extractor 22 is an example of a video acquisition means and an identification element extraction means, and the identification element reference non-local layer 23 is an example of an identification element reference conversion means.
[0055] [Modifications] Next, a description will be given of modifications of the first embodiment. The following modifications can be applied to the first embodiment in appropriate combinations.
[0056] (Variation 1) In the feature transformation device 100 described above, the discriminant element extractor 22 extracts one type of discriminant element, such as a person. However, the feature transformation device 100 can also perform feature transformation using multiple types of discriminant elements. Multiple types of discriminant elements are discriminant elements with different dimensions or numbers of discriminant element features. Examples of using multiple types of discriminant elements include using a person feature and an object feature in combination, or using a person and object feature in combination with a person-object interaction feature. In this case, the discriminant element reference non-local layer 23 is provided with a gaze unit 50 in parallel for each type of discriminant element. Figures 5 and 6 show configuration examples of the discriminant element reference non-local layer 23 when multiple gaze units 50 are included. The gaze unit 50 includes a video feature query generation unit 31, a resolution conversion unit 32, a video feature key generation unit 33, a video feature value generation unit 34, connection units 35a, 35b, 35c, an identification element query generation unit 41, an identification element key generation unit 43, an identification element value generation unit 44, a gaze weight generation unit 45, and a weighted sum calculation unit 46.
[0057] 5 shows a first example of an identification element non-local layer according to Modification 1. In the first example, the identification element reference non-local layer 23c includes a plurality of gaze units 50a to 50n and a combination conversion unit 51. The plurality of gaze units 50a to 50n correspond to different types of identification elements. For example, when person features, object features, and person-object interaction features are used as the plurality of types of identification elements, the gaze unit 50a is configured to include a gaze mechanism for the person features, the gaze unit 50b is configured to include a gaze mechanism for the object features, and the gaze unit 50c is configured to include a gaze mechanism for the person-object interaction features.
[0058] The response values generated by the attention units 50a to 50n are input to the combination transformation unit 51. The combination transformation unit 51 combines the response values and converts the combined response values into a feature map residual and a discrimination element feature residual. The combination transformation unit 51 outputs the feature map residual to the first adder 48 and outputs the discrimination element feature residual to the second adder 49. The first adder 48 adds the feature map residual to the input feature map and outputs the added feature map to the subsequent transformation layer 21. The second adder 49 adds the discrimination element feature residual to the discrimination element feature and outputs the added discrimination element feature to the subsequent discrimination element reference non-local layer 23.
[0059] FIG. 6 shows a second example of a discrimination element non-local layer according to Modification 1. In this second example, the discrimination element reference non-local layer 23d includes multiple attention units 50a-50n and multiple response value converters 47a-47n. That is, in the example of FIG. 6, response value converters 47a-47n are provided corresponding to the attention units 50a-50n, respectively. Each of the response value converters 47a-47n converts the response values input from the attention units 50a-50n into a feature map residual and a discrimination element feature residual. The response value converters 47a-47n then output the feature map residuals to a first adder 48 and the discrimination element feature residuals to a second adder 49. The first adder 48 adds each feature map residual to the input feature map and outputs the resulting feature map to the subsequent conversion layer 21. The second adder 49 adds each discrimination element feature residual to the discrimination element feature, and outputs the added discrimination element feature to the subsequent discrimination element reference non-local layer 23 .
[0060] It is also possible to combine the first example of the discrimination element non-local layer shown in Fig. 5 with the second example of the discrimination element non-local layer shown in Fig. 6. That is, among the multiple observation units, the outputs of some observation units 50 may be input to the first adder 48 and the second adder 49 via a combined conversion unit, and the outputs of the other observation units 50 may be input to the first adder 48 and the second adder 49 via individually provided response value conversion units 47.
[0061] (Variation 2) In the first embodiment described above, the discriminant element extractor 22 extracts discriminant element features based on a video input from an external device. Alternatively, the discriminant element extractor 22 may extract discriminant element features using a video input from an external device and a feature map output by any transformation layer located before the discriminant element reference non-local layer 23. FIG. 7 shows the configuration of a feature transformation device 100x according to Variation 2. In the example of FIG. 7 , the discriminant element extractor 22 extracts discriminant element features based on a video input from an external device and a feature map input from the transformation layer 21c. For example, the discriminant element extractor 22 detects discriminant elements, such as objects, from the video and extracts feature values of regions corresponding to the discriminant elements from the feature map. That is, the discriminant element extractor 22 detects the position of the target discriminant element in the image based on the input video and extracts feature values of portions corresponding to the position in the input feature map as discriminant element features. Then, the discrimination element extractor 22 outputs the extracted feature amount to the discrimination element reference non-local layer 23 as a discrimination element feature amount.
[0062] 8 is a block diagram showing the functional configuration of a moving image processing device according to Embodiment 1. The moving image processing device 200 includes a moving image acquisition unit 201, a plurality of conversion units 202, an identification element extraction unit 203, and an identification element reference conversion unit 204.
[0063] 9 is a flowchart of processing by the video processing device of the second embodiment. A video acquisition unit 201 acquires a video (step S201). A plurality of consecutively arranged conversion units 202 convert the feature map of the input video and output the converted feature map (step S202). A discriminant element extraction unit 203 detects discriminant elements from the video and outputs discriminant element features (step S203). A discriminant element reference conversion unit 204 is arranged between the plurality of conversion units and generates and outputs a converted feature map and converted discriminant element features based on the input feature map and the discriminant element features (step S204).
[0064] According to the moving image processing device 200 of the second embodiment, it is possible to perform optimal feature conversion by using the features of the discrimination elements.
[0065] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.
[0066] (Supplementary Note 1) A video processing device comprising: a video acquisition means for acquiring a video; a plurality of conversion means arranged successively and converting a feature map of the input video and outputting a converted feature map; a discrimination element extraction means for detecting discrimination elements from the video and outputting discrimination element features; and a discrimination element reference conversion means arranged between the plurality of conversion means and generating and outputting a converted feature map and a converted discrimination element features based on the input feature map and the discrimination element features.
[0067] (Supplementary Note 2) The video processing device according to Supplementary Note 1, wherein the discriminant element reference conversion means comprises: a feature map acquisition means for acquiring an input feature map; a discrimination element feature acquisition means for acquiring a discrimination element feature; a generation means for generating a query, a key, and a value from the feature map and the discrimination element feature, respectively; a gaze weight generation means for generating a gaze weight from the query and the key; a weighted sum calculation means for calculating a response value that is a weighted sum of the values according to the gaze weight; a response value conversion means for converting the response value into a feature map residual and a discrimination element feature residual; a first addition means for calculating a converted feature map as the sum of the feature map and the feature map residual; and a second addition means for calculating a converted discrimination element feature as the sum of the discrimination element feature and the discrimination element feature residual.
[0068] (Supplementary Note 3) The video processing device according to Supplementary Note 2, wherein the discrimination element reference conversion means includes a plurality of gaze means provided for each type of discrimination element, and each of the plurality of gaze means converts the feature map residual and the discrimination element feature residual based on an individual discrimination element feature.
[0069] (Supplementary Note 4) The video processing device according to Supplementary Note 1, wherein the discriminant element extraction means outputs discriminant element features based on the video and a feature map output by a conversion layer located before the discriminant element reference conversion means.
[0070] (Supplementary Note 5) The moving image processing device according to Supplementary Note 1, wherein the plurality of conversion means and the identification element reference conversion means are configured by a neural network.
[0071] (Supplementary Note 6) A video processing method executed by a computer, comprising: acquiring a video; converting a feature map of the input video and outputting the converted feature map; detecting discriminant elements from the video and outputting discriminant element features; and generating and outputting a feature map and discriminant element features based on the converted feature map and the discriminant element features.
[0072] (Supplementary Note 7) A recording medium having recorded thereon a program that causes a computer to execute the following processes: acquire a video; convert a feature map of the input video and output the converted feature map; detect discriminant elements from the video and output discriminant element features; and generate and output a feature map and discriminant element features based on the converted feature map and the discriminant element features.
[0073] Although the present disclosure has been described above with reference to the embodiments and examples, the present disclosure is not limited to the above-described embodiments and examples. Various modifications that can be understood by a person skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure.
[0074] REFERENCE SIGNS LIST 1 Video processing device 5 Prediction device 21 Conversion layer 22 Discrimination element extractor 23, 23a, 23b, 23c, 23d Discrimination element reference non-local layer 31 Video feature query generation unit 32 Resolution conversion unit 33 Video feature key generation unit 34 Video feature value generation unit 35a, 35b, 35c Connection unit 41 Discrimination element query generation unit 43 Discrimination element key generation unit 44 Discrimination element value generation unit 45 Attention weight generation unit 46 Weighted sum calculation unit 47 Response value conversion unit 48 First addition unit 49 Second addition unit 100, 100x Feature conversion device
Claims
1. a video acquisition means for acquiring a video; a plurality of conversion means arranged consecutively, which convert the feature map of the input video and output the converted feature map; an identification element extraction means for detecting an identification element from the video and outputting an identification element feature amount; a discriminant element reference conversion means disposed between the plurality of conversion means, which generates and outputs a converted feature map and a converted discriminant element feature based on the input feature map and the discriminant element feature; A video processing device comprising:
2. The identification element reference conversion means a feature map acquisition means for acquiring an input feature map; a discrimination element feature amount acquiring means for acquiring a discrimination element feature amount; a generation means for generating a query, a key, and a value from the feature map and the discriminant element feature, respectively; a gaze weight generating means for generating a gaze weight from the query and the key; a weighted sum calculation means for calculating a response value, which is a weighted sum of the values, according to the gaze weights; a response value conversion means for converting the response value into a feature map residual and a discrimination element feature residual; a first summing means for calculating a transformed feature map as a sum of the feature map and the feature map residual; a second adding means for calculating a transformed discrimination element feature as a sum of the discrimination element feature and the discrimination element feature residual; The video processing device according to claim 1 , comprising:
3. the identification element reference conversion means includes a plurality of gaze means provided for each type of identification element; The video processing device according to claim 2 , wherein each of the plurality of gazing means converts the feature map residual and the discriminant element feature residual based on an individual discriminant element feature.
4. The video processing device according to claim 1 , wherein the discriminant element extraction means outputs discriminant element features based on the video and a feature map output by a conversion layer located before the discriminant element reference conversion means.
5. 2. The video processing device according to claim 1, wherein the plurality of conversion means and the identification element reference conversion means are configured by a neural network.
6. is executed by a computer, Get the video, Transform the feature map of the input video and output the transformed feature map; Detecting an identification element from the video and outputting an identification element feature amount; A moving image processing method for generating and outputting a feature map and a discriminant element feature based on the converted feature map and the discriminant element feature.
7. Get the video, Transform the feature map of the input video and output the transformed feature map; Detecting an identification element from the video and outputting an identification element feature amount; A program that causes a computer to execute a process of generating and outputting a feature map and a discrimination element feature based on the transformed feature map and the discrimination element feature.