Depth feature characterization method based on high-dimensional shift field

By introducing high-dimensional shift field technology into deep convolutional networks, feature extraction and structured characterization of spectral video data is solved, and the problem that the existing technology is difficult to deal with high-dimensional data is achieved, achieving more efficient feature extraction and characterization effects.

CN120107623APending Publication Date: 2025-06-06NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510184499.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing deep convolutional networks are difficult to effectively process high-dimensional spectral video data, resulting in poor performance in feature extraction and structured characterization.

Method used

The depth feature characterization method based on a high-dimensional shift field is adopted, and effective feature extraction and structured characterization of high-dimensional data is achieved by performing dimensional recombination, spectral channel mapping, three-dimensional space-time and spatial shift prediction and depth feature map enhancement processing on spectral video data.

Benefits of technology

Without changing the classic deep convolutional network structure, the efficient up-dimensional processing of spectral video data is improved, which improves the fine-grained feature extraction and the accuracy of structured characterization, and has good migration and computing resource friendliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107623A_ABST
    Figure CN120107623A_ABST
Patent Text Reader

Abstract

The invention provides a depth feature characterization method based on a high-dimensional shift field, and the method comprises the steps: 1, carrying out the dimension recombination and spectrum channel mapping processing of spectrum video data, carrying out the feature extraction through a deep convolution network, and obtaining a multi-dimensional depth feature map of the spectrum video data; step 2, predicting a shift offset in a three-dimensional space-time space for each voxel in the multi-dimensional depth feature map; step 3, sampling the multi-dimensional depth feature map in a spectral video multi-dimensional space, calculating an interpolation weight, and obtaining a depth shift feature map after voxel shift field enhancement processing; and step 4, processing the depth shift feature map by using a convolution kernel, and performing residual connection on the processed depth shift feature map and an input multi-dimensional depth feature map to obtain a final output feature map. According to the method, effective extraction of the depth features of the high-dimensional data of the spectrum video and structural characterization with more fine granularity can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of hyperspectral image processing and video understanding, and in particular relates to a deep feature characterization method based on a high-dimensional shift field. Background Art

[0002] Due to the widespread and successful application of deep learning technology, the field of computer vision has made great progress in recent years. However, the objects processed in the current mainstream visual research field are mainly RGB images. With the development of computational spectral imaging technology, it is possible to obtain spectral images (spectral video high-dimensional data) at video-level frame rates, which has important applications in industrial detection, combustion diagnosis and analysis, hazardous gas leak warning, public transportation safety and other fields. Compared with conventional RGB images, spectral video data has higher data dimensions in time and spectral dimensions, which poses a great challenge to data processing and analysis. At present, deep convolutional networks mostly perform feature extraction on RGB images. How to use existing deep convolutional networks to process and deeply structure higher-dimensional data (hyperspectral images, videos, spectral video data, etc.) is an urgent problem to be solved. Summary of the invention

[0003] Purpose of the invention: The technical problem to be solved by the present invention is to provide a deep feature characterization method based on a high-dimensional shift field in view of the deficiencies of the prior art. The present invention designs a voxel shift field specifically for structured characterization of spectral video data, which increases the processing dimension of the deep network without changing the main structure of the current deep convolutional network.

[0004] The method of the present invention comprises the following steps:

[0005] Step 1: Reorganize the dimensions of the spectral video data and map the spectral channels, extract features through a deep convolutional network, and obtain a multi-dimensional deep feature map of the spectral video data.

[0006] Step 2: Multi-dimensional deep feature map Each voxel in predicts a shift bias O in the three-dimensional spatiotemporal space;

[0007] Step 3: Multi-dimensional deep feature map Sampling is performed in the multidimensional space of the spectral video, and the interpolation weights are calculated to obtain the depth shift feature map after the voxel shift field enhancement processing.

[0008] Step 4: Use convolution kernel to calibrate the depth shift feature map After processing, the multi-dimensional depth feature map of the input Perform residual connection to obtain the final output feature map

[0009] Step 1 includes:

[0010] Step 1-1, input spectral video data Among them, m and n represent the length and width of the spatial dimension of the spectral video, c represents the number of channels in the spectral dimension, and z represents the number of frames of the spectral video. represents the real number space;

[0011] Step 1-2, combine the time dimension of the spectral video data with the training batch dimension B in the deep network Φ (such as resnet); at the same time, replace the number of input channels of the first convolutional layer of the deep network Φ with the number of channels c of the spectral dimension in the spectral video data S;

[0012] Step 1-3, extract features from the spectral video data S through the deep network Φ to obtain a multi-dimensional deep feature map of the spectral video Among them, B represents the batch size of the deep network Φ training, T represents the time dimension size of the deep network feature map, H and W represent the multi-dimensional deep feature map The length and width of the spatial dimension, C represents the multi-dimensional depth feature map The number of channels.

[0013] Step 2 includes:

[0014] Step 2-1: Multi-dimensional depth feature map Considered as a set of B*T*H*W*C voxels in total;

[0015] Step 2-2, for the multi-dimensional depth feature map in the three-dimensional spatiotemporal space For each voxel, predict the three-dimensional displacement bias O of the voxel in the three-dimensional space-time space: the three-dimensional displacement bias O of each voxel is determined by the pre-initialized displacement bias O init and the displacement O predicted by learning pred composition.

[0016] Step 2-2 includes: pre-initializing the displacement offset O init It is used to assign different timing shift values ​​to different channels in the depth feature map, where the multidimensional depth feature map The displacement offset O of the t-th frame in the time dimension is the i-th channel in init (i,t) is defined as:

[0017]

[0018] Based on multi-dimensional deep representation graph The displacement O predicted by learning pred Through the three-dimensional convolution Conv3D(·) operation:

[0019]

[0020] For the final 3D displacement offset The displacement offset O set by the pre-initialization init and the displacement O predicted by learning pred Adding them together gives:

[0021] O=O init +O pred .

[0022] Step 3 includes:

[0023] Step 3-1, multi-dimensional deep feature map With five dimensions B, T, H, W, C, for the feature map of the b-th batch and the i-th channel The voxel of the i-th channel at the spatiotemporal position coordinate (x, y, t) is represented as Where x and y represent the horizontal and vertical coordinates of the spatial dimension respectively, and t represents the coordinate of the time dimension. For the voxel at (x, y, t) of the i-th channel, the process of the voxel shift field is expressed as:

[0024]

[0025] in Represents the voxel representation at the adjacent spatial (x+dx,y+dy,t+dt) position, The feature map represents the three-dimensional displacement of voxels. (dx, dy, dt) is the offset corresponding to each voxel, where dx, dy, and dt represent the offset sizes along the x, y, and t directions respectively. The sizes of dx, dy, and dt are given by the three-dimensional displacement bias O.

[0026] in represents the feature map after being processed by the voxel shift field, Represents the feature map before voxel shift field processing; Step 3-2, obtain by bilinear interpolation The formula is:

[0027]

[0028] Where N = 8, representing the 8 corner points of the sampling grid cube around the sampling point; (Δx, Δy, Δt) represents the nth corner point among the 8 adjacent corner points. th The distance from the corner point, where Δx represents the distance from the corner point along the x direction, Δy represents the distance from the corner point along the y direction, Δt represents the distance from the corner point along the t direction, and w(Δx, Δy, Δt) represents the weight coefficient of bilinear interpolation;

[0029] Step 3-3, use GPU parallel computing to perform multi-dimensional depth feature maps Each voxel in the image is subjected to the operation of step 3-1 to step 3-2 to obtain a depth shift feature map after the voxel shift field enhancement processing.

[0030] Step 4 includes:

[0031] Step 4-1, by using differential attention at the channel level Original feature map After enhancement, the final output of the voxel shift field is obtained:

[0032]

[0033] Where GAP stands for global average pooling, FC stands for fully connected layer, represents element-wise multiplication, Represents element-wise addition.

[0034] The present invention also provides an application of the method, wherein the process of the voxel shift field is nested in each layer of a deep network for target detection, semantic segmentation or video action positioning tasks.

[0035] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the described method.

[0036] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction is run on a computer, the steps of the method described are executed.

[0037] In summary, the present invention realizes effective feature extraction and enhancement processing of spectral video data through processing of spectral video data, including dimensional reorganization, spectral channel mapping, three-dimensional spatiotemporal shift prediction and enhancement processing of depth feature maps, providing subsequent applications with richer and more fine-grained information optimized for task characteristics.

[0038] The present invention has the following beneficial effects: (1) The voxel shift field proposed in the present invention has good portability and can be plug-and-play without changing the original network structure. Any feature map with structural adaptation can be embedded to enhance multi-dimensional information.

[0039] (2) The voxel shift field has a small amount of computation and video memory usage, and has a relatively friendly computational overhead and resource usage.

[0040] (3) The present invention can increase the dimension of the data objects processed by the classic deep convolutional network without changing its structure, and can achieve effective deep feature extraction and more fine-grained structured representation of high-dimensional spectral video data. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.

[0042] Figure 1 It is a flow chart of the method of the present invention.

[0043] Figure 2 is a schematic diagram of the voxel displacement field.

[0044] Figure 3 This is a schematic diagram of the voxel shift field in the ResNet Block.

[0045] Figure 4 It is a schematic diagram comparing the existing target detector Faster RCNN and the method of the present invention in the task of spectral video gaseous target detection. DETAILED DESCRIPTION

[0046] like Figure 1 As shown, an embodiment of the present invention provides a method for characterizing depth features based on a high-dimensional shift field, comprising the following steps:

[0047] Step 1, dimension reorganization and spectral channel mapping processing are performed on the spectral video data, so that feature extraction can be performed through a conventional deep convolutional network to obtain a multi-dimensional deep feature map F of the spectral video data;

[0048] Step 2: Multi-dimensional deep feature map Each voxel in predicts a displacement in the three-dimensional space-time space. Specifically, the three-dimensional voxel displacement field is initialized by the pre-set temporal displacement bias O init and the displacement O predicted by learning pred composition;

[0049] Step 3: Multi-dimensional deep feature map Sampling is performed in the multidimensional space of the spectral video, and the interpolation weights are calculated to obtain the depth shift feature map after the voxel shift field enhancement processing.

[0050] Step 4: Use conventional convolution kernels to calibrate the depth shift feature map After processing, the multi-dimensional depth feature map of the input Perform residual connection to obtain the final output feature map The overall process diagram of the voxel shift field is as follows: Figure 2shown.

[0051] Step 1 includes:

[0052] Step 1-1, for spectral video data Where m and n represent the length and width of the spatial dimension of the spectral video, c represents the number of channels in the spectral dimension, and z represents the number of frames of the spectral video. Since the spectral video data S is expanded in spectral dimension and time dimension compared with the conventional RGB image, it is necessary to perform dimension reorganization and spectral channel mapping on the spectral video data S;

[0053] Step 1-2, combine the time dimension of the spectral video data with the training batch dimension B in the deep network Φ; at the same time, replace the number of input channels of the first convolutional layer of the deep network Φ with the spectral channel c in the spectral video data S. Through the above operations, the spectral video data can be trained with minimal modifications to conventional deep networks, such as ResNet, ShuffleNet, etc. Perform feature extraction;

[0054] Step 1-3, extract features from the spectral video data S through the deep network Φ to obtain the deep feature map of the spectral video Among them, B represents the training batch size of the deep network training, T represents the number of frames of the spectral video, and H and W represent the deep feature maps respectively. The length and width of the spatial dimensions of C represent the depth The number of channels of the feature map.

[0055] Step 2 includes:

[0056] Step 2-1: Deep feature map of spectral video Consider it as a collection of B*T*H*W*C voxels in total;

[0057] Step 2-2, for the depth feature map in the three-dimensional spatiotemporal space For each voxel of , a three-dimensional displacement bias O in the three-dimensional space-time space is predicted for it. Specifically, the three-dimensional displacement bias O of each voxel is composed of the pre-initialized temporal displacement bias O init and the displacement O predicted by learning pred composition;

[0058] Step 2-3, pre-initialize the timing shift offset O init It mainly assigns different time shift values ​​to different channels in the depth feature map, among which the multidimensional depth feature map The i-th channel in the time dimension has a temporal displacement offset O of the t-th frame. init (i,t) is defined as:

[0059]

[0060] Step 2-4, deep feature map based on spectral video The displacement O predicted by learning pred O pred Through the 3D convolution Conv3D(·) operation:

[0061]

[0062] Steps 2-5, for the final 3D displacement offset The timing offset O set by the pre-initialization init and the displacement O predicted by learning pred Adding them together gives:

[0063] O=O init +O pred .

[0064] Step 3 includes:

[0065] Step 3-1, due to the feature map extracted by the deep convolutional network With five dimensions B, H, W, C, T, for the feature map of the b-th batch and the i-th channel The voxel at the spatiotemporal position coordinate (x, y, t) of the b-th batch and the i-th channel is represented as Where x and y represent the coordinates of the spatial dimension, and t represents the coordinate of the time dimension. For the voxel at (x, y, t) in the bth batch and the i-th channel, the process of the voxel shift field can be expressed as:

[0066]

[0067] in Represents the voxel representation at the adjacent spatial (x+dx,y+dy,t+dt) position, The feature map represents the three-dimensional displacement of the voxel. (dx, dy, dt) is the offset corresponding to each voxel. This offset is learnable. The adjusted sampling position is obtained by adding the original coordinate position (on the input feature map) to the corresponding offset.

[0068] Step 3-2, due to the generation of sampling positions The predicted offset values ​​in the x, y, and t directions are floating point values. The sampling point position is not located at the original voxel position, so it needs to be obtained by bilinear interpolation. Bilinear interpolation is an interpolation method based on the weighted average of the four nearest neighbor points, which is used to estimate the intermediate value in two dimensions. As shown in the following formula:

[0069]

[0070] Where N = 8, representing the 8 corner points of the sampling grid cube around the sampling point, (Δx, Δy, Δt) represents the nth corner point among the 8 adjacent corner points. th The distance from the corner point, where Δx represents the distance from the corner point along the x direction, Δy represents the distance from the corner point along the y direction, Δt represents the distance from the corner point along the t direction, w(Δx,Δy,Δt) represents the weight coefficient of bilinear interpolation, which is given according to (Δx,Δy,Δt). The interpolated eigenvalues ​​are multiplied by the corresponding convolution kernel weights, and these values ​​are summed to obtain the final convolution output;

[0071] Step 3-3, feature map The above operations are performed on each voxel in the image, and GPU parallel computing is used to achieve high-speed implementation.

[0072] Step 4 includes:

[0073] Step 4-1, considering that the shifted sampling area exceeds the boundary and the feature value becomes 0, the feature map processed by the voxel shift field is processed by differential attention at the channel level Original feature map Perform enhancement to obtain the output of the voxel shift field:

[0074]

[0075] Where GAP stands for global average pooling, FC stands for fully connected layer, represents element-wise multiplication, Represents element-wise addition;

[0076] In step 4-2, the voxel shift field mentioned above is embedded in each layer of the deep network. Figure 3 The application of voxel shift field in ResNet backbone network is demonstrated. For ResNet backbone network, voxel shift field is embedded into each ResNet Block. The deep feature map enhanced by voxel shift features can be trained through supervised learning to obtain a more robust and universal feature representation. Therefore, it has higher accuracy and recall in various downstream tasks such as target detection, semantic segmentation, video action localization, etc.

[0077] Specifically, if Figure 4As shown in the figure, in the task of detecting gaseous targets in spectral video, the existing classic target detector Faster RCNN cannot locate the target boundary well. Adding the voxel displacement field proposed in the present invention to the Faster RCNN backbone network ResNet Block can potentially model the changing shape of the target in three-dimensional space, so the method proposed in the present invention can better locate the boundary of gaseous targets.

[0078] The present invention provides a method for characterizing deep features based on a high-dimensional shift field. There are many methods and approaches to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention. All components not specified in this embodiment can be implemented using existing technologies.

Claims

1. A deep feature characterization method based on high-dimensional shift field, characterized in that: The following steps are involved: Step 1: Reorganize the dimensions of the spectral video data and map the spectral channels, extract features through a deep convolutional network, and obtain a multi-dimensional deep feature map of the spectral video data. Step 2: Multi-dimensional deep feature map Each voxel in predicts a shift bias O in the three-dimensional spatiotemporal space; Step 3: Multi-dimensional deep feature map Sampling is performed in the multidimensional space of the spectral video, and the interpolation weights are calculated to obtain the depth shift feature map after the voxel shift field enhancement processing. Step 4: Use convolution kernel to calibrate the depth shift feature map After processing, the multi-dimensional depth feature map of the input Perform residual connection to obtain the final output feature map 2. The method according to claim 1, characterized in that Step 1 includes: Step 1-1, input spectral video data Among them, m and n represent the length and width of the spatial dimension of the spectral video, c represents the number of channels in the spectral dimension, and z represents the number of frames of the spectral video. represents the real number space; Step 1-2, combining the time dimension of the spectral video data with the training batch dimension B in the deep network Φ; at the same time, replacing the number of input channels of the first convolutional layer of the deep network Φ with the number of channels c of the spectral dimension in the spectral video data S; Step 1-3, extract features from the spectral video data S through the deep network Φ to obtain a multi-dimensional deep feature map of the spectral video Among them, B represents the batch size of the deep network Φ training, T represents the time dimension size of the deep network feature map, H and W represent the multi-dimensional deep feature map The length and width of the spatial dimension, C represents the multi-dimensional depth feature map The number of channels.

3. The method according to claim 2, characterized in that Step 2 includes: Step 2-1: Multi-dimensional depth feature map Considered as a set of B*T*H*W*C voxels in total; Step 2-2, for the multi-dimensional depth feature map in the three-dimensional spatiotemporal space For each voxel, predict the three-dimensional displacement bias O of the voxel in the three-dimensional space-time space: the three-dimensional displacement bias O of each voxel is determined by the pre-initialized displacement bias O init and the displacement O predicted by learning pred composition.

4. The method according to claim 3, characterized in that Step 2-2 includes: pre-initializing the displacement offset O init It is used to assign different timing shift values ​​to different channels in the depth feature map, where the multidimensional depth feature map The displacement offset O of the t-th frame in the time dimension is the i-th channel in init (i,t) is defined as: Based on multi-dimensional deep representation graph The displacement O predicted by learning pred Through the three-dimensional convolution Conv3D(·) operation: For the final 3D displacement offset The displacement offset O set by the pre-initialization init and the displacement O predicted by learning pred Adding them together gives: O=O init +O pred 。 5. The method according to claim 4, characterized in that Step 3 includes: Step 3-1, multi-dimensional deep feature map With five dimensions B, T, H, W, C, for the feature map of the b-th batch and the i-th channel The voxel of the i-th channel at the spatiotemporal position coordinate (x, y, t) is represented as Where x and y represent the horizontal and vertical coordinates of the spatial dimension respectively, and t represents the coordinate of the time dimension. For the voxel at (x, y, t) of the i-th channel, the process of the voxel shift field is expressed as: in Represents the voxel representation at the adjacent spatial (x+dx,y+dy,t+dt) position, The feature map represents the three-dimensional displacement of voxels. (dx, dy, dt) is the offset corresponding to each voxel, where dx, dy, and dt represent the offset sizes along the x, y, and t directions respectively. The sizes of dx, dy, and dt are given by the three-dimensional displacement bias O. in represents the feature map after being processed by the voxel shift field, Represents the feature map before voxel shift field processing; Step 3-2: Obtain by bilinear interpolation The formula is: Where N = 8, representing the 8 corner points of the sampling grid cube around the sampling point; (Δx, Δy, Δt) represents the nth corner point among the 8 adjacent corner points. th The distance from the corner point, where Δx represents the distance from the corner point along the x direction, Δy represents the distance from the corner point along the y direction, Δt represents the distance from the corner point along the t direction, and w(Δx, Δy, Δt) represents the weight coefficient of bilinear interpolation; Step 3-3, use GPU parallel computing to perform multi-dimensional depth feature maps Each voxel in the image is subjected to the operation of step 3-1 to step 3-2 to obtain a depth shift feature map after the voxel shift field enhancement processing.

6. The method according to claim 5, characterized in that Step 4 includes: Step 4-1, by using differential attention at the channel level Original feature map After enhancement, the final output of the voxel shift field is obtained: Where GAP stands for global average pooling, FC stands for fully connected layer, represents element-wise multiplication, Represents element-wise addition.

7. Use of the method according to claim 6, characterized in that The voxel shift field process is nested in each layer of the deep network for object detection, semantic segmentation or video action localization tasks.

8. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 6 are executed.