Unmanned aerial vehicle video coding and decoding model based on deep learning and processing method thereof
Through the deep learning-based drone video encoding and decoding model, combined with CNN and Transformer modules, multi-dimensional adaptive codec and video supplement module are designed, the universality and efficiency of drone video encoding and decoding are solved, and efficient compression and rapid codec of high-quality videos are achieved.
Patent Information
- Application Number
- CN202510260500.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-06
AI Technical Summary
The existing drone video encoding and codec technology lacks versatility and efficiency in unmanned aircraft, and it is difficult to adapt to the changing shooting environment and the video encoding and codec requirements of wide coverage areas.
A drone video encoding and decoding model based on deep learning is adopted, combined with the CNN network structure, Transformer learning module and traditional compression algorithm, a multi-dimensional adaptive codec module and video supplement module are designed to realize differentiated encoding of image frame dimensions and spatial dimensions, and the video missing information is supplemented through the U-Net network module.
It realizes efficient compression of high-quality videos, supports video encoding and decoding under multi-environmental conditions, has fast frame processing speed, high compression ratio and wide applicability, and solves the versatility and efficiency of drone video encoding and decoding.
Smart Images

Figure CN120111244A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a compression model, in particular to a UAV video encoding and decoding model based on deep learning and a processing method thereof. Background Art
[0002] Drones (or unmanned aerial vehicles, UAVs) have received increasing attention over the past decade due to their flexible, extensive, and dynamic spatial sensing availability. The Commercial Drone Market Report describes that the global commercial drone market will reach 501.4 billion by 2028. These drones can be effectively used for autonomous facility maintenance, scientific discovery, or agricultural monitoring through the cameras they are equipped with. Therefore, higher requirements are placed on efficient compression and low-cost storage solutions for visual data collected by drones, which makes video coding more closely related to drones.
[0003] Video coding aims to represent video signals with fewer bits while ensuring high-quality reconstruction of signals, so as to achieve efficient storage and transmission. Specifically, there are a large number of semantically related or invalid areas in image and video information, which in turn cause the storage and transmission of redundant information. Therefore, video data can be compressed by removing such redundant information. The traditional hybrid coding framework, which is based on the combination of motion compensation and transform coding of statistical characteristics, has achieved remarkable results in the field of data compression. This framework forms the basis of many common video compression standards at home and abroad, such as the H.26L series and the MPEG series. However, with the continuous evolution of deep learning technology and the continuous diversification of application needs, the research on video coding technology is no longer limited to improving compression performance. On the contrary, the focus of research is gradually shifting to new dimensions such as adaptability and user interactivity. Therefore, the development of video coding technology in recent years has shown a dual trend: on the one hand, in-depth exploration of the potential of deep learning in video coding and decoding to seek further improvement of compression performance; on the other hand, active exploration of the development of emerging branches such as scalable coding and multi-view coding. Although deep learning-based video coding and decoding solutions have achieved great success in computer network applications, existing methods have relatively little research on video coding and decoding technology in unmanned aerial vehicles, and all of them are implemented by initially expanding existing coding methods for a few perspective features, lacking the versatility and efficiency of unmanned aerial vehicle video coding and decoding. Summary of the invention
[0004] The purpose of the present invention is to provide a UAV video encoding and decoding model based on deep learning. The present invention has the characteristics of wide versatility, strong applicability and high video quality.
[0005] The technical solution of the present invention is as follows: a drone video coding and decoding model based on deep learning, comprising an image compression benchmark module based on deep learning, a multidimensional adaptive coding and decoding module based on an attention mechanism, and a video supplementation module based on diffusion; the image compression benchmark module is used to extract high-fidelity, low-redundancy image potential representations, and restore the image potential representations into reconstructed video frames; the multidimensional adaptive coding and decoding module is used to fuse image frame dimensions with image space dimensions to obtain a multidimensional feature map; the video supplementation module is used to generate a more complete and accurate feature representation as video supplementary information, thereby improving the quality and integrity of the entire video.
[0006] In the aforementioned deep learning-based drone video encoding and decoding model, the image compression benchmark module includes a CNN network structure for local semantic clues, a Transformer learning module for global semantic features, and a traditional compression algorithm module; the CNN network structure is used to capture local features in the image and use local features to understand more complex global semantic information; the Transformer learning module is used to capture global features in the image and construct a core encoder and a core decoder; the traditional compression algorithm module includes a quantization module, an entropy encoder and an entropy decoder, which are used to effectively remove redundant information, thereby achieving efficient compression while maintaining high quality.
[0007] In the aforementioned deep learning-based drone video encoding and decoding model, the multi-dimensional adaptive encoding and decoding module includes an image frame dimension attention module and an image space dimension attention module; the image frame dimension attention module is used to realize differentiated encoding of image frame sequences, and solve the problem that the background of image frames changes with time due to position variability during drone shooting; the image space dimension attention module is used to realize differentiated encoding of image frame areas, and solve the problem that the importance of different areas of the same image varies greatly due to the fact that the aerial view during drone shooting usually covers a wide range of ground.
[0008] In the aforementioned deep learning-based drone video encoding and decoding model, the image frame dimension attention module includes a channel attention module and an inter-frame correlation analysis module; the channel attention module contains a global average pooling layer and two fully connected layers to obtain the weight coefficient of each channel; the inter-frame correlation analysis module contains a linear layer, a softmax layer and an inter-frame correlation matrix module to obtain the inter-frame correlation probability by passing the image potential representation through the linear layer and the softmax layer, and multiply the inter-frame correlation probability by the weight coefficient to obtain the frame timing weight, and then multiply the frame timing weight by the image potential representation to obtain the frame timing feature map.
[0009] In the aforementioned deep learning-based drone video encoding and decoding model, the image spatial dimension attention module includes a local feature efficient extractor based on CNN and containing four convolutional layers and a spatial attention learning module; the local feature efficient extractor is used to enhance the attention of key effective pixels, achieve high-fidelity compression in semantic information-rich areas and high-ratio compression in other unimportant areas; the spatial attention learning module is used to reduce the number of channels of the frame temporal feature map through three linear layers, pull the dimensions except the number of channels into a vector, transpose and multiply two of them to obtain an attention score matrix, and then perform a softmax operation to obtain the correlation weight of each pixel point to other pixels. Finally, the attention score matrix is multiplied with the untransposed vector, and the convolutional layer is used to increase the dimension and multiply it with the frame temporal feature map to obtain a multidimensional feature map, thereby realizing the mining of the weight relationship between different pixels in the spatial neighborhood.
[0010] In the aforementioned deep learning-based drone video encoding and decoding model, the video supplement module includes a U-Net network module, a conditional encoder module and a noise module; the U-Net network module includes an upsampling module and a downsampling module, and both the upsampling module and the downsampling module include a convolution block and a cross-attention block, which are used to downsample and upsample the input information and predict noise; the conditional encoder module uses the reconstructed feature map restored after entropy decoding and inverse quantization as a conditional input to better supplement the missing or incomplete parts of the video, thereby improving the quality and integrity of the video; the noise module is used to generate random noise according to a Gaussian distribution in the training stage, and add the random noise to the video supplement information for learning; in the generation stage, an initial random noise is first generated, and then in the subsequent iteration process, the noise is gradually removed and the video supplement information that meets the requirements is generated.
[0011] The processing method of the above-mentioned drone video encoding and decoding model includes the following steps:
[0012] S1, input original video sequence X = {x 1 ,x 2 ,...,x T}, where each x t is a video frame;
[0013] S2, video frame x t The image potential representation is obtained by processing through the core encoder of the image compression benchmark module, and the image potential representation is processed by the multi-dimensional adaptive coding module to obtain a multi-dimensional feature map;
[0014] S3, converting the multidimensional feature map into a bit stream through a quantization module and an entropy encoder;
[0015] S4, performing entropy decoding and dequantization on the bit stream to obtain a reconstructed feature map;
[0016] S5, input the reconstructed feature map into the video supplementation module to obtain a supplementary feature map, then splice the supplementary feature map with the reconstructed feature map to obtain video supplementary information, and input it into the core decoder to obtain a reconstructed video frame;
[0017] S6. Repeat steps S1-S5 to obtain a series of reconstructed video sequences consisting of reconstructed video frames.
[0018] In the aforementioned processing method, in step S2, the multi-dimensional adaptive coding module processing includes image frame dimension attention module processing and image space dimension attention module processing, and the image frame dimension attention module processing includes the following steps:
[0019] S2311, performing global average pooling on the potential representation of the image through the global average pooling layer of the channel attention module to obtain a global feature description representing the channel;
[0020] S2312, passing the global feature description through two fully connected layers; the first fully connected layer reduces the number of channels, and the second fully connected layer restores the dimension to the original dimension to obtain the weight coefficient of each channel;
[0021] S2313, the image potential representation is simultaneously transformed through a linear layer, and the correlation between different image frames is calculated through a softmax layer to obtain the inter-frame correlation probability;
[0022] S2314, performing a dot multiplication of the weight coefficient of each channel and the inter-frame association probability to obtain a frame timing weight;
[0023] S2315. The frame time series weight is normalized through a softmax layer and point-multiplied with the image potential representation to obtain a frame time series feature map.
[0024] In the above processing method, in step S2, the image space dimension attention module processes as follows:
[0025] S2321, performing a convolution operation on the frame time series feature map through a local feature efficient extractor having four convolution layers to obtain feature maps at different spatial positions and scales;
[0026] S2322, reducing the number of channels of the feature map through three linear layers;
[0027] S2323, pull the dimensions except the number of channels into a vector, and transpose and multiply two of them to obtain an attention score matrix;
[0028] S2324, performing a softmax operation on the attention score matrix to obtain the relevance weight of each pixel to other pixels;
[0029] S2325. Multiply the attention score matrix with the untransposed vector, and use the convolution layer to perform a dimensionality increase operation to restore the number of channels to the same dimension as the feature map, and then perform a dot multiplication with the frame timing feature map to obtain a multidimensional feature map that integrates the image frame dimension and image space dimension information.
[0030] In the above processing method, step S5 is specifically as follows:
[0031] S51, taking the reconstructed feature map as a condition, inputting it into a conditional encoder module to obtain a conditional guidance vector;
[0032] S52, the noise module generates random initial noise and inputs it into the UNet network module, and gradually downsamples and upsamples to remove the noise;
[0033] S53, UNet network module first passes through the convolution block to reduce the number of channels in the process of step-by-step downsampling, and then passes through the cross-attention block and the conditional guidance vector to perform cross-attention; in the process of step-by-step upsampling, the number of channels is first increased through the convolution block, and then passes through the cross-attention block and the conditional guidance vector to perform cross-attention; finally, the predicted noise is obtained;
[0034] S54, subtract the predicted noise from the random initial noise as the secondary noise, and input it into the UNet network module again to remove the noise, so as to obtain a supplementary feature map;
[0035] S55. Concatenate the supplementary feature map with the reconstructed feature map to obtain video supplementary information, and input the video supplementary information into a core decoder to obtain a reconstructed video frame.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] The present invention designs a UAV video encoding and decoding model. Firstly, a deep learning-based image compression benchmark module combining a CNN network structure, a Transformer learning module and a traditional compression algorithm module is developed to basically realize the encoding and decoding of high-quality video frames and effectively remove redundant information, thereby achieving efficient compression while maintaining high quality, and reaching an advanced level on a public data set; a multi-dimensional adaptive encoding and decoding module based on an attention mechanism is designed to fuse the image frame dimension with the image space dimension to realize differentiated encoding of image frame sequences and image frame regions, thereby solving the problem that the background of image frames changes with time due to position variability during UAV shooting, and solving the problem that the importance of different regions of the same image varies greatly due to the relatively wide ground range covered by the bird's-eye view during the shooting of UAVs; finally, a diffusion-based video supplement module is introduced to merge with the image compression benchmark module to generate a more complete and accurate feature representation as video supplementary information, thereby improving the quality and integrity of the entire video.
[0038] The present invention supports video encoding and decoding under multiple environmental conditions, with a frame processing speed of 20 frames per second. In conventional scenarios, the MJPEG standard is adopted, and the compression ratio is not less than 200:1. It has strong versatility and wide applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of the overall operation steps of the present invention;
[0040] Figure 2 It is a working flow chart of the core encoder and the core decoder in the image compression benchmark module of the present invention;
[0041] Figure 3 It is a working flow chart of the multi-dimensional adaptive coding and decoding module in the present invention;
[0042] Figure 4 It is a workflow diagram of the diffusion-based video supplementation module described in the present invention. DETAILED DESCRIPTION
[0043] The present invention will be further described below in conjunction with the embodiments, but they are not intended to limit the present invention.
[0044] Example:
[0045] A deep learning-based unmanned aerial vehicle video encoding and decoding model comprises a deep learning-based image compression benchmark module, an attention mechanism-based multidimensional adaptive encoding and decoding module and a diffusion-based video supplementation module; the image compression benchmark module is used to extract high-fidelity, low-redundancy image potential representations, and restore the image potential representations into reconstructed video frames; the multidimensional adaptive encoding and decoding module is used to fuse image frame dimensions with image space dimensions to obtain a multidimensional feature map; the video supplementation module is used to generate a more complete and accurate feature representation as video supplementation information, thereby improving the quality and integrity of the entire video.
[0046] The image compression benchmark module includes a CNN network structure for local semantic clues, a Transformer learning module for global semantic features, and a traditional compression algorithm module; the CNN network structure is used to capture local features in the image and use the local features to understand more complex global semantic information; the Transformer learning module is used to capture global features in the image and construct a core encoder and a core decoder; the core encoder is used to extract high-fidelity, low-redundancy potential image representation, and the core decoder is used to restore the potential image representation into a reconstructed video frame; the traditional compression algorithm module includes a quantization module, an entropy encoder and an entropy decoder, which are used to effectively remove redundant information, thereby achieving efficient compression while maintaining high quality.
[0047] The multi-dimensional adaptive coding and decoding module includes an image frame dimension attention module and an image space dimension attention module; the image frame dimension attention module is used to realize differential encoding of image frame sequences, so as to solve the problem that the background of image frames changes with time due to the variability of positions during drone shooting; the image space dimension attention module is used to realize differential encoding of image frame areas, so as to solve the problem that the importance of different areas of the same image varies greatly due to the fact that the aerial view during drone shooting usually covers a wide range of ground.
[0048] The image frame dimension attention module includes a channel attention module and an inter-frame correlation analysis module; the channel attention module contains a global average pooling layer and two fully connected layers to obtain the weight coefficient of each channel; the inter-frame correlation analysis module contains a linear layer, a softmax layer and an inter-frame correlation matrix module to obtain the inter-frame correlation probability by passing the image potential representation extracted from the core encoder through the linear layer and the softmax layer; and multiplying the inter-frame correlation probability with the weight coefficient of each channel to obtain the frame timing weight, and then multiplying the frame timing weight with the image potential representation after passing through the softmax layer to obtain the frame timing feature map.
[0049] The image spatial dimension attention module includes a local feature efficient extractor based on CNN and containing four convolutional layers and a spatial attention learning module; the local feature efficient extractor is used to improve the attention of key effective pixels, realize high-fidelity compression of semantic information-rich areas and high-ratio compression of other unimportant areas; the spatial attention learning module is used to reduce the number of channels of the frame time series feature map obtained from the image frame dimension attention module through three linear layers, pull the dimensions except the number of channels into a vector, transpose two of them and multiply them to obtain an attention score matrix, and then perform a softmax operation to obtain the correlation weight of each pixel point to other pixels, and finally multiply the attention score matrix with the untransposed vector, and use the convolution layer to increase the dimension and multiply it with the frame time series feature map to obtain a multidimensional feature map, so as to realize the mining of the weight relationship between different pixels in the spatial neighborhood.
[0050] The video supplement module includes a U-Net network module, a conditional encoder module and a noise module; the U-Net network module includes an upsampling module and a downsampling module, both of which include a convolution block and a cross-attention block, which are used to downsample and upsample the input information and predict noise; the conditional encoder module uses the reconstructed feature map restored from entropy decoding and dequantization as a conditional input to better supplement the missing or incomplete parts of the video, thereby improving the quality and integrity of the video; the noise module is used to generate random noise according to Gaussian distribution in the training stage, and add the random noise to the video supplement information to allow the U-Net network module to learn how to recover the video supplement information therefrom; in the generation stage, an initial random noise is first generated, and then in the subsequent iteration process, the noise is gradually removed and the video supplement information that meets the requirements is generated.
[0051] The processing method of the above-mentioned drone video encoding and decoding model is as follows: Figure 1-4 As shown, the following steps are included:
[0052] S1, input original video sequence X = {x 1 ,x 2 ,...,x T}, where each x t It is a video frame, which is obtained by splitting the video sequence into frames.
[0053] S2, video frame x t The image potential representation is obtained by encoding through the core encoder of the image compression benchmark module, and the image potential representation is processed by the multi-dimensional adaptive encoding module to obtain a multi-dimensional feature map;
[0054] Step S2 is specifically as follows:
[0055] S21, video frame xt After normalization operation, the pixel values are normalized to the interval [0,1] to obtain a normalized video frame;
[0056] S22, the normalized video frame is encoded by the core encoder of the image compression benchmark module, the core encoder includes multiple convolutional layers and generalized divisive normalization (GDN), and the image potential representation y is obtained. t ;
[0057] S23, the image potential representation y t After being processed by the multi-dimensional adaptive coding module, an enhanced multi-dimensional feature map is obtained that combines the image frame dimension and image space dimension information.
[0058] Step S23 is specifically as follows:
[0059] S231, the image potential representation y t After being processed by the image frame dimension attention module, the frame temporal feature map is obtained;
[0060] S232. Process the frame temporal feature map through the image space dimension attention module to obtain a multi-dimensional feature map.
[0061] Step S231 is specifically as follows:
[0062] S2311, the image potential representation y t Input channel attention module, perform global average pooling through the global average pooling layer to obtain a global feature description representing the channel;
[0063] S2312, pass the global feature description through two fully connected layers; the first fully connected layer reduces the number of channels, and the second fully connected layer restores the dimension to the original dimension, and limits the output value between 0 and 1 through a Sigmoid activation function to obtain the weight coefficient of each channel;
[0064] S2313, image potential representation y t At the same time, a linear layer of the inter-frame correlation analysis module performs feature transformation, and the correlation degree between different image frames is calculated through the softmax layer to obtain the inter-frame correlation probability;
[0065] S2314, performing a dot multiplication of the weight coefficient of each channel and the inter-frame association probability to obtain a frame timing weight;
[0066] S2315, normalize the frame time sequence weight through the softmax layer and compare it with the image potential representation y t Point product to obtain the frame temporal feature map.
[0067] Step S232 is specifically as follows:
[0068] S2321, performing a convolution operation on the frame time series feature map through a local feature efficient extractor having four convolution layers to obtain feature maps at different spatial positions and scales;
[0069] S2322, reducing the number of channels of the feature map through three linear layers;
[0070] S2323, pull the dimensions except the number of channels into a vector, and transpose and multiply two of them to obtain an attention score matrix;
[0071] S2324, performing a softmax operation on the attention score matrix to obtain the relevance weight of each pixel to other pixels;
[0072] S2325. Multiply the attention score matrix with the untransposed vector, and use the convolution layer to perform a dimensionality increase operation to restore the number of channels to the same dimension as the feature map, and then dot-multiply it with the frame temporal feature map to obtain an enhanced multi-dimensional feature map that combines the image frame dimension and image space dimension information.
[0073] S3. Convert the multi-dimensional feature map into a bit stream through the quantization module and entropy encoder of the traditional compression algorithm module.
[0074] Step S3 is specifically as follows:
[0075] S31, mapping the continuous multi-dimensional feature map to a finite number of discrete quantization levels through a quantization module, reducing the value possibilities of the data, thereby reducing the amount of data, and obtaining a quantized multi-dimensional feature map;
[0076] S32, counting the distribution of the quantized multidimensional feature graphs, and calculating the frequency or probability of occurrence of each quantized multidimensional feature graph;
[0077] S33. According to the frequency or probability distribution obtained by statistics, the quantized multi-dimensional feature map is converted into a binary bit stream through an entropy encoder according to coding rules.
[0078] S4. When it is necessary to restore the video frame, the binary bit stream is entropy decoded and dequantized through a quantization module and an entropy decoder to obtain a reconstructed feature map.
[0079] S5, input the reconstructed feature map into the video supplementation module to obtain a supplementary feature map, then splice the supplementary feature map with the reconstructed feature map to obtain video supplementary information, input the video supplementary information into the core decoder to reconstruct the video, and obtain the reconstructed video frame x t .
[0080] Step S5 is specifically as follows:
[0081] S51, taking the reconstructed feature map as a condition, inputting it into a conditional encoder to obtain a conditional guidance vector;
[0082] S52, the noise module generates random initial noise that obeys Gaussian distribution, and uses it as initial supplementary information, inputs it into the UNet network module of the video supplementation module, and gradually downsamples and upsamples to remove noise;
[0083] S53, UNet network module first passes through the convolution block to reduce the number of channels in the process of step-by-step downsampling, and then passes through the cross-attention block and the conditional guidance vector to perform cross-attention; in the process of step-by-step upsampling, the number of channels is first increased through the convolution block, and then passes through the cross-attention block and the conditional guidance vector to perform cross-attention; finally, the predicted noise is obtained;
[0084] S54, subtract the predicted noise output by the UNet network module from the random initial noise as the secondary noise, input the noise into the UNet network module again to remove the noise, repeat the iteration, and obtain a supplementary feature map;
[0085] S55, splicing the supplementary feature map with the reconstructed feature map to obtain video supplementary information, inputting the video supplementary information into the core decoder to reconstruct the video, and obtaining the reconstructed video frame x t .
[0086] S6, repeat steps S1-S5 to obtain a series of reconstructed video frames x t The reconstructed video sequence.
Claims
1. A drone video encoding and decoding model based on deep learning, characterized by: It includes an image compression benchmark module based on deep learning, a multi-dimensional adaptive coding and decoding module based on attention mechanism, and a video supplement module based on diffusion; the image compression benchmark module is used to extract high-fidelity, low-redundancy image potential representation, and restore the image potential representation into a reconstructed video frame; the multi-dimensional adaptive coding and decoding module is used to fuse the image frame dimension with the image space dimension to obtain a multi-dimensional feature map; the video supplement module is used to generate a more complete and accurate feature representation as video supplement information, thereby improving the quality and integrity of the entire video.
2. The deep learning-based drone video encoding and decoding model according to claim 1, characterized in that: The image compression benchmark module includes a CNN network structure for local semantic clues, a Transformer learning module for global semantic features, and a traditional compression algorithm module; the CNN network structure is used to capture local features in the image and use the local features to understand more complex global semantic information; the Transformer learning module is used to capture global features in the image and construct a core encoder and a core decoder; the traditional compression algorithm module includes a quantization module, an entropy encoder and an entropy decoder, which are used to effectively remove redundant information, thereby achieving efficient compression while maintaining high quality.
3. The deep learning-based drone video encoding and decoding model according to claim 1, characterized in that: The multi-dimensional adaptive coding and decoding module includes an image frame dimension attention module and an image space dimension attention module; the image frame dimension attention module is used to realize differential encoding of image frame sequences, so as to solve the problem that the background of image frames changes with time due to the variability of positions during drone shooting; the image space dimension attention module is used to realize differential encoding of image frame areas, so as to solve the problem that the importance of different areas of the same image varies greatly due to the fact that the aerial view during drone shooting usually covers a wide range of ground.
4. The deep learning-based drone video encoding and decoding model according to claim 3 is characterized in that: The image frame dimension attention module includes a channel attention module and an inter-frame correlation analysis module; the channel attention module contains a global average pooling layer and two fully connected layers to obtain the weight coefficient of each channel; the inter-frame correlation analysis module contains a linear layer, a softmax layer and an inter-frame correlation matrix module to obtain the inter-frame correlation probability through the linear layer and the softmax layer of the image potential representation, and multiply the inter-frame correlation probability by the weight coefficient to obtain the frame timing weight, and then multiply the frame timing weight by the image potential representation to obtain the frame timing feature map.
5. The deep learning-based drone video encoding and decoding model according to claim 4 is characterized in that: The image spatial dimension attention module includes a local feature efficient extractor based on CNN and containing four convolutional layers and a spatial attention learning module; the local feature efficient extractor is used to improve the attention of key effective pixels, realize high-fidelity compression of semantic information-rich areas and high-ratio compression of other unimportant areas; the spatial attention learning module is used to reduce the number of channels of the frame temporal feature map through three linear layers, pull the dimensions except the number of channels into a vector, transpose and multiply two of them to obtain an attention score matrix, and then perform a softmax operation to obtain the correlation weight of each pixel point to other pixels. Finally, the attention score matrix is multiplied with the untransposed vector, and the convolutional layer is used to increase the dimension and multiply it with the frame temporal feature map to obtain a multidimensional feature map, so as to realize the mining of the weight relationship between different pixels in the spatial neighborhood.
6. The deep learning-based drone video encoding and decoding model according to claim 1, characterized in that: The video supplementation module includes a U-Net network module, a conditional encoder module and a noise module; the U-Net network module includes an upsampling module and a downsampling module, and the upsampling module and the downsampling module both include a convolution block and a cross attention block, which are used to downsample and upsample input information and predict noise; the conditional encoder module uses the reconstructed feature map restored after entropy decoding and dequantization as a conditional input to better supplement the missing or incomplete parts of the video, thereby improving the quality and integrity of the video; The noise module is used to generate random noise according to Gaussian distribution during the training phase, and add the random noise to the video supplementary information for learning; In the generation stage, an initial random noise is generated first, and then in the subsequent iteration process, the noise is gradually removed and the video supplementary information that meets the requirements is generated.
7. A method for processing a drone video codec model, using the drone video codec model based on deep learning according to any one of claims 1 to 6, characterized in that: The following steps are involved: S1, input original video sequence X = {x1, x2, ..., x T }, where each x t is a video frame; S2, video frame x t The image potential representation is obtained by processing through the core encoder of the image compression benchmark module, and the image potential representation is processed by the multi-dimensional adaptive coding module to obtain a multi-dimensional feature map; S3, converting the multidimensional feature map into a bit stream through a quantization module and an entropy encoder; S4, performing entropy decoding and dequantization on the bit stream to obtain a reconstructed feature map; S5, input the reconstructed feature map into the video supplementation module to obtain a supplementary feature map, then splice the supplementary feature map with the reconstructed feature map to obtain video supplementary information, and input it into the core decoder to obtain a reconstructed video frame; S6. Repeat steps S1-S5 to obtain a series of reconstructed video sequences consisting of reconstructed video frames.
8. The processing method of the drone video encoding and decoding model according to claim 7 is characterized in that: In step S2, the multi-dimensional adaptive coding module processing includes image frame dimension attention module processing and image space dimension attention module processing, and the image frame dimension attention module processing includes the following steps: S2311, performing global average pooling on the potential representation of the image through the global average pooling layer of the channel attention module to obtain a global feature description representing the channel; S2312, passing the global feature description through two fully connected layers; the first fully connected layer reduces the number of channels, and the second fully connected layer restores the dimension to the original dimension to obtain the weight coefficient of each channel; S2313, the image potential representation is simultaneously transformed through a linear layer, and the correlation between different image frames is calculated through a softmax layer to obtain the inter-frame correlation probability; S2314, performing a dot multiplication of the weight coefficient of each channel and the inter-frame association probability to obtain a frame timing weight; S2315. The frame timing weight is normalized through a softmax layer and point-multiplied with the image potential representation to obtain a frame timing feature map.
9. The processing method of the drone video encoding and decoding model according to claim 8 is characterized in that: In step S2, the steps of the image space dimension attention module processing are specifically as follows: S2321, performing a convolution operation on the frame time series feature map through a local feature efficient extractor having four convolution layers to obtain feature maps at different spatial positions and scales; S2322, reducing the number of channels of the feature map through three linear layers; S2323, pull the dimensions except the number of channels into a vector, and transpose and multiply two of them to obtain an attention score matrix; S2324, performing a softmax operation on the attention score matrix to obtain the relevance weight of each pixel to other pixels; S2325. Multiply the attention score matrix with the untransposed vector, and use the convolution layer to perform a dimensionality increase operation to restore the number of channels to the same dimension as the feature map, and then perform a dot multiplication with the frame timing feature map to obtain a multidimensional feature map that integrates the image frame dimension and image space dimension information.
10. The processing method of the drone video encoding and decoding model according to claim 7 is characterized in that: Step S5 is specifically as follows: S51, taking the reconstructed feature map as a condition, inputting it into a conditional encoder module to obtain a conditional guidance vector; S52, the noise module generates random initial noise and inputs it into the UNet network module, and gradually downsamples and upsamples to remove the noise; S53, UNet network module first passes through the convolution block to reduce the number of channels in the process of step-by-step downsampling, and then passes through the cross-attention block and the conditional guidance vector to perform cross-attention; in the process of step-by-step upsampling, the number of channels is first increased through the convolution block, and then passes through the cross-attention block and the conditional guidance vector to perform cross-attention; finally, the predicted noise is obtained; S54, subtract the predicted noise from the random initial noise as the secondary noise, and input it into the UNet network module again to remove the noise, so as to obtain a supplementary feature map; S55. Concatenate the supplementary feature map with the reconstructed feature map to obtain video supplementary information, and input the video supplementary information into a core decoder to obtain a reconstructed video frame.
Citation Information
Cited By
Image data transmission method, image data processing method and device
CN120512548A
Photovoltaic image dynamic compression and reconstruction method and system based on auto-encoder
CN120825594A