Video coding and decoding method and system for earth observation video of low earth orbit satellite, and medium
By constructing a global model of a hybrid architecture of variational autoencoders and deep neural networks, combined with satellite motion vectors and long-term background reference frame sequences, the low-orbit satellite video encoding and decoding is optimized, solving the problem of insufficient reconstructed video quality caused by inefficient encoding, and achieving high-definition real-time analysis.
Patent Information
- Application Number
- CN202510988688.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-17
AI Technical Summary
When the data acquisition capacity of a single satellite exceeds 500Gbps in low-orbit satellite earth observation, the existing video coding technology is inefficient, resulting in the reconstructed video quality hovering at 720p. This makes it difficult to meet the core requirements of high-definition real-time analysis in scenarios such as disaster prevention and mitigation and safe production.
A global model with a hybrid architecture of variational autoencoders and deep neural networks is constructed, trained at the ground station and incrementally updated at the satellite end. By combining satellite motion vectors and long-term background reference frame sequences, the encoding and decoding process is optimized through a lightweight end-to-end video encoding and decoding model.
It significantly improves the stability and image fidelity of reconstructed videos in weak network environments, meets the requirements of high compression rate and low computational complexity, and realizes the application scenario of high-definition real-time analysis.
Smart Images

Figure CN120812282A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of satellite remote sensing data compression and processing, in particular to a video coding method and system for low-orbit satellite earth observation video and a medium. BACKGROUND
[0002] In recent years, satellite video dynamic monitoring is widely used in the fields of emergency command, disaster prevention and reduction, and land surveying, but the data volume of ultra-high resolution is far beyond the upper limit of on-board storage and processing capacity. Traditional video coding technology mainly realizes compression by eliminating short-term spatio-temporal redundancy within a single video stream. However, when a satellite repeatedly observes the same area, the topography background shows high similarity in a long time span. The traditional encoder lacks the ability to make global reference across videos, and cannot efficiently utilize such redundancy, resulting in low compression efficiency. With the breakthrough progress of low-orbit satellite Internet and high-throughput technology, the large-scale application of satellite dynamic video monitoring in key fields such as emergency command and land surveying is also facing severe constraints in underlying coding efficiency; the inherent ultra-high code rate of ultra-high definition video forms a sharp contradiction with the limited on-board resources: on the one hand, the topography background shows high stability in long-term repeated observation, but the traditional coding technology lacks the ability to make global reference across time periods, and cannot efficiently reuse such long-range background redundancy; on the other hand, the existing coding schemes lack adaptability to complex environments such as dramatic light changes and seasonal changes, and are prone to quality degradation in reconstruction in application in key coverage areas of national new infrastructure such as ocean-going ships and remote mountainous areas.
[0003] Although the existing technology attempts to improve efficiency through multi-modal background reference frame construction and semantic communication compression, it has significant limitations in actual deployment: long-term background reference frames are easily disturbed by dramatic light changes or seasonal changes, causing image blurring and color deviation, and reducing the effectiveness of compression; semantic coding based on deep learning can improve compression rate, but it is difficult to process high-resolution video streams in real time due to limited on-board computing power. More importantly, the national "Regulations on Satellite Services for Terminal Equipment Direct Connection" clearly requires "giving equal importance to development and security", and the existing technology lacks robustness in complex observation environments, and complex interaction processes are required for cross-ephemeris data collaboration, increasing the security risk of the system. These defects directly restrict the policy implementation of space-ground integration network - when the data acquisition capacity of a single satellite breaks through 500 Gbps, inefficient coding makes the reconstructed video quality hover at the 720p level, making it difficult to meet the core needs of "high-definition real-time analysis" in disaster prevention and reduction, safety production and other scenarios. Therefore, there is an urgent need for a new satellite video coding mechanism that takes into account high compression rate, environmental robustness and low computational complexity, to truly release the technical potential of satellite data processing industry and support the strategic advancement of national new infrastructure to full coverage. SUMMARY
[0004] The technical problem to be solved by the present application is that the defects of the existing video coding technology make the reconstructed video quality hover at the 720p level when the single-satellite data acquisition capacity breaks through 500Gbps, which is difficult to meet the core needs of high-definition real-time analysis in disaster prevention and reduction, safety production and other scenarios; The present application aims to provide a video coding method, system and medium for low-orbit satellite earth observation video, which improves the method and coding model architecture based on existing technology. The present application can effectively solve the inherent defects of traditional coding technology that cannot reuse long-range spatiotemporal redundancy, significantly improve the stability and fidelity of reconstructed video in weak network environment, and meet the application scenarios where inefficient coding makes the reconstructed video quality hover at the 720p level when the single-satellite data acquisition capacity breaks through 500Gbps.
[0005] The present application is realized by the following technical solutions: The present application provides a video coding method for low-orbit satellite earth observation video, which includes: Constructing a global model of variational autoencoder and deep neural network hybrid architecture, training the global model at the ground station, and incrementally updating the global model at the satellite end; Predicting future background frames based on the global model, constructing a long-term background reference frame sequence, and extracting long-term background features from the long-term background reference frame sequence; Obtaining satellite motion features and extracting satellite motion vectors in combination with long-term background reference frame sequences; Twist transformation operation is performed on the satellite motion vector and long-term background features, and then refined grid operation is performed to obtain high-precision context information; The high-precision context information is input into an end-to-end video coding model based on deep neural network for coding and decoding.
[0006] Further optimization scheme is that the global model of variational autoencoder and deep neural network hybrid architecture is constructed, the global model is trained at the ground station, and the global model is incrementally updated at the satellite end; The method includes: The global model of variational autoencoder and deep neural network hybrid architecture is constructed, the global model uses variational autoencoder as the encoder to convert the input variable into hidden variable, and uses deep neural network as the decoder for decoding prediction; The total loss function of the global model includes reconstruction error and KL divergence term; The local change area of the global model is finely updated at the satellite end with a period T, and the fine updating range includes a buffer zone formed by the outward expansion of the local change area by a proportion n; the boundary artifact is eliminated by combining the edge smoothing processing technology in the fine updating process; the satellite end also analyzes the historical local change area in real time to predict the future change area of the global model.
[0007] A further optimization scheme is that the future background frame is predicted based on the global model to construct a long-term background reference frame sequence; the method includes: inputting satellite orbit parameters and multi-temporal ground observation data into the global model to construct a basic long-term background reference frame sequence; The global model first performs preliminary prediction on the current frame based on a traditional intra-frame prediction mode, and then searches for a long-term background reference frame in the basic long-term background reference frame sequence as supplementary information based on the satellite orbit parameters, and adaptively fuses the preliminary prediction result to obtain a prediction value:
[0008] Among them, The preliminary prediction result is represented by p(x, y, t); The compensation part is represented by p(x, y, t); The dynamic fusion coefficient is represented by a background complexity and a size of a coding tree unit; t represents time; x represents a horizontal direction coordinate of a coding unit in the current frame; and y represents a vertical direction coordinate of the coding unit in the current frame; The long-term background reference frame sequence is constructed based on all adaptive fusion prediction values.
[0009] A further optimization scheme is that the size of the coding tree unit is dynamically divided according to the satellite pitch angle change rate, and the specific method includes: A first threshold A and a second threshold B are set, and the first threshold A is greater than the second threshold B; When the satellite pitch angle change rate is greater than the first threshold A, the coding tree unit is divided into a first size; When the satellite pitch angle change rate is greater than the second threshold B and less than or equal to the first threshold A, the coding tree unit is divided into a second size; When the satellite pitch angle change rate is less than or equal to the second threshold B, the coding tree unit is divided into a third size; the first size is less than the second size, and the second size is less than the third size.
[0010] A further optimization scheme is that the satellite motion feature is obtained, and a satellite motion vector is extracted in combination with the long-term background reference frame sequence; the method includes: According to the orbit parameters and the motion model of the satellite, the displacement parameters of the satellite between two time points are calculated; The displacement parameters are converted into image coordinates, and the long-term background reference frame is translated according to the image coordinates to obtain a predicted frame of the current frame; The motion vector of a local window around each pixel or block is estimated by minimizing the difference in the local window between the current frame and the predicted frame.
[0011] A further optimization scheme is that the satellite motion vector and the long-term background feature are first subjected to a warping transformation operation, and then subjected to a refined grid operation to obtain high-precision context information; the method comprises: The long-term background feature is warped to the current frame perspective according to the satellite motion vector; The context information is obtained by eliminating the warping artifacts and fusing dynamic information based on a residual attention mechanism, and the background structure, texture details and motion information are carried based on a channel attention mechanism in the process.
[0012] A further optimization scheme is that the end-to-end video coding model based on the deep neural network is an end-to-end encoder-decoder of a double-branch progressive reconstruction architecture, the encoding end takes the context information of the feature domain as the conditional information, the decoding end first performs primary reconstruction, and then performs background fusion reconstruction; the deep neural network based on the adaptive spatio-temporal hybrid entropy model is used as the entropy model.
[0013] A further optimization scheme is that the entropy model comprises:
[0014]
[0015] wherein, refers to a spatio-temporal prior encoder, and the input is , and refers to a prior fusion network; represents an adaptive weight coefficient, and the range is [0, 1], which is used for dynamically fusing multi-source prior information; represents a super-prior decoding feature; represents an autoregressive prior feature; represents a spatio-temporal prior feature; represents a weight matrix, which is used for linear transformation; represents a bias term, which is used for adjusting the offset of the output distribution; represents a Sigmoid function, which maps the output to the interval [0, 1].
[0016] The scheme also provides a video coding system for low-orbit satellite earth observation video, which is used for implementing the video coding method for low-orbit satellite earth observation video; the system comprises: The model training updating module is configured to construct a global model of a variational autoencoder and a deep neural network hybrid architecture, train the global model at a ground station, and incrementally update the global model at a satellite end; The first extraction module is configured to predict a future background frame based on the global model, construct a long-term background reference frame sequence, and extract long-term background features from the long-term background reference frame sequence; The second extraction module is configured to obtain satellite motion features and extract satellite motion vectors in combination with the long-term background reference frame sequence; The feature operation module is configured to perform a warping transformation operation on the satellite motion vectors and the long-term background features, and then perform a refined grid operation to obtain high-precision context information. The output module is configured to input the high-precision context information into an end-to-end video coding and decoding model based on a deep neural network for coding and decoding.
[0017] The present application also provides a computer readable medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the low-orbit satellite earth observation video coding and decoding method according to the above-mentioned solution.
[0018] Compared with the prior art, the present application has the following advantages and beneficial effects: 1. The low-orbit satellite earth observation video coding and decoding method, system and medium provided by the present application improve the method and the coding and decoding model architecture on the basis of the prior art. By constructing a long-term background reference frame sequence on the ground, using satellite motion vectors, and using a lightweight end-to-end video coding and decoding model based on a deep neural network, the inherent defects of the traditional coding technology that cannot reuse long-range spatiotemporal redundancy can be effectively solved, and the stability and picture fidelity of the reconstructed video in a weak network environment can be significantly improved. The application scenario in which the single-satellite data acquisition capacity breaks through 500 Gbps, but the inefficient coding makes the reconstructed video quality hover at the 720p level can be met.
[0019] 2. The low-orbit satellite earth observation video coding and decoding method, system and medium provided by the present application. The end-to-end video coding and decoding model based on a deep neural network uses a self-adaptive spatiotemporal hybrid entropy model as the entropy model, which can adaptively distinguish between static backgrounds and motion regions in the video. For the static region, the spatiotemporal prior provides strong constraints to reduce the code rate. For the region with intense motion, the distribution constraints are relaxed to retain details.
[0020] 3. The low-orbit satellite earth observation video coding and decoding method, system and medium provided by the present application. The global model for constructing the long-term background reference frame sequence on the ground has a coding tree unit size that is dynamically divided according to the satellite pitch angle change rate. In different dynamic scenarios, different refinement processing strategies are used to retain detail information while improving coding efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present application, the drawings required to be used in the examples will be briefly introduced as follows, and it should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings. In the drawings: Figure 1 Flowchart of a video coding method for low-orbit satellite earth observation video; Figure 2 Schematic diagram of video coding principle for low-orbit satellite earth observation video; Figure 3 Schematic diagram of video coding system architecture for low-orbit satellite earth observation video; Figure 4 Schematic diagram of video coding system structure for low-orbit satellite earth observation video. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solutions and advantages of the present application more clear and obvious, the present application will be further described in detail below in combination with examples and drawings, and the exemplary embodiments of the present application and their descriptions are only used to explain the present application, and should not be regarded as a limitation on the present application.
[0023] The defects of the existing video coding technology are that when the single-satellite data acquisition capacity breaks through 500Gbps, the inefficient coding makes the reconstructed video quality hover at the level of 720p, which is difficult to meet the core needs of high-definition real-time analysis in disaster prevention and reduction, safety production and other scenarios; In view of this, the present application provides the following examples to solve the above technical problems.
[0024] Example 1 The present embodiment provides a video coding method for low-orbit satellite earth observation video, as shown in Figure 1 and Figure 2 , comprising: Step 1: build a global model of a variational autoencoder and a deep neural network hybrid architecture, train the global model on the ground station, and incrementally update the global model on the satellite end; This step specifically includes the method: S11, build a global model of a variational autoencoder and a deep neural network hybrid architecture, the global model uses a variational autoencoder as an encoder to convert input variables into latent variables, and uses a deep neural network as a decoder for decoding prediction; The total loss function of the global model contains a reconstruction error and a KL divergence term; Specific operation is to collect a large number of star observation data, and then the star observation data is cleaned, pretreated and feature extracted, then based on the star observation data and its features, the global model is trained by using large-scale unsupervised learning algorithm, and the coding ability and generalization ability of the global model are improved in this way. Among them, the variational autoencoder VAE as the encoder, its role is to convert the input image sequence into hidden variables, and the deep neural network decoder decodes to predict the future background frame, and the decoder adopts the multi-head self-attention mechanism in the prediction process.
[0025] The expression of variational autoencoder VAE is: ; Among them, and represent the mean network and the covariance network; represents the approximate posterior distribution defined by the encoder; z represents the latent variable; x represents the original data to be encoded or reconstructed.
[0026] The deep neural network decoder adopts the multi-head self-attention mechanism :
[0027] Among them, Q represents the focus position to be paid attention to at present, which is used to calculate the similarity with other positions; K represents the identification of all positions in the sequence, which is used to match the query vector; V represents the actual information of each position, and the weighted sum forms the output; T represents transposition; d k represents the dimension of Key; S12, the satellite end updates the local change area of the global model with a period T, and the fine update range includes the buffer zone formed by the outward expansion of the local change area by n; In the fine update process, the edge smoothing processing technology is combined to eliminate the boundary artifacts; The satellite end also analyzes the historical local change area in real time to predict the future change area of the global model.
[0028] In this scheme, the satellite end adopts the incremental update strategy to maintain a lightweight global model, and every 5 minutes, the satellite will fine update the local change area; In order to capture changes more comprehensively, the fine update range of this embodiment is set to the buffer zone formed by the outward expansion of the change area by 10%; When the ground structure change of the newly collected video exceeds 10% pixel difference, the dynamic optimization of the feature library is automatically triggered to ensure the timeliness of the global model; In the adjustment process, in order to avoid edge artifacts, edge smoothing processing technology will be used. In addition, the satellite will analyze the historical background sequence in real time, so as to predict the future change area in advance, so as to update the global model in time.
[0029] Specifically, the fine update result is: ; where, represents the expansion of the feature effective area in the input through the convolution kernel; represents the original feature map or binary mask; represents the convolution kernel; represents the updated feature map and region mask; The embodiment is based on a Gaussian filter to eliminate boundary artifacts:
[0030] where, represents the standard deviation of the Gaussian kernel, used to control the smoothing strength; = 1.0; represents the background frame after Gaussian smoothing; represents the Gaussian weight kernel function; i, j represents the horizontal and vertical offset relative to the current pixel; represents the pixel value of the original background frame at coordinate (i, j); Total loss function of global model contains reconstruction error and KL divergence term:
[0031] where, KL() divergence term represents the constraint that the implicit variable distribution is close to the standard normal distribution; represents the balance coefficient; = 0.01; T represents the total number of time steps; represents the reconstructed background of the t-th frame; represents the true background of the t-th frame; represents the regularization coefficient; q() represents the approximate posterior distribution defined by the encoder; p() represents the prior distribution.
[0032] Step two: predict future background frames based on the global model, construct a long-term background reference frame sequence, and extract long-term background features from the long-term background reference frame sequence; In step two, the global model is used to predict future background frames and construct a long-term background reference frame sequence. The method includes: inputting satellite orbit parameters and multi-temporal ground observation data into the global model to construct a basic long-term background reference frame sequence. The global model first performs preliminary prediction on the current frame based on the traditional intra-frame prediction mode, then searches for a long-term background reference frame in the basic long-term background reference frame sequence as supplementary information based on the satellite orbit parameters, and adaptively fuses the preliminary prediction result to obtain a prediction value:
[0033] where, represents the initial prediction result; represents the compensation part, which is the search value of the long-term background reference frame; represents the dynamic fusion coefficient, which is determined by the background complexity (the variance of the long-term background reference frame) and the size of the coding tree unit; t represents time; x, y represent the position of the prediction frame; The long-term background reference frame sequence is constructed based on the prediction values obtained by all adaptive fusion. The above initial prediction result is: ; wherein M represents a candidate prediction mode set, represents the mode weight; represents the intra-frame prediction value at time t and position (x, y); m represents the time offset.
[0034] Dynamic fusion coefficient is:
[0035] wherein, represents the background complexity constant, which is defined as , i.e. the maximum variance value of the reference frame; represents the size of the coding tree unit; k represents the slope coefficient.
[0036] Size of the coding tree unit According to the dynamic division of the satellite elevation angle change rate, the specific method comprises: Setting a first threshold A and a second threshold B, the first threshold A is greater than the second threshold B; When the satellite elevation angle change rate is greater than the first threshold A, the coding tree unit is divided into a first size; When the satellite elevation angle change rate is greater than the second threshold B and less than or equal to the first threshold A, the coding tree unit is divided into a second size; When the satellite elevation angle change rate is less than or equal to the second threshold B, the coding tree unit is divided into a third size; the first size is less than the second size, and the second size is less than the third size.
[0037] In the intra-frame prediction link, the traditional DVC coding mode is improved; first, dynamic coding unit CU division is carried out based on satellite orbit parameters, and the size of the coding tree unit is determined by the satellite elevation angle change rate
[0038] wherein, is the satellite elevation angle change rate; A=1° / s; B=0.5° / s; In high dynamic scenes (such as high-speed moving vehicles, etc.), the satellite elevation angle change rate is large, and the size of the coding tree unit is large; In the low dynamic scene (B), large coding tree units (CTU) are suitable for satellite attitude stabilization, improving the coding efficiency.
[0039] In step two, long-term background features are extracted from the long-term background reference frame sequence, including methods: extracting low-redundancy background feature representation from long-term background reference frames provides key prior knowledge for subsequent long-range reference frame generation; using a lightweight convolutional neural network (CNN) to extract multi-scale deep features: ; wherein, is a lightweight encoding network of parameters , and is a deep feature vector of the background frame .
[0040] Step three: obtain satellite motion features and combine long-term background reference frame sequence to extract satellite motion vectors; this step specifically includes methods: According to the satellite's orbital parameters and motion model, the displacement parameters of the satellite between two time points are calculated; Convert the displacement parameters to image coordinates, and translate the long-term background reference frame according to the image coordinates to obtain the predicted frame of the current frame; Minimize the difference between the current frame and the predicted frame within each pixel or block local window, and estimate the motion vector of each pixel or block local window.
[0041] In satellite video, the motion of the satellite is usually predictable because its orbital path and speed are known, and this characteristic can be used to estimate the overall motion of the background; first, according to the satellite's orbital parameters and motion model, the displacement between two time points is calculated. Then, this displacement is applied to the long-term background reference frame to predict the background position of the current frame. By comparing the actual current frame with the predicted frame, the motion vector is extracted. These motion vectors mainly reflect the dynamic changes in the scene, rather than the predictable motion of the background; by analyzing the differences within each pixel or block local window, and combining the known motion information of the satellite to optimize the motion estimation. Specifically: assuming the position of the satellite at time is ; the displacement between time and is :
[0042] Convert the displacement between time and to image coordinates ; predicted frame is the current frame By translating the background reference frame :
[0043] Then, by minimizing the difference between the current frame and the predicted frame , the motion vector of each pixel or block is estimated :
[0044] where, is the local window, is the displacement searched within a limited range near the predicted motion; represents the brightness value of the current frame at pixel position .
[0045] Step four: first warp the satellite motion vector and long-term background features, and then perform a refined grid operation to obtain high-precision context information; this step specifically includes the method: Warp the long-term background features to the current frame perspective according to the satellite motion vector; there
[0046] where, represents the long-term background features, represents the satellite motion vector; adopt differentiable bilinear sampling to support gradient rotation; represents the motion-compensated background frame estimate; Warp() represents the warp operation.
[0047] Based on the residual attention mechanism, eliminate warp artifacts and fuse dynamic information to obtain context information; during the process, based on the channel attention mechanism, carry background structure, texture details and motion information; the obtained context information is:
[0048] where, represents the residual attention mechanism, represents the feature splicing and fusion operation; represents the fully connected network, which extracts appearance features; At the same time, increase the channel attention mechanism:
[0049] Output high-dimensional context features , carrying background structure, texture details and motion information, used to guide the end-to-end video encoding and decoding of deep neural networks. Among them, is the activation function; It is a fully connected network used to generate attention weights; represents global average pooling; Step 5: Input the high-precision context information into the end-to-end video encoding and decoding model based on deep neural network for encoding and decoding.
[0050] The end-to-end video codec model based on deep neural networks is an end-to-end codec with a dual-branch progressive reconstruction architecture. The encoding end uses the context information of the feature domain as conditional information, and the decoding end first performs primary reconstruction and then background fusion reconstruction. The deep neural network-based end-to-end video codec model uses a hybrid entropy model of a spatiotemporal prior path and a super prior path as the entropy model.
[0051] For traditional video coding, only simple pixel subtraction operations are used, which cannot fully utilize the complex correlation between frames. Traditional coding schemes use pixel subtraction to remove inter-frame redundancy, and their entropy lower bound is higher than that of conditional coding:
[0052] Among them, represents the Shannon entropy, Indicates the current frame, represents the predicted frame; In view of this, the encoding framework of this solution defines the feature domain context It is conditional information and carries rich spatiotemporal information through high-dimensional features.
[0053]
[0054] in, ,context() is the context extraction module; Represents contextual information; Indicates reconstructed frame; dec() indicates decoding; enc() indicates encoding; The channel dimension of is much higher than the three-dimensional dimension of RGB, which can well separate high-frequency texture information and color information, automatically switch to intra-frame coding mode for new content on the motion boundary, and reduce the residual amplitude.
[0055] At the same time, in order to lightweight the end-to-end codec, this scheme adopts a dual-branch progressive reconstruction architecture and achieves efficient compression through the feature domain background fusion mechanism. The encoder uses four-level downsampling to splice the input current frame and context features and compress them into a low-dimensional potential representation. Its mathematical expression is as follows:
[0056] in, It consists of 4 convolutional layers, each followed by a GDN normalization and residual block; Indicates the current frame input; represents the background feature.
[0057] The decoding end adopts progressive reconstruction, first performing primary reconstruction:
[0058] wherein D represents the primary reconstructed feature map; represents 4-level sub-pixel convolution up-sampling, each level followed by inverse GDN and residual block, output size: .
[0059] Then, background fusion reconstruction is performed:
[0060] wherein, represents down-sampling to the same resolution of the background feature, output is the reconstructed frame; represents a second-stage decoder network.
[0061] The entropy model of the existing deep video encoder (such as DVC) only considers spatial correlation and static prior, the spatial correlation refers to point-by-point prediction through an autoregressive model, and the static prior refers to modeling global statistical properties using hyper-prior; this leads to two problems: high computational complexity and ignoring the time redundancy specific to videos, and the satellite long-term background reference frame cannot be combined; therefore, the adaptive spatio-temporal hybrid entropy model is used as the entropy model in the present scheme:
[0062]
[0063] wherein, refers to a spatio-temporal prior encoder, the input is , and refers to a prior fusion network; represents an adaptive weight coefficient, ranging from 0 to 1, used for dynamically fusing multi-source prior information; represents a hyper-prior decoded feature; represents an autoregressive prior feature; represents a spatio-temporal prior feature; represents a weight matrix, used for linear transformation; represents a bias term, used for adjusting the offset of the output distribution; represents a Sigmoid function, mapping the output to the interval [0, 1].
[0064] This design enables the model to adaptively distinguish between static background and motion regions in the video - for static regions, the spatio-temporal prior provides strong constraints to reduce the bit rate; for regions with intense motion, the distribution constraints are relaxed to preserve details. To adapt to the limited computing resources on the satellite side, this model provides a configurable lightweight mode that significantly outperforms the traditional residual coding framework under the same complexity by removing the autoregressive module of serial computation and only retaining the parallelized hyper-prior and spatio-temporal prior paths.
[0065] Embodiment 2 The present embodiment provides a video coding system for low-orbit satellite earth observation video, as shown in Figure 3 and Figure 4 , for implementing the video coding method for low-orbit satellite earth observation video described in Embodiment 1; the system comprises: a model training and updating module for constructing a global model of a variational autoencoder and deep neural network hybrid architecture, training the global model at a ground station, and incrementally updating the global model at the satellite side; a first extraction module for predicting future background frames based on the global model, constructing a long-term background reference frame sequence, and extracting long-term background features from the long-term background reference frame sequence; a second extraction module for obtaining satellite motion features and extracting satellite motion vectors in combination with the long-term background reference frame sequence; a feature operation module for performing warping transformation operation on the satellite motion vectors and long-term background features, and then performing refinement grid operation to obtain high-precision context information; an output module for inputting the high-precision context information into an end-to-end video coding model based on deep neural network for coding and decoding.
[0066] Embodiment 3 The present embodiment provides a computer readable medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the video coding method for low-orbit satellite earth observation video as described in Embodiment 1; the following steps are specifically executed: Step one: construct a global model of a variational autoencoder and deep neural network hybrid architecture, train the global model at a ground station, and incrementally update the global model at the satellite side; Step two: predict future background frames based on the global model, construct a long-term background reference frame sequence, and extract long-term background features from the long-term background reference frame sequence; Step three: obtain satellite motion features and extract satellite motion vectors in combination with the long-term background reference frame sequence; Step four: perform warping transformation operation on the satellite motion vectors and long-term background features, and then perform refinement grid operation to obtain high-precision context information; Step five: input high-precision context information into a deep neural network-based end-to-end video coding model for coding and decoding.
[0067] The above detailed description of the specific implementation of the application, the purpose, technical solutions and beneficial effects have been further detailed, it should be understood that the above-mentioned only for the specific implementation of the present application, and not used to limit the scope of protection of the present application, any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application, should be included within the scope of protection of the present application.
Claims
1. A video encoding and decoding method for low-orbit satellite earth observation videos, characterized in that: include: Constructing a global model with a hybrid architecture of a variational autoencoder and a deep neural network, training the global model at the ground station, and incrementally updating the global model at the satellite end; Predict future background frames based on the global model, construct a long-term background reference frame sequence, and extract long-term background features from the long-term background reference frame sequence; Obtain satellite motion features and extract satellite motion vectors in combination with long-term background reference frame sequences; The satellite motion vector and long-term background features are first distorted and then meshed to obtain high-precision context information. High-precision context information is input into an end-to-end video encoding and decoding model based on deep neural networks for encoding and decoding.
2. The video encoding and decoding method for low-orbit satellite earth observation video according to claim 1, characterized in that: The global model of the hybrid architecture of the variational autoencoder and the deep neural network is constructed, the global model is trained at the ground station, and the global model is incrementally updated at the satellite end; Includes methods: Constructing a global model with a hybrid architecture of a variational autoencoder and a deep neural network. The global model uses a variational autoencoder as an encoder to convert input variables into latent variables, and a deep neural network as a decoder for decoding prediction. The total loss function of the global model includes reconstruction error and KL divergence terms; The satellite end performs fine updating of the local change area of the global model with a period T, and the range of the fine update includes a buffer zone formed by an outward expansion ratio n of the local change area; the edge smoothing technology is combined with the edge smoothing processing technology to eliminate boundary artifacts during the fine update process; the satellite end also analyzes the historical local change area in real time to predict the future change area of the global model.
3. The video encoding and decoding method for low-orbit satellite earth observation video according to claim 2, characterized in that: The method of predicting future background frames based on the global model to construct a long-term background reference frame sequence includes: obtaining satellite orbit parameters and multi-phase surface observation data and inputting them into the global model to construct a basic long-term background reference frame sequence; The global model first makes a preliminary prediction of the current frame based on the traditional intra-frame prediction mode, and then searches for a long-term background reference frame in the basic long-term background reference frame sequence as supplementary information in combination with the satellite orbit parameters. The preliminary prediction results are adaptively fused to obtain the predicted value: in, Indicates the preliminary prediction results; It represents the compensation part, which is the search value of the long-term background reference frame; represents the dynamic fusion coefficient, which is determined by the background complexity and the size of the coding tree unit; t represents time; x represents the horizontal coordinate of the coding unit in the current frame; y represents the vertical coordinate of the coding unit in the current frame; A long-term background reference frame sequence is constructed based on all adaptively fused prediction values.
4. The video encoding and decoding method for low-orbit satellite earth observation video according to claim 3, characterized in that: The size of the coding tree unit is dynamically divided according to the satellite pitch angle change rate, and the specific method includes: Setting a first threshold value A and a second threshold value B, wherein the first threshold value A is greater than the second threshold value B; When the satellite elevation angle change rate is greater than a first threshold A, dividing the coding tree unit into a first size; When the satellite pitch angle change rate is greater than a second threshold value B and less than or equal to a first threshold value A, dividing the coding tree unit into a second size; When the satellite pitch angle change rate is less than or equal to a second threshold value B, the coding tree unit is divided into a third size; the first size is smaller than the second size, and the second size is smaller than the third size.
5. The video encoding and decoding method for low-orbit satellite earth observation video according to claim 1, characterized in that: The method of obtaining satellite motion features and extracting satellite motion vectors in combination with a long-term background reference frame sequence includes: Calculate the satellite's displacement parameters between two time points based on the satellite's orbital parameters and motion model; The displacement parameters are converted into image coordinates, and the long-term background reference frame is translated according to the image coordinates to obtain the predicted frame of the current frame; Minimize the difference in the local window around each pixel or block between the current frame and the predicted frame, and estimate the motion vector of the local window around each pixel or block.
6. The video encoding and decoding method for low-orbit satellite earth observation video according to claim 1, characterized in that: The satellite motion vector and the long-term background features are first subjected to a distortion transformation operation, and then subjected to a grid refinement operation to obtain high-precision context information; Includes methods: Warp the long-term background features to the current frame perspective based on the satellite motion vector; The residual attention mechanism is used to eliminate distortion artifacts and fuse dynamic information to obtain contextual information. In the process, the channel attention mechanism is used to carry background structure, texture details and motion information.
7. The video encoding and decoding method for low-orbit satellite earth observation video according to claim 1, characterized in that: The end-to-end video encoding and decoding model based on deep neural networks is an end-to-end codec with a dual-branch progressive reconstruction architecture. The encoding end uses the context information of the feature domain as conditional information, and the decoding end first performs primary reconstruction and then background fusion reconstruction; the deep neural network-based end uses an adaptive spatiotemporal hybrid entropy model as the entropy model.
8. The video encoding and decoding method for low-orbit satellite earth observation video according to claim 7, characterized in that: The entropy model includes: in, refers to the spatiotemporal prior encoder, the input is ,and refers to the prior fusion network; Represents the adaptive weight coefficient, ranging from [0,1], which is used to dynamically fuse multi-source prior information; represents the super-prior decoding feature; represents the autoregressive prior feature; Represents spatiotemporal prior features; Represents the weight matrix, used for linear transformation; Represents the bias term, which is used to adjust the offset of the output distribution; Represents the Sigmoid function, which maps the output to the [0,1] interval.
9. A video encoding and decoding system for low-orbit satellite earth observation videos, characterized in that: A method for encoding and decoding low-orbit satellite earth observation videos according to any one of claims 1 to 8; the system comprising: A model training and update module is used to build a global model with a hybrid architecture of variational autoencoders and deep neural networks, train the global model at the ground station, and perform incremental updates on the satellite end; The first extraction module predicts future background frames based on the global model, constructs a long-term background reference frame sequence, and extracts long-term background features from the long-term background reference frame sequence; The second extraction module extracts the satellite motion vector based on the acquired satellite motion features and combined with the long-term background reference frame sequence; The feature operation module is used to first perform a distortion transformation operation on the satellite motion vector and long-term background features, and then perform a grid refinement operation to obtain high-precision context information; The output module is used to input high-precision context information into the end-to-end video encoding and decoding model based on deep neural networks for encoding and decoding.
10. A computer-readable medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the video encoding and decoding method for low-orbit satellite earth observation video as described in any one of claims 1 to 8.