Satellite video eigen decomposition method based on non-training network
Through the satellite video intrinsic decomposition method based on non-training network, the pyramid structure and bilateral filtering technology are used to solve the problems of large computational complexity and limited decomposition accuracy of satellite video, and achieve efficient spatial and color separation and high-quality reflectivity and shadow component extraction.
Patent Information
- Application Number
- CN202510746349.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-12
AI Technical Summary
Existing satellite video intrinsic decomposition methods are computationally intensive and have limited decomposition accuracy. In addition, deep learning models rely on large amounts of training data and cannot effectively separate spatial and color information in complex environments.
A satellite video intrinsic decomposition method based on a non-trained network is adopted to achieve spatial and color separation at the physical level through a pyramid structure. The spatiotemporal characteristics of background information are globally modeled in combination with a deep network, and bilateral filtering and deconvolution techniques are used for feature extraction and fusion.
It achieves efficient decomposition, improves the signal-to-noise ratio, enhances the extraction accuracy of reflectivity and shadow components, overcomes the problems of local inconsistency and temporal fluctuation, and is suitable for high-quality decomposition in dynamic environments.
Smart Images

Figure CN120635772A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of satellite video processing, and relates to a non-training satellite video eigendecomposition network method, in particular to a non-training network-based satellite video eigendecomposition method. Background Art
[0002] Satellite video has important applications in monitoring the dynamics of the Earth's surface and is widely used in target detection, change analysis, disaster monitoring, and other fields. Due to its high temporal resolution, satellite video can capture surface changes in real time. However, its imaging characteristics are characterized by a stable static background and a changing dynamic target, which makes aliasing between the target and the background easy to occur. Furthermore, to maintain high temporal resolution, satellite video often sacrifices signal-to-noise ratio, resulting in weak spatial and color information of the target, increasing the difficulty of target detection and tracking. This is especially true in complex lighting and low signal-to-noise ratio environments, where target boundaries are blurred and noise interference in the image is exacerbated, further reducing the accuracy of analysis.
[0003] Existing methods primarily rely on iterative optimization and one-shot decomposition strategies. While iterative optimization methods can provide high-quality decomposition results, they are computationally expensive and struggle to meet real-time processing requirements. While one-shot decomposition methods are computationally faster, their accuracy is often limited under complex lighting conditions, preventing high-quality target extraction. Furthermore, deep learning methods are increasingly being applied to the decomposition and processing of satellite video, potentially improving decomposition accuracy and reducing the time required for subsequent iterations. Existing deep learning methods typically rely on large amounts of satellite video training data, and their performance may be unstable in the absence of sufficient data. Furthermore, most deep learning models are based on data-driven optimization and fail to achieve effective spatial and color separation at the physical modeling level, limiting their application in complex environments. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems of large computational complexity, limited decomposition accuracy of the first decomposition mode, and dependence of deep models on a large amount of training data in traditional eigendecomposition methods, and to propose a satellite video eigendecomposition method based on a non-training network.
[0005] A satellite video intrinsic decomposition method based on a non-training network has the following specific steps: Step 1: Get satellite video images, input the satellite video images into the first convolution layer, the first pooling layer, and the first bilateral filter structure layer in sequence. The first bilateral filter structure layer outputs the feature ; Output features of the first bilateral filtering structure layer Input the second convolution layer, the second pooling layer, and the second bilateral filter structure layer in sequence, and the second bilateral filter structure layer outputs the features ; The output features of the second bilateral filtering structure layer Input the third convolution layer, the third pooling layer, and the third bilateral filter structure layer in sequence, and the third bilateral filter structure layer outputs features ; The output features of the third bilateral filtering structure layer Input the fourth convolution layer, the fourth pooling layer, and the fourth bilateral filter structure layer in sequence, and the fourth bilateral filter structure layer outputs the feature ; Output features of the fourth bilateral filtering structure layer Input the fifth convolution layer, the fifth pooling layer, and the fifth bilateral filter structure layer in sequence, and the fifth bilateral filter structure layer outputs the features ; Step 2: Calculate the output features of the first bilateral filtering structure layer Laplace difference of ; Calculate the output features of the second bilateral filtering structure layer Laplace difference of ; Calculate the output features of the third bilateral filtering structure layer Laplace difference of ; Calculate the output features of the fourth bilateral filtering structure layer Laplace difference of ; Step 3: Output features of the fifth bilateral filtering structure layer Input the first deconvolution layer, the first deconvolution layer outputs features ; The first deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; The second deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; The third deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; The fourth deconvolution layer output features and Laplace difference Perform element-by-element summation, input the fifth deconvolution layer after summation, and the fifth deconvolution layer outputs the feature ,feature As the final reflectivity component video; Step 4: Perform the first upsampling, second upsampling, third upsampling, fourth upsampling and fifth upsampling on the satellite video image obtained in step 1 to obtain the feature ; The features Divide by the output feature of the fifth bilateral filtering structure layer , get the features ; The features Input the first deconvolution layer, the first deconvolution layer outputs features ; The first deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; The second deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; The third deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; The fourth deconvolution layer output features and Laplace difference Perform element-by-element summation, input the fifth deconvolution layer after summation, and the fifth deconvolution layer outputs the feature ,feature As the final shadow component video.
[0006] The beneficial effects of the present invention are: The purpose of this research is to address the low signal-to-noise ratio (SNR) caused by the high temporal resolution of satellite video, which results in weak spatial and color information of targets, hindering detection and tracking. Existing intrinsic decomposition methods are computationally intensive, rely on large amounts of training data, and fail to achieve spatial-color separation at the physical level. This invention proposes a satellite video intrinsic decomposition method based on a non-training network. This method does not require extensive computation or training data, but achieves spatial and color separation at the physical level through a pyramid structure. Furthermore, it constructs a non-training network to achieve efficient decomposition and improve the SNR.
[0007] This method employs a novel initialization-decomposition strategy, combined with a deep network, to globally model the spatiotemporal characteristics of background information in satellite video. This architecture fully captures subtle variations in dynamic scenes while maintaining high temporal consistency of reflectivity components. This method improves temporal stability during reflectivity extraction, successfully overcoming the local inconsistencies and temporal fluctuations encountered in other methods, demonstrating its broad applicability and robustness in dynamic environments.
[0008] To validate the performance of the proposed algorithm, experiments were conducted on six Jilin-1 satellite video feeds. The experimental results confirmed that the proposed method, based on an untrained network, demonstrates excellent performance in extracting reflectance components from dynamic scenes. Its innovative model architecture enables comprehensive capture of dynamic illumination changes, effectively eliminating patch effects and temporal inconsistencies. Its precise boundary restoration significantly improves scene clarity and dynamic expression. Results from two metrics, mean squared error (MSE) and local mean squared error (LMSE), demonstrate the effectiveness of the proposed method, based on an untrained network, for eigendecomposition of satellite video.
[0009] The present invention proposes a satellite video intrinsic decomposition method based on a non-training network. First, the reflectance component is extracted in dynamic scenes through the PSVID model, making full use of the dynamic changes of illumination and background stability. On this basis, a multi-dimensional collaborative optimization mechanism is used to ensure the consistency of reflectance results between different frames, ultimately achieving the purpose of decomposition and extraction of high-quality reflectance and shadow components, and improving the decomposition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 It is a schematic diagram of the implementation process of the present invention; Figure 2 It is a network structure diagram of the present invention; Figure 3 is a frame from the original satellite video of the experiment; Figure 4 is a frame in the experimental reflectance component video; Figure 5 is a frame from the experimental shadow component video; Figure 6 Comparison of visualization results between the proposed method and other satellite video intrinsic decomposition methods. From left to right and from top to bottom, the order is: true value, SVID, MTE-ISVD, USVIDNet, and visualization results of the proposed method. Table 1 compares the MSE and LMSE values of the proposed method and other satellite video intrinsic decomposition methods. From left to right and from top to bottom, the order is: SVID, MTE-ISVD, USVIDNet, and the proposed method. DETAILED DESCRIPTION
[0011] Specific implementation method 1: This implementation method is a satellite video intrinsic decomposition method based on a non-training network. The specific process is as follows: Step 1: Get satellite video images, input the satellite video images into the first convolution layer, the first pooling layer, and the first bilateral filter structure layer in sequence. The first bilateral filter structure layer outputs the feature ; Output features of the first bilateral filtering structure layer Input the second convolution layer, the second pooling layer, and the second bilateral filter structure layer in sequence, and the second bilateral filter structure layer outputs the features ; The output features of the second bilateral filtering structure layer Input the third convolution layer, the third pooling layer, and the third bilateral filter structure layer in sequence, and the third bilateral filter structure layer outputs features ; The output features of the third bilateral filtering structure layer Input the fourth convolution layer, the fourth pooling layer, and the fourth bilateral filter structure layer in sequence, and the fourth bilateral filter structure layer outputs the feature ; Output features of the fourth bilateral filtering structure layer Input the fifth convolution layer, the fifth pooling layer, and the fifth bilateral filter structure layer in sequence, and the fifth bilateral filter structure layer outputs the features ;(Extract and enhance boundary features, filter details); Step 2: Calculate the output features of the first bilateral filtering structure layer Laplace difference of ; Calculate the output features of the second bilateral filtering structure layer Laplace difference of ; Calculate the output features of the third bilateral filtering structure layer Laplace difference of ; Calculate the output features of the fourth bilateral filtering structure layer Laplace difference of ; Step 3: Output features of the fifth bilateral filtering structure layer Input the first deconvolution layer, the first deconvolution layer outputs features ; The first deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; The second deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; The third deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; The fourth deconvolution layer output features and Laplace difference Perform element-by-element summation, input the fifth deconvolution layer after summation, and the fifth deconvolution layer outputs the feature ,feature As the final reflectivity component video; Step 4: Perform the first upsampling, second upsampling, third upsampling, fourth upsampling and fifth upsampling on the satellite video image obtained in step 1 to obtain the feature ; The features Divide by the output feature of the fifth bilateral filtering structure layer , get the features ; The features Input the first deconvolution layer, the first deconvolution layer outputs features ; The first deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; The second deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; The third deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; The fourth deconvolution layer output features and Laplace difference Perform element-by-element summation, input the fifth deconvolution layer after summation, and the fifth deconvolution layer outputs the feature ,feature As the final shadow component video.
[0012] Specific embodiment 2: The difference between this embodiment and specific embodiment 1 is that: in step 1, a satellite video image is obtained, and the satellite video image is sequentially input into the first convolution layer, the first pooling layer, and the first bilateral filter structure layer. The first bilateral filter structure layer outputs the feature ; Output features of the first bilateral filtering structure layer Input the second convolution layer, the second pooling layer, and the second bilateral filter structure layer in sequence, and the second bilateral filter structure layer outputs the features ; The output features of the second bilateral filtering structure layer Input the third convolution layer, the third pooling layer, and the third bilateral filter structure layer in sequence, and the third bilateral filter structure layer outputs features ; The output features of the third bilateral filtering structure layer Input the fourth convolution layer, the fourth pooling layer, and the fourth bilateral filter structure layer in sequence, and the fourth bilateral filter structure layer outputs the feature ; Output features of the fourth bilateral filtering structure layer Input the fifth convolution layer, the fifth pooling layer, and the fifth bilateral filter structure layer in sequence, and the fifth bilateral filter structure layer outputs the features ;(Extract and enhance boundary features, filter details); The specific process is: Get satellite video images, input the satellite video images into the first convolutional layer, and the first convolutional layer outputs features ; The first convolutional layer output features Input the first pooling layer, the first pooling layer outputs features ; The first pooling layer output features Input the first bilateral filtering structure layer, the first bilateral filtering structure layer outputs features ; The first bilateral filtering structure layer outputs features Input the second convolutional layer, the second convolutional layer outputs features ; The second convolutional layer output features Input the second pooling layer, the second pooling layer outputs features ; The second pooling layer output features Input the second bilateral filtering structure layer, the second bilateral filtering structure layer outputs features ; The second bilateral filtering structure layer outputs features Input the third convolution layer, the third convolution layer outputs features ; The third convolutional layer output features Input the third pooling layer, the third pooling layer outputs features ; The output features of the third pooling layer Input the third bilateral filtering structure layer, the third bilateral filtering structure layer outputs features ; The output features of the third bilateral filtering structure layer Input the fourth convolution layer, the fourth convolution layer outputs features ; The fourth convolutional layer output features Input the fourth pooling layer, the fourth pooling layer outputs features ; The fourth pooling layer output features Input the fourth bilateral filtering structure layer, the fourth bilateral filtering structure layer outputs features ; The fourth bilateral filtering structure layer outputs features Input the fifth convolutional layer, the fifth convolutional layer outputs features ; The fifth convolutional layer output features Input the fifth pooling layer, the fifth pooling layer outputs features ; The fifth pooling layer output features Input the fifth bilateral filtering structure layer, the fifth bilateral filtering structure layer outputs features .
[0013] Other steps and parameters are the same as those in the first embodiment.
[0014] Specific embodiment three: The difference between this embodiment and specific embodiment one or two is that: the satellite video image is input into the first convolution layer, and the first convolution layer outputs the feature ; The first bilateral filtering structure layer outputs features Input the second convolutional layer, the second convolutional layer outputs features ; The second bilateral filtering structure layer outputs features Input the third convolution layer, the third convolution layer outputs features ; The output features of the third bilateral filtering structure layer Input the fourth convolution layer, the fourth convolution layer outputs features ; The fourth bilateral filtering structure layer outputs features Input the fifth convolutional layer, the fifth convolutional layer outputs features ; The specific process is: Input satellite video images into Convolutional layer, Convolutional layer output features , ; The expression is: (1) in, Indicates the Convolutional layer output features; Indicates passing Input features after layer convolution (the first layer for satellite video images); is the position coordinate on the satellite video image, Number the timeframe; For the layer convolution kernel, is the convolution kernel input channel index, Output channel index of the convolution kernel; is the bias term of the convolution, which adjusts the convolution output; is the convolution operation; For the activation function, ReLU is used; Other steps and parameters are the same as those in the first or second embodiment.
[0015] Specific embodiment 4: This embodiment differs from any one of the specific embodiments 1 to 3 in that: the first convolutional layer outputs features Input the first pooling layer, the first pooling layer outputs features ; The second convolutional layer output features Input the second pooling layer, the second pooling layer outputs features ; The third convolutional layer output features Input the third pooling layer, the third pooling layer outputs features ; The fourth convolutional layer output features Input the fourth pooling layer, the fourth pooling layer outputs features ; The fifth convolutional layer output features Input the fifth pooling layer, the fifth pooling layer outputs features ; The specific process is: The resolution of the feature map is gradually reduced through the pooling operation. The pooling operation reduces the size of the feature map by downsampling. The expression is: (2) in, Indicates the pooling operation, using average pooling; Indicates the The pooling layer outputs features.
[0016] The other steps and parameters are the same as those in the first to third embodiments.
[0017] Specific embodiment 5: This embodiment differs from any one of specific embodiments 1 to 4 in that: the first pooling layer outputs features Input the first bilateral filtering structure layer, the first bilateral filtering structure layer outputs features ; The second pooling layer output features Input the second bilateral filtering structure layer, the second bilateral filtering structure layer outputs features ; The output features of the third pooling layer Input the third bilateral filtering structure layer, the third bilateral filtering structure layer outputs features ; The fourth pooling layer output features Input the fourth bilateral filtering structure layer, the fourth bilateral filtering structure layer outputs features ; The fifth pooling layer output features Input the fifth bilateral filtering structure layer, the fifth bilateral filtering structure layer outputs features ; The specific process is: Joint bilateral filtering is performed on the processed feature maps, combining spatial distance and pixel value similarity to further remove noise and enhance the separation of reflectance and shadow components. This process preserves edge information and details while removing background noise, allowing for more refined processing of reflectance and shadow components in complex scenes, ultimately resulting in clearer and more accurate component output.
[0018] The convolution and pooling feature maps are then output, and bilateral filtering is used to further optimize the detail recovery of reflectance and shadow components. This step effectively improves the accuracy of object boundaries and ensures the stability and consistency of reflectance and shadow components under varying lighting conditions and background dynamics.
[0019] Based on joint bilateral filtering to reduce noise and optimize the distinction between reflectance and shadow components, the expression is: (3) in, Represents pixels Neighborhood; are the standard deviations in the spatial domain, respectively; is the normalization factor to ensure the balance of the filtering results; It is The pooling layer outputs features; It is Pooling layer output features, position It's location The coordinates of the neighborhood points; Indicates the The output features of the bilateral filtering structure layer.
[0020] The output of each layer of operation is the input of the output of the previous layer.
[0021] (4) in, : The final output feature video; : The output feature map of the fourth layer is passed to the next layer as input; : Convolution kernel, weight matrix used to extract features; : The bias term of the convolution, which adjusts the convolution output.
[0022] Other steps and parameters are the same as those in Specific Embodiments 1 to 4-1.
[0023] Specific embodiment 6: This embodiment differs from any one of the specific embodiments 1 to 5 in that: in step 2 Calculate the output features of the first bilateral filtering structure layer Laplace difference of ; Calculate the output features of the second bilateral filtering structure layer Laplace difference of ; Calculate the output features of the third bilateral filtering structure layer Laplace difference of ; Calculate the output features of the fourth bilateral filtering structure layer Laplace difference of ; The specific process is: Layer-by-layer upsampling is used to gradually restore high-resolution feature maps, mitigating the loss of detail and information caused by low-resolution processing. Compared to traditional upsampling methods, this method uniquely combines the feature information of reflectance and shadow components, smoothing and enhancing them through Laplacian interpolation. By gradually upsampling and restoring high-resolution feature maps, the demand for high-precision reflectance and shadow component extraction can be more effectively met, ensuring accurate and consistent detail recovery.
[0024] To improve the accuracy and stability of detail recovery, Laplacian interpolation is used for enhancement. This effectively improves the boundary accuracy of reflectance and shadow components and reduces distortion during image processing. This method enhances the details of reflectance and shadow components during the recovery of high-resolution feature maps, thereby improving overall visual quality and object detection accuracy.
[0025] 1) Output features of the first bilateral filtering structure layer Upsample and get the upsampled features , gradually recover the high-resolution feature map; expressed as: (4) in, Represents an upsampling operation; Based on the output features of the first bilateral filtering structure layer And the upsampled features , calculate the Laplace difference , enhanced detail recovery, used to capture the difference between each layer feature map and its upsampling result, thereby improving the detail recovery ability; expressed as: (5) 2) Output features of the second bilateral filtering structure layer Upsample and get the upsampled features , gradually recover the high-resolution feature map; expressed as: (6) in, Represents an upsampling operation; Based on the output features of the second bilateral filtering structure layer And the upsampled features , calculate the Laplace difference , enhanced detail recovery, used to capture the difference between each layer feature map and its upsampling result, thereby improving the detail recovery ability; expressed as: (7) 3) Output features of the third bilateral filtering structure layer Upsample and get the upsampled features , gradually recover the high-resolution feature map; expressed as: (8) in, Represents an upsampling operation; Based on the output features of the third bilateral filtering structure layer And the upsampled features , calculate the Laplace difference , enhanced detail recovery, used to capture the difference between each layer feature map and its upsampling result, thereby improving the detail recovery ability; expressed as: (9) 4) Output features of the fourth bilateral filtering structure layer Upsample and get the upsampled features , gradually recover the high-resolution feature map; expressed as: (10) in, Represents an upsampling operation; Based on the output features of the fourth bilateral filtering structure layer And the upsampled features , calculate the Laplace difference , enhanced detail recovery, used to capture the difference between each layer feature map and its upsampling result, thereby improving the detail recovery ability; expressed as: (11) 5) Output features of the fourth bilateral filtering structure layer Upsample and get the upsampled features , gradually recover the high-resolution feature map; expressed as: (12) in, Represents an upsampling operation; Based on the output features of the fourth bilateral filtering structure layer and the upsampled features , calculate the Laplace difference , enhanced detail recovery, used to capture the difference between each layer feature map and its upsampling result, thereby improving the detail recovery ability; expressed as: (13).
[0026] Other steps and parameters are the same as those in Specific Implementations 1 to 5-1.
[0027] Specific embodiment seven: This embodiment differs from any one of the specific embodiments one to six in that: in step 3, the output feature of the fifth bilateral filtering structure layer is Input the first deconvolution layer, the first deconvolution layer outputs features ; The first deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; The second deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; The third deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; The fourth deconvolution layer output features and Laplace difference Perform element-by-element summation, input the fifth deconvolution layer after summation, and the fifth deconvolution layer outputs the feature ,feature As the final reflectivity component video; The specific process is: Deconvolution and pixel-by-pixel fusion are used to combine the reflectance and shadow components to obtain the final output, mitigating the information loss and detail blurring experienced in traditional decomposition methods. Compared to traditional single decomposition methods, this method uniquely fuses the reflectance and shadow components layer by layer and applies weighting coefficients to ensure the precise integration of features at different levels. This layer-by-layer optimization approach more effectively meets the requirements for high-quality reflectance and shadow component extraction, ensuring the accuracy and clarity of the final output. Improved the extraction accuracy of reflectance and shadow components, effectively enhancing the feature expression of reflectance and shadow components at different levels and reducing information loss. This strategy can effectively improve the accuracy of detail recovery and ensure the high quality of the final reflectance and shadow component output; 1) Output features of the fifth bilateral filtering structure layer Input the first deconvolution layer, the first deconvolution layer outputs features ; Expressed as: (14) in, represents the first deconvolution; represents the bias term; represents the convolution kernel of the first deconvolution; 2) Output features of the first deconvolution layer and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; Expressed as: (15) in, represents the second deconvolution; represents the bias term; represents the convolution kernel of the second deconvolution; 3) Output features of the second deconvolution layer and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; Expressed as: (16) in, represents the third deconvolution; represents the bias term; represents the convolution kernel of the third deconvolution; 4) Output features of the third deconvolution layer and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; Expressed as: (17) in, represents the fourth deconvolution; represents the bias term; represents the convolution kernel of the second deconvolution; 5) Output features of the fourth deconvolution layer and Laplace difference Perform element-by-element summation, which is then fed into the fifth deconvolution layer, which then outputs the final reflectivity component video. Expressed as: (18) in, represents the fifth deconvolution; represents the bias term; represents the convolution kernel of the fifth deconvolution; The fifth deconvolution layer outputs the final reflectivity component video.
[0028] The other steps and parameters are the same as those in the first to sixth embodiments.
[0029] Specific embodiment eight: This embodiment differs from any one of specific embodiments one to seven in that: in step 4, the satellite video image obtained in step 1 is sequentially upsampled for the first time, the second time, the third time, the fourth time, and the fifth time to obtain the feature ; The features Divide by the output feature of the fifth bilateral filtering structure layer , get the features ; The features Input the first deconvolution layer, the first deconvolution layer outputs features ; The first deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; The second deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; The third deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; The fourth deconvolution layer output features and Laplace difference Perform element-by-element summation, input the fifth deconvolution layer after summation, and the fifth deconvolution layer outputs the feature ,feature As the final shadow component video; The specific process is: 1) The satellite video image obtained in step 1 is upsampled for the first time, the second time, the third time, the fourth time and the fifth time in sequence to obtain the feature ; The features Divide by the output feature of the fifth bilateral filtering structure layer , get the features ; 2) The features Input the first deconvolution layer, the first deconvolution layer outputs features ; Expressed as: (19) in, represents the first deconvolution; represents the bias term; represents the convolution kernel of the first deconvolution; 3) Output features of the first deconvolution layer and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; Expressed as: (20) in, represents the second deconvolution; represents the bias term; represents the convolution kernel of the second deconvolution; 4) Output features of the second deconvolution layer and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; Expressed as: (twenty one) in, represents the third deconvolution; represents the bias term; represents the convolution kernel of the third deconvolution; 5) Output features of the third deconvolution layer and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; Expressed as: (twenty two) in, represents the fourth deconvolution; represents the bias term; represents the convolution kernel of the fourth deconvolution; 6) Output features of the fourth deconvolution layer and Laplace difference Perform element-by-element summation, input the fifth deconvolution layer after summation, and the fifth deconvolution layer outputs the feature ,feature As the final shadow component video; Expressed as: (twenty three) in, represents the fifth deconvolution; represents the bias term; Represents the convolution kernel of the fifth deconvolution.
[0030] Other steps and parameters are the same as those in Specific Embodiments 1 to 7-1.
[0031] Specific embodiment 9: This embodiment is a storage medium, which stores at least one instruction. The at least one instruction is loaded and executed by a processor to implement the satellite video intrinsic decomposition method based on a non-training network.
[0032] It should be understood that any method described herein may be provided as a computer program product, software, or computerized method, which may include a non-transitory machine-readable medium having instructions stored thereon, the instructions being used to program a computer system or other electronic device. The storage medium may include, but is not limited to, magnetic storage media, optical storage media; magneto-optical storage media including: read-only memory (ROM), random access memory (RAM), erasable programmable memory (e.g., EPROM and EEPROM), and flash memory layers; or other types of media suitable for storing electronic instructions.
[0033] Specific embodiment 10: This embodiment is a satellite video intrinsic decomposition device based on a non-training network, the device including a processor and a memory. It should be understood that including any device including a processor and a memory described in the present invention, the device may also include other units and modules that perform display, interaction, processing, control, and other functions through signals or instructions; At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the satellite video intrinsic decomposition method based on a non-training network.
[0034] The following examples are used to verify the beneficial effects of the present invention: Example 1: The data used in the experiment comes from six videos captured by the Jilin-1 satellite. This dataset contains different background scenes, such as woodlands, lakes, and cities, as well as a variety of moving objects, such as cars, trains, ships, and airplanes. The scenes in the videos involve various challenges, such as changing lighting, dark backgrounds, and densely packed objects. The videos are diverse and representative, making them a comprehensive and highly challenging experimental dataset. Figure 3shows a frame from the original satellite video used in the experiment. Figure 4 The processed reflectance component video results are shown. As can be seen from the image, the reflectance component is well extracted, and background information is effectively preserved and enhanced, with background details becoming clearer. The experimental results in Table 1 also confirm that after processing, the MSE (mean square error) and LMSE (logarithmic mean square error) values are significantly reduced, demonstrating that our method demonstrates significant superiority in reflectance component extraction. Figure 5 A frame from the shadow component video in the experiment is shown. The shadow component in the image is clearly visible and the boundary of the target is accurately restored. Figure 6 The visualization comparison of the proposed method and other satellite video motion target detection methods is shown. As can be seen from the figure, the proposed method has significant advantages in motion target detection accuracy and detail restoration, can more effectively extract and separate reflectance and shadow components, and demonstrates excellent performance in complex dynamic scenes.
[0035] Table 1 The present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.
Claims
1. A satellite video intrinsic decomposition method based on a non-trained network, characterized by: The specific process of the method is: Step 1: Get satellite video images, input the satellite video images into the first convolution layer, the first pooling layer, and the first bilateral filter structure layer in sequence. The first bilateral filter structure layer outputs the feature ; The output features of the first bilateral filtering structure layer Input the second convolution layer, the second pooling layer, and the second bilateral filter structure layer in sequence, and the second bilateral filter structure layer outputs the features ; The output features of the second bilateral filtering structure layer Input the third convolution layer, the third pooling layer, and the third bilateral filter structure layer in sequence, and the third bilateral filter structure layer outputs features ; The output features of the third bilateral filtering structure layer Input the fourth convolution layer, the fourth pooling layer, and the fourth bilateral filter structure layer in sequence, and the fourth bilateral filter structure layer outputs the feature ; Output features of the fourth bilateral filtering structure layer Input the fifth convolution layer, the fifth pooling layer, and the fifth bilateral filter structure layer in sequence, and the fifth bilateral filter structure layer outputs the features ; Step 2: Calculate the output features of the first bilateral filtering structure layer Laplace difference of ; Calculate the output features of the second bilateral filtering structure layer Laplace difference of ; Calculate the output features of the third bilateral filtering structure layer Laplace difference of ; Calculate the output features of the fourth bilateral filtering structure layer Laplace difference of ; Step 3: Output features of the fifth bilateral filtering structure layer Input the first deconvolution layer, the first deconvolution layer outputs features ; The first deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; The second deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; The third deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; The fourth deconvolution layer output features and Laplace difference Perform element-by-element summation, input the fifth deconvolution layer after summation, and the fifth deconvolution layer outputs the feature ,feature As the final reflectivity component video; Step 4: Perform the first upsampling, second upsampling, third upsampling, fourth upsampling and fifth upsampling on the satellite video image obtained in step 1 to obtain the feature ; The features Divide by the output feature of the fifth bilateral filtering structure layer , get the features ; The features Input the first deconvolution layer, the first deconvolution layer outputs features ; The first deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; The second deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; The third deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; The fourth deconvolution layer output features and Laplace difference Perform element-by-element summation, input the fifth deconvolution layer after summation, and the fifth deconvolution layer outputs the feature ,feature As the final shadow component video.
2. The satellite video eigendecomposition method based on a non-training network according to claim 1, wherein: In step 1, a satellite video image is obtained, and the satellite video image is sequentially input into the first convolution layer, the first pooling layer, and the first bilateral filter structure layer. The first bilateral filter structure layer outputs the feature ; The output features of the first bilateral filtering structure layer Input the second convolution layer, the second pooling layer, and the second bilateral filter structure layer in sequence, and the second bilateral filter structure layer outputs the features ; The output features of the second bilateral filtering structure layer Input the third convolution layer, the third pooling layer, and the third bilateral filter structure layer in sequence, and the third bilateral filter structure layer outputs features ; The output features of the third bilateral filtering structure layer Input the fourth convolution layer, the fourth pooling layer, and the fourth bilateral filter structure layer in sequence, and the fourth bilateral filter structure layer outputs the feature ; Output features of the fourth bilateral filtering structure layer Input the fifth convolution layer, the fifth pooling layer, and the fifth bilateral filter structure layer in sequence, and the fifth bilateral filter structure layer outputs the features ; The specific process is: Get satellite video images, input the satellite video images into the first convolutional layer, and the first convolutional layer outputs features ; The first convolutional layer output features Input the first pooling layer, the first pooling layer outputs features ; The first pooling layer output features Input the first bilateral filtering structure layer, the first bilateral filtering structure layer outputs features ; The first bilateral filtering structure layer outputs features Input the second convolutional layer, the second convolutional layer outputs features ; The second convolutional layer output features Input the second pooling layer, the second pooling layer outputs features ; The second pooling layer output features Input the second bilateral filtering structure layer, the second bilateral filtering structure layer outputs features ; The second bilateral filtering structure layer outputs features Input the third convolution layer, the third convolution layer outputs features ; The third convolutional layer output features Input the third pooling layer, the third pooling layer outputs features ; The output features of the third pooling layer Input the third bilateral filtering structure layer, the third bilateral filtering structure layer outputs features ; The output features of the third bilateral filtering structure layer Input the fourth convolution layer, the fourth convolution layer outputs features ; The fourth convolutional layer output features Input the fourth pooling layer, the fourth pooling layer outputs features ; The fourth pooling layer output features Input the fourth bilateral filtering structure layer, the fourth bilateral filtering structure layer outputs features ; The fourth bilateral filtering structure layer outputs features Input the fifth convolutional layer, the fifth convolutional layer outputs features ; The fifth convolutional layer output features Input the fifth pooling layer, the fifth pooling layer outputs features ; The fifth pooling layer output features Input the fifth bilateral filtering structure layer, the fifth bilateral filtering structure layer outputs features .
3. The satellite video eigendecomposition method based on a non-training network according to claim 2, wherein: The satellite video image is input into the first convolution layer, and the first convolution layer outputs the feature ; The first bilateral filtering structure layer outputs features Input the second convolutional layer, the second convolutional layer outputs features ; The second bilateral filtering structure layer outputs features Input the third convolution layer, the third convolution layer outputs features ; The output features of the third bilateral filtering structure layer Input the fourth convolution layer, the fourth convolution layer outputs features ; The fourth bilateral filtering structure layer outputs features Input the fifth convolutional layer, the fifth convolutional layer outputs features ; The specific process is: Input satellite video images into Convolutional layer, Convolutional layer output features , ; The expression is: (1) in, Indicates the Convolutional layer output features; Indicates passing Input features after layer convolution processing; is the position coordinate on the satellite video image, Number the timeframe; For the layer convolution kernel, is the convolution kernel input channel index, Output channel index of the convolution kernel; is the bias term of convolution; is the convolution operation; As the activation function, ReLU is used.
4. The satellite video eigendecomposition method based on a non-training network according to claim 3, wherein: The first convolutional layer outputs features Input the first pooling layer, the first pooling layer outputs features ; The second convolutional layer output features Input the second pooling layer, the second pooling layer outputs features ; The third convolutional layer output features Input the third pooling layer, the third pooling layer outputs features ; The fourth convolutional layer output features Input the fourth pooling layer, the fourth pooling layer outputs features ; The fifth convolutional layer output features Input the fifth pooling layer, the fifth pooling layer outputs features ; The specific process is: The resolution of the feature map is gradually reduced through the pooling operation. The pooling operation reduces the size of the feature map by downsampling. The expression is: (2) in, Indicates the pooling operation, using average pooling; Indicates the The pooling layer outputs features.
5. The satellite video eigendecomposition method based on a non-training network according to claim 4, characterized in that: The first pooling layer output features Input the first bilateral filtering structure layer, the first bilateral filtering structure layer outputs features ; The second pooling layer output features Input the second bilateral filtering structure layer, the second bilateral filtering structure layer outputs features ; The output features of the third pooling layer Input the third bilateral filtering structure layer, the third bilateral filtering structure layer outputs features ; The fourth pooling layer output features Input the fourth bilateral filtering structure layer, the fourth bilateral filtering structure layer outputs features ; The fifth pooling layer output features Input the fifth bilateral filtering structure layer, the fifth bilateral filtering structure layer outputs features ; The expression is: (3) in, Represents pixels Neighborhood; are the standard deviations in the spatial domain, respectively; is the normalization factor; It is The pooling layer outputs features; It is Pooling layer output features, position It's location The coordinates of the neighborhood points; Indicates the The output features of the bilateral filtering structure layer.
6. The satellite video eigendecomposition method based on a non-training network according to claim 5, characterized in that: In step 2, the output features of the first bilateral filtering structure layer are calculated Laplace difference of ; Calculate the output features of the second bilateral filtering structure layer Laplace difference of ; Calculate the output features of the third bilateral filtering structure layer Laplace difference of ; Calculate the output features of the fourth bilateral filtering structure layer Laplace difference of ; The specific process is: 1) Output features of the first bilateral filtering structure layer Upsample and get the upsampled features ; expressed as: (4) in, Represents an upsampling operation; Based on the output features of the first bilateral filtering structure layer And the upsampled features , calculate the Laplace difference ; expressed as: (5) 2) Output features of the second bilateral filtering structure layer Upsample and get the upsampled features ; expressed as: (6) in, Represents an upsampling operation; Based on the output features of the second bilateral filtering structure layer And the upsampled features , calculate the Laplace difference ; expressed as: (7) 3) Output features of the third bilateral filtering structure layer Upsample and get the upsampled features ; expressed as: (8) in, Represents an upsampling operation; Based on the output features of the third bilateral filtering structure layer And the upsampled features , calculate the Laplace difference ; expressed as: (9) 4) Output features of the fourth bilateral filtering structure layer Upsample and get the upsampled features ; expressed as: (10) in, Represents an upsampling operation; Based on the output features of the fourth bilateral filtering structure layer And the upsampled features , calculate the Laplace difference ; expressed as: (11) 5) Output features of the fourth bilateral filtering structure layer Upsample and get the upsampled features ; expressed as: (12) in, Represents an upsampling operation; Based on the output features of the fourth bilateral filtering structure layer and the upsampled features , calculate the Laplace difference ; expressed as: (13)。 7. The satellite video eigendecomposition method based on a non-training network according to claim 6, characterized in that: In step 3, the fifth bilateral filtering structure layer outputs the feature Input the first deconvolution layer, the first deconvolution layer outputs features ; The first deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; The second deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; The third deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; The fourth deconvolution layer output features and Laplace difference Perform element-by-element summation, input the fifth deconvolution layer after summation, and the fifth deconvolution layer outputs the feature ,feature As the final reflectivity component video; The specific process is: 1) Output features of the fifth bilateral filtering structure layer Input the first deconvolution layer, the first deconvolution layer outputs features ; expressed as: (14) in, represents the first deconvolution; represents the bias term; represents the convolution kernel of the first deconvolution; 2) Output features of the first deconvolution layer and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; expressed as: (15) in, represents the second deconvolution; represents the bias term; represents the convolution kernel of the second deconvolution; 3) Output features of the second deconvolution layer and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; expressed as: (16) in, represents the third deconvolution; represents the bias term; represents the convolution kernel of the third deconvolution; 4) Output features of the third deconvolution layer and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; expressed as: (17) in, represents the fourth deconvolution; represents the bias term; represents the convolution kernel of the second deconvolution; 5) Output features of the fourth deconvolution layer and Laplace difference Perform element-by-element summation, which is then fed into the fifth deconvolution layer, which then outputs the final reflectivity component video. This is expressed as: (18) in, represents the fifth deconvolution; represents the bias term; represents the convolution kernel of the fifth deconvolution; The fifth deconvolution layer outputs the final reflectivity component video.
8. The satellite video eigendecomposition method based on a non-training network according to claim 7, characterized in that: In step 4, the satellite video image obtained in step 1 is sequentially upsampled for the first time, the second time, the third time, the fourth time, and the fifth time to obtain the feature ; The features Divide by the output feature of the fifth bilateral filtering structure layer , get the features ; The features Input the first deconvolution layer, the first deconvolution layer outputs features ; The first deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; The second deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; The third deconvolution layer output features and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; The fourth deconvolution layer output features and Laplace difference Perform element-by-element summation, input the fifth deconvolution layer after summation, and the fifth deconvolution layer outputs the feature ,feature As the final shadow component video; The specific process is: 1) The satellite video image obtained in step 1 is upsampled for the first time, the second time, the third time, the fourth time and the fifth time in sequence to obtain the feature ; The features Divide by the output feature of the fifth bilateral filtering structure layer , get the features ; 2) The features Input the first deconvolution layer, the first deconvolution layer outputs features ; Expressed as: (19) in, represents the first deconvolution; represents the bias term; represents the convolution kernel of the first deconvolution; 3) Output features of the first deconvolution layer and Laplace difference Perform element-by-element summation, input the summation into the second deconvolution layer, and the second deconvolution layer outputs the feature ; Expressed as: (20) in, represents the second deconvolution; represents the bias term; represents the convolution kernel of the second deconvolution; 4) Output features of the second deconvolution layer and Laplace difference Perform element-by-element summation, input the summation into the third deconvolution layer, and the third deconvolution layer outputs the feature ; Expressed as: (21) in, represents the third deconvolution; represents the bias term; represents the convolution kernel of the third deconvolution; 5) Output features of the third deconvolution layer and Laplace difference Perform element-by-element summation, input the summation into the fourth deconvolution layer, and the fourth deconvolution layer outputs the feature ; Expressed as: (22) in, represents the fourth deconvolution; represents the bias term; represents the convolution kernel of the fourth deconvolution; 6) Output features of the fourth deconvolution layer and Laplace difference Perform element-by-element summation, input the fifth deconvolution layer after summation, and the fifth deconvolution layer outputs the feature ,feature As the final shadow component video; Expressed as: (23) in, represents the fifth deconvolution; represents the bias term; Represents the convolution kernel of the fifth deconvolution.
9. A storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the satellite video eigendecomposition method based on a non-training network as described in any one of claims 1 to 8.
10. A satellite video intrinsic decomposition device based on a non-training network, characterized in that: The device includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement a non-training network-based satellite video eigendecomposition method as described in any one of claims 1 to 8.