A video snow removal method, apparatus, equipment and medium

CN119007086BActive Publication Date: 2026-09-01HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411216748.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2026-09-01
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

[0002]降雪作为一种恶劣天气,经常出现在户外视频中,雪粒和条纹造成的退化效应严重降低了视频帧的可视性,进而阻碍了资助系统中视频处理算法的高级性能

Benefits of technology

获取雪视频序列,将雪视频序列输入至编码器中,提取视频帧特征图;将视频帧特征图输入至第一解耦模块,获取构成雪视频序列的物理特征,并去除物理特征中无用的退化层,得到空间特征图;构建第一网络和第二网络,获取预测结果的标签数据,根据预设的监督阶段将标签数据输入至第一网络和/或第二网络,以执行视频帧特征图的去雪任务,得到第一去雪结果和第二去雪结果;根据第一去雪结果和第二去雪结果构建正样本、锚样本和负样本,根据正样本、锚样本和负样本对视频帧特征图进行替换,得到超正样本;获取超正样本合成雪层,将超正样本合成雪层输入至解码器中,得到预测结果图像。根据本实施例的技术方案,能够去除视频中的雪图样,并提高视频中雪景场景的清晰度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119007086B_ABST
    Figure CN119007086B_ABST
Patent Text Reader

Abstract

This invention proposes a video desnowing method, including: acquiring a snow video sequence; inputting the snow video sequence into an encoder to extract video frame feature maps; inputting the video frame feature maps into a first decoupling module to acquire the physical features constituting the snow video sequence and removing useless degradation layers from the physical features to obtain a spatial feature map; constructing a first network and a second network to acquire label data of the prediction results; inputting the label data into the first network and / or the second network according to a preset supervision stage to perform the desnowing task of the video frame feature maps, obtaining a first desnowing result and a second desnowing result, and constructing positive samples, anchor samples, and negative samples to replace the video frame feature maps to obtain superpositive samples; acquiring a synthesized snow layer from the superpositive samples, and inputting the synthesized snow layer from the superpositive samples into a decoder to obtain the prediction result image. According to the technical solution of this embodiment, snow patterns in videos can be removed, improving the clarity of snow scenes in videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, device, and medium for desnowing videos. Background Technology

[0002] Snowfall, as a severe weather phenomenon, frequently appears in outdoor videos. The degradation effect caused by snow particles and streaks severely reduces the visibility of video frames, thereby hindering the advanced performance of video processing algorithms in funding systems. In existing technologies, due to the domain distribution differences between synthetic data and real-world data, current learning methods inevitably suffer from unrealistic training data and cannot handle real snowfall with unpredictable shapes and movements, resulting in blurry snow scenes in videos with low perceptual fidelity. Summary of the Invention

[0003] This invention aims to at least solve one of the technical problems existing in the prior art. To this end, this invention proposes a video desnow removal method, apparatus, device, and medium, which can remove snow patterns from videos and improve the clarity of snow scenes in videos.

[0004] In a first aspect, embodiments of the present invention provide a video desnowing method applied to a video desnowing system, the video desnowing system including an encoder, a first decoupling module, and a decoder, the first decoupling module being connected to the encoder and the decoder, and the video desnowing method including: Obtain a snow video sequence, input the snow video sequence into the encoder, and extract video frame feature maps; The video frame feature map is input to the first decoupling module to obtain the physical features constituting the snow video sequence, and the useless degradation layer in the physical features is removed to obtain the spatial feature map. Construct a first network and a second network to obtain the label data of the spatial feature map. Input the label data into the first network and / or the second network according to a preset supervision stage to perform the desnowing task of the video frame feature map and obtain a first desnowing result and a second desnowing result. Based on the first snow removal result and the second snow removal result, positive samples, anchor samples and negative samples are constructed. The video frame feature map is replaced based on the positive samples, the anchor samples and the negative samples to obtain super positive samples. The superpositive sample synthesized snow layer is obtained, and the superpositive sample synthesized snow layer is input into the decoder to obtain the prediction result image.

[0005] In some embodiments of the present invention, the video desnow removal system further includes a second converter, the second converter including a multilayer sensor block, and the decoupling module including a physical converter block and a first processing unit. The step of acquiring the physical features constituting the snow video sequence includes: Overlapping block embedding is performed on the video frame feature map to obtain linear blocks of the video frame feature map; Obtain the number of tags and the number of tags in the feature map of the video frame, and embed a first tag in the linear block according to the number of tags and the number of tags to obtain a tagged linear block; The marked linear block is input to the physical converter block for physical information separation, so as to integrate the first decoupling module into the second converter and replace the multilayer sensor block to obtain the second decoupling module.

[0006] In some embodiments of the present invention, the video desnowing system is further provided with a time-series decomposition router, and after replacing the multilayer sensing module with the first decoupling module, the method further includes: Multiple input markers are obtained from the feature map of the video frame, wherein each input marker is set with a corresponding parameter vector; Obtain the temporal dimension of the temporal decomposition router, and obtain the adaptive weights of the video frame feature map based on the input label and the temporal dimension; The first convex combination of the video frame feature map is obtained based on the adaptive weights and the input labels, wherein the first convex combination includes adaptive temporal information of the temporal dimension; Based on the first function preset in the second decoupling module and the adaptive timing information, the component labels of the video frame feature map are obtained; The time decomposition router dynamically decodes the physical features based on the component labels and the adaptive weights to obtain a first output label, and uses the first output label as the second convex combination of the video frame feature map.

[0007] In some embodiments of the present invention, the step of inputting the tag data to the first network and / or the second network according to a preset supervision phase includes: Semi-supervised processing is performed on the spatial feature map based on the first network and the second network to obtain labeled data and first unlabeled data of the first network, and second unlabeled data of the second network; Obtain the first background feature and the first snow feature of the labeled data, the second background feature and the second snow feature of the first unlabeled data, and the third background feature and the third snow feature of the second unlabeled data; The second background feature is combined with the first snow feature to form a positive sample, the third background feature and the second snow feature are combined to form an anchor sample, and the first background feature and the third snow feature are combined to form a negative sample. The contrast loss is calculated based on the positive sample, the anchor sample, and the negative sample.

[0008] In some embodiments of the present invention, the semi-supervised processing of the spatial feature map based on the first network and the second network includes: Complementary information is obtained based on the labeled data and the first unlabeled data; The first supervision loss and the second supervision loss of the first network, and the average mobility index of the second network are obtained based on the complementary information. The first network is updated based on the first supervision loss and the second supervision loss, and the second network is updated based on the average mobility index.

[0009] In some embodiments of the present invention, obtaining the synthetic snow layer of the hyperpositive sample includes: The hyperpositive samples are input into a preset Gaussian mixture model, and the hyperpositive samples are quantized to obtain the real snow layer distribution and the synthetic snow layer distribution. Calculate the divergence between the real snow layer distribution and the synthetic snow layer distribution to obtain the superpositive sample synthetic snow layer; The contrast loss of the spatial feature map is calculated based on the synthesized snow layer from the super-positive samples, the positive samples, and the negative samples.

[0010] In some embodiments of the present invention, after calculating the contrast loss based on the positive sample, the anchor sample, and the negative sample, the method further includes: When the supervision stage is fully supervised training, the label data is input into the first network to obtain the restored frame and clean frame of the spatial feature map, and the pixel-wise supervision loss of the spatial feature map is calculated based on the restored frame and the clean frame. When the supervision stage is semi-supervised training, the first unlabeled data is input into the first network and the second network respectively, and the pixel-wise consistency loss, perceptual contrast loss and prior loss of the spatial feature map are calculated.

[0011] In a second aspect, embodiments of the present invention provide a video desnow removal device, including at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, the instructions being executed by the at least one control processor to enable the at least one control processor to perform the video desnow removal method as described in the first aspect above.

[0012] Thirdly, embodiments of the present invention provide an electronic device including the video desnow removal device as described in the second aspect above.

[0013] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions for performing the video desnowing method as described in the first aspect above.

[0014] The video desnow removal method according to embodiments of the present invention has at least the following beneficial effects: A snow video sequence is acquired and input into an encoder to extract video frame feature maps. The video frame feature maps are then input into a first decoupling module to obtain the physical features constituting the snow video sequence, and useless degradation layers are removed from the physical features to obtain spatial feature maps. A first network and a second network are constructed to obtain label data for the prediction results. According to a preset supervision stage, the label data is input into the first network and / or the second network to perform a snow removal task on the video frame feature maps, obtaining a first snow removal result and a second snow removal result. Positive samples, anchor samples, and negative samples are constructed based on the first and second snow removal results. The video frame feature maps are then replaced using the positive samples, anchor samples, and negative samples to obtain super-positive samples. A snow layer synthesized from the super-positive samples is obtained and input into a decoder to obtain the prediction result image. According to the technical solution of this embodiment, snow patterns in videos can be removed, and the clarity of snow scenes in videos can be improved. Attached Figure Description

[0015] Figure 1 This is a flowchart of a video snow removal method provided in one embodiment of the present invention; Figure 2 This is a flowchart of obtaining the physical features constituting a snow video sequence according to an embodiment of the present invention; Figure 3 This is a flowchart of the process after replacing the multilayer sensing module according to one embodiment of the present invention; Figure 4 This is a flowchart provided by one embodiment of the present invention, showing how tag data is input to a first network and / or a second network according to a preset supervision stage; Figure 5 This is a flowchart of a semi-supervised processing of a spatial feature map based on a first network and a second network, provided in one embodiment of the present invention. Figure 6 This is a flowchart of obtaining a synthetic snow layer from a superpositive sample according to an embodiment of the present invention; Figure 7 This is a flowchart illustrating the calculation of contrast loss based on positive samples, anchor samples, and negative samples, provided in one embodiment of the present invention. Figure 8This is a structural diagram of a video snow removal device provided in another embodiment of the present invention. Detailed Implementation

[0016] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0017] In the description of this invention, it should be understood that the orientation descriptions, such as up, down, front, back, left, right, etc., are based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.

[0018] In the description of this invention, "several" means one or more, "more than" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.

[0019] In the description of this invention, unless otherwise explicitly defined, terms such as "set up," "install," and "connect" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this invention in conjunction with the specific content of the technical solution.

[0020] This invention provides a video desnowing method applied to a video desnowing system. The video desnowing system includes an encoder, a first decoupling module, and a decoder. The first decoupling module is connected to the encoder and the decoder. The video desnowing method includes: acquiring a snow video sequence; inputting the snow video sequence into the encoder to extract video frame feature maps; inputting the video frame feature maps into the first decoupling module to acquire the physical features constituting the snow video sequence and removing useless degradation layers from the physical features to obtain a spatial feature map; constructing a first network and a second network to acquire label data of the spatial feature map; inputting the label data into the first network and / or the second network according to a preset supervision stage to perform the desnowing task of the video frame feature map, obtaining a first desnowing result and a second desnowing result; constructing positive samples, anchor samples, and negative samples based on the first and second desnowing results; replacing the video frame feature map based on the positive samples, anchor samples, and negative samples to obtain super-positive samples; acquiring a synthesized snow layer from the super-positive samples; inputting the synthesized snow layer from the super-positive samples into the decoder to obtain a predicted result image. According to the technical solution of this embodiment, snow patterns in videos can be effectively removed, and the clarity of snow scenes in videos can be improved. A video desnowing system is developed based on a Mean-Teacher architecture. The video desnowing system consists of an encoder, a first decoupling module, and an encoder. Using a hybrid machine composed of synthetic and real data, based on the Mean-Teacher architecture, unlabeled real-world data is introduced in a semi-supervised manner during the training phase to enhance the system's generalization ability in various real-world scenarios. A distribution-driven contrastive regularization method is introduced to prevent the deep model from being disturbed by the different shapes and motion variations of snow in synthetic and real data. Furthermore, the likelihood algorithm of a Gaussian mixture model is used to capture synthetic snow layers most similar to real snow, and super-positive samples are obtained by approximating the distribution of real snow components. By replacing snow layers in positive samples with super-positive sample snow layers and conversely replacing the background in negative samples, snow-invariant information is preserved and highlighted. This improves the restoration of snow in video sequences, providing clearer content and enhanced perceptual fidelity in both synthetic and real snow scenes.

[0021] The control method of the present invention will be further described below with reference to the accompanying drawings.

[0022] Reference Figure 1 , Figure 1 This is a flowchart of a video snow removal method provided in an embodiment of the present invention. The video snow removal method includes, but is not limited to, the following steps: Step S11: Obtain the snow video sequence, input the snow video sequence into the encoder, and extract the video frame feature map; It should be noted that when the video desnowing system performs desnowing operations on a video, it first inputs a snow video sequence into the encoder. The encoder uses a general ConvNeXt backbone network to extract video frame feature maps. In this embodiment, the video desnowing system based on the ConvNeXt encoder can accurately identify snow areas in the video and effectively remove snowflake interference areas. Through accurate feature extraction and efficient computational processing, the system can restore clearer and more realistic video content, improving the visual quality of the video.

[0023] Step S12: Input the video frame feature map into the first decoupling module to obtain the physical features that constitute the snow video sequence, and remove the useless degradation layer in the physical features to obtain the spatial feature map. It should be noted that, in order to enhance the decoupling capability of the first decoupling module, a physical transformer block in the video system is first defined. The physical transformer block is equipped with multiple processing units. Through these multiple processing units, the snow video sequence is decomposed into multiple physical features and input to the first decoupling module. The first decoupling module can identify and separate the physical features in the feature map of the video frame, such as the dynamic features of snowflakes (falling speed, direction), static features (snow thickness, snow distribution), and background features (such as buildings, trees, etc.). It also removes the useless degradation layer in the physical features to obtain the spatial feature map of the snow video.

[0024] It should be noted that after obtaining the spatial feature map, the features of each component of the second decoupling module are concatenated with the corresponding merged features, and then input into their respective decoders to enhance the component features. These three decoders each consist of three convolutional layers with upsampling. Finally, these enhanced features are used for the final recovery in the subsequent prior-guided recovery module. We use the first formula to simultaneously remove snow from the frame, thereby decomposing the frame into three distinct components S, A, and T in the feature space.

[0025] The first formula is: ; in, The images represent videos damaged by snowfall. J represents the corresponding clear video, T represents the transmission image, A represents atmospheric light, and S represents the snowfall image.

[0026] According to the first formula, the prior-based recovery process can be described as follows: ; Wherein, S is the first component feature, A is the second component feature, and T is the third component feature.

[0027] Step S13: Construct a first network and a second network, obtain the label data of the prediction results, and input the label data into the first network and / or the second network according to the preset supervision stage to perform the desnowing task of the video frame feature map and obtain the first desnowing result and the second desnowing result. It should be noted that by constructing a first network and a second network, and processing the video frames through different paths or processing strategies, the first and second desnow removal results of the background features are obtained. At the same time, supervised learning of the labeled data ensures the effectiveness and accuracy of the desnow removal task.

[0028] Step S14: Construct positive samples, anchor samples, and negative samples based on the first and second snow removal results. Replace the video frame feature map based on the positive samples, anchor samples, and negative samples to obtain super-positive samples. It should be noted that positive samples, anchor samples, and negative samples are constructed using the first and second desnowing results. Based on multiple samples, the feature map of the video frame is replaced to obtain super-positive samples, which further improves the generalization ability and desnowing effect of the video desnowing system and generates super-positive samples to improve the processing quality of video desnowing.

[0029] Step S15: Obtain the superpositive sample synthesized snow layer, input the superpositive sample synthesized snow layer into the decoder, and obtain the prediction result image.

[0030] It should be noted that by synthesizing a snow layer and inputting it into the decoder, a snowflake layer in the real environment is simulated, and the snowflake layer is removed to make the predicted image closer to the actual scene. This achieves effective snow removal processing of snow video sequences and significantly improves the quality and clarity of the predicted image.

[0031] It should be noted that this embodiment learns and preserves the invariant information of snow by replacing the snow layer in positive samples with the snow layer in super-positive samples, and conversely replacing the background in negative samples. This improves the restoration reflected in snow video sequences, providing clearer content and enhanced perceptual fidelity in both synthetic and real snow scenes.

[0032] In another embodiment, the video desnow removal system further includes a second converter, which includes a multilayer sensing module. The decoupling module includes a physical converter block and a first processing unit. (Refer to...) Figure 2 ,exist Figure 1 Step S12 of the illustrated embodiment also includes, but is not limited to, the following steps: Step S21: Overlapping block embedding is performed on the video frame feature map to obtain linear blocks of the video frame feature map; Step S22: Obtain the number of tags and the number of tags in the feature map of the video frame; embed the first tag in the linear block according to the number of tags and the number of tags to obtain the tag linear block; Step S23: Input the marked linear block into the physical converter block for physical information separation, so as to integrate the first decoupling module into the second converter, replace the multilayer sensor block, and obtain the second decoupling module.

[0033] It should be noted that the encoder acquires video frame feature maps and performs overlapping block embedding on each video frame feature map to obtain linear blocks of the video frame feature map. All linear blocks are then embedded into a first marker, where the first marker is... Where m is the number of tags in a frame and d is the number of tag channels, Y is then input into the transformer block for separation of physical dependency information. In the physical transformer block, the feedforward network of the first transformer is first improved by fusing the feedforward network to enhance feature fusion. The first decoupling module is integrated into the second transformer of the physical transformer block, replacing the multilayer perceptron block. The proposed module contains specific units corresponding to different physical components, namely snow units, transmission map units, and atmospheric light units.

[0034] Additionally, in one embodiment, the video desnowing system also includes a time-series decomposition router, as referred to Figure 3 ,exist Figure 2 Following step 23 in the illustrated embodiment, the following steps may also be included, but are not limited to: Step S31: Obtain multiple input markers of the video frame feature map, wherein each input marker is set with a corresponding parameter vector; Step S32: Obtain the temporal dimension of the temporal decomposition router, and obtain the adaptive weights of the video frame feature map based on the input label and the temporal dimension; Step S33: Obtain the first convex combination of video frame feature maps based on adaptive weights and input labels, wherein the first convex combination includes adaptive temporal information in the temporal dimension; Step S34: Obtain component labels of video frame feature maps based on the preset first function and adaptive timing information within the second decoupling module; Step S35: The time decomposition router dynamically decodes the physical features based on the component labels and adaptive weights to obtain the first output label, and uses the first output label as the second convex combination of the video frame feature map.

[0035] It should be noted that the input label of a sequence is represented as... The first function is specifically represented as {fj:Rd→Rd}j=1n. The snow unit, transmission graph unit, and atmospheric optical unit each process a time-adaptive marker, and each marker has a corresponding d-dimensional parameter vector, represented as Γ∈Rd×n. Based on the time-series decomposition, the time dimension in the router is determined. The time-series adaptive weights Qij are obtained, and the adaptive weights can be expressed by the following formula: ; Therefore, the weighted label is obtained based on the input label, the first convex combination, and the adaptive timing information. The weighted label can be expressed by the following formula: Z; in, For weighted labeling, Let Z be the first convex combination, and Z be the input marker.

[0036] Then, the first function is applied to each input tag to obtain the output component tag, which can be expressed by the following formula: ; Based on the time-decomposition router, physical features are dynamically decoded according to component labels and decomposition weights to obtain output labels. The output label C is calculated as the second convex combination of all output component labels, which can be expressed by the following formula: ; Here, D is the decomposition weight, i.e., the softmax result of Z·Γ in the expert dimension. Existing sparse hybrid expert algorithms are usually discrete and therefore non-differentiable. However, the second decoupling module in this embodiment employs continuous and differentiable operations. It effectively utilizes all time markers and physical feature information acquired by the units to extract physics-specific features.

[0037] Additionally, in one embodiment, reference is made to Figure 4 ,exist Figure 1 Step S13 of the illustrated embodiment also includes, but is not limited to, the following steps: Step S41: Perform semi-supervised processing on the spatial feature map according to the first network and the second network to obtain labeled data and first unlabeled data of the first network, and second unlabeled data of the second network; Step S42: Obtain the first background feature and the first snow feature of the labeled data, the second background feature and the second snow feature of the first unlabeled data, and the third background feature and the third snow feature of the second unlabeled data; Step S43: Combine the second background feature with the first snow feature to form a positive sample, combine the third background feature with the second snow feature to form an anchor sample, and combine the first background feature with the third snow feature to form a negative sample. Step S44: Calculate the contrast loss based on the positive sample, anchor sample, and negative sample.

[0038] It should be noted that the first background features are obtained from the labeled data in the first network. and the characteristics of the first snow And to obtain second background features from the first unlabeled data in the first network. Second snow characteristics Furthermore, we also extract a third background feature from the second unlabeled data in the second network. and the characteristics of the third snow According to the snow synthesis formula The expression J(x) = J(x) + S(x) recombines background and snow features from labeled data, the first unlabeled data, and the second unlabeled data. To preserve and highlight information unrelated to snow, we replace the corresponding parts with snow features in positive samples and vice versa in negative samples. Specifically, we will... and The combination is set as a positive sample. and The combination is set as the anchor sample. and enhanced The combination of positive samples, anchor samples, and negative samples is set as negative samples, and the contrast loss is calculated based on the recombined positive samples, anchor samples, and negative samples.

[0039] Additionally, in one embodiment, reference is made to Figure 5 ,exist Figure 4 Step S41 in the illustrated embodiment also includes, but is not limited to, the following steps: Step S51: Obtain complementary information based on the labeled data and the first unlabeled data; Step S52: Obtain the first supervision loss and the second supervision loss of the first network, and the average mobility index of the second network, based on complementary information; In step S53, the first network is updated based on the first supervision loss and the second supervision loss, and the second network is updated based on the average mobility index.

[0040] It should be noted that, to enhance the system's generalization ability on real data, we introduce semi-supervised learning in the video snow removal task. Semi-supervised learning enables the learning system to obtain complementary information from labeled synthetic data and unlabeled real data. Our semi-supervised learning framework follows typical settings. During training, the first network is updated by minimizing the supervised loss and unsupervised loss, while the second network updates the supervised loss using an exponential moving average, which can be expressed by the following formula: ) ; To constrain the output of the first network, Charbonnier loss and perceptual loss are employed to improve the visual quality of the recovered results. Features are extracted from layers 3, 8, and 15 of the pre-trained VGG-16 to compute the perceptual loss. FocalFrequency loss is introduced to address different artifacts in the model's spectral response to different regions of the image, and can be expressed by the following formula: 1 2 ; The unsupervised loss uses pixel-level Charbonnier loss as the consistency loss between the first and second unsupervised networks to ensure that the two networks generate consistent results. This can be expressed by the following formula: 3 ; Additionally, in one embodiment, reference is made to Figure 6 ,exist Figure 1 Step S15 of the illustrated embodiment also includes, but is not limited to, the following steps: Step S61: Input the hyperpositive samples into the preset Gaussian mixture model, quantize the hyperpositive samples, and obtain the real snow layer distribution and the synthetic snow layer distribution. Step S62: Calculate the divergence between the real snow layer distribution and the synthetic snow layer distribution to obtain the superpositive sample synthetic snow layer; Step S63: Calculate the contrast loss of the spatial feature map based on the synthesized snow layer from the super-positive samples, the positive samples, and the negative samples.

[0041] It should be noted that due to the distributional differences between synthetic and real snow, the snow layers generated by our network often exhibit significant differences. Therefore, our goal is to obtain a super-positive sample, representing the synthetic snow layer whose characteristics are very close to those of real snow, to be used as a positive sample. Since real snow has inherently diverse structures due to its different generation states and viewing angles, it can be represented using a Gaussian Mixture Model (GMM). Using GMM can accurately approximate the distribution of real snow layers, effectively capturing multiple patterns in the data. The distribution of the snow layer can be represented as: ; Contrast loss can be expressed by the following formula: ; Additionally, in one embodiment, reference is made to Figure 7 ,exist Figure 1 Step S11 in the illustrated embodiment also includes, but is not limited to, the following steps: Step S71: When the supervision stage is fully supervised training, the label data is input into the first network to obtain the restored frame and clean frame of the spatial feature map. The pixel-wise supervision loss of the spatial feature map is calculated based on the restored frame and clean frame. Step S72: When the supervised training stage is semi-supervised, the first unlabeled data is input into the first network and the second network respectively, and the pixel-wise consistency loss, perceptual contrast loss and prior loss of the spatial feature map are calculated.

[0042] It should be noted that, since the comparison is based directly on pixel values, the pixel-by-pixel supervised loss can ensure that the recovered frame is as close as possible to the real frame visually. It is quite effective for detail preservation and color accuracy. By comparing the outputs of the first network and the second network to the same input, the supervised effect of the labeled data is simulated to a certain extent, which increases the robustness and generalization ability of the model.

[0043] Furthermore, during supervised training, labeled data is fed into the student network, and pixel-wise supervised loss is computed between restored and clean frames. During semi-supervised training, unlabeled data is fed into both the first and second networks, and pixel-wise consistency loss, perceptual contrastive loss, and prior loss are computed to normalize the student network. Additionally, we utilize distribution-driven contrastive regularization loss to prevent the model from being negatively impacted by distributional differences between synthetic and real data.

[0044] like Figure 8 As shown, Figure 8 This is a structural diagram of a video snow removal device provided in one embodiment of the present invention. The present invention also provides a video snow removal device, comprising: The processor 801 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 802 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 802 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 using the video desnowing method of the embodiments of this application. The 803 input / output interface is used to implement information input and output. The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804); The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.

[0045] This application also provides an electronic device, including the video desnow removal device described above.

[0046] This application embodiment also provides a storage medium, which is a computer-readable storage medium, storing a computer program that, when executed by a processor, implements the above-described video desnowing method.

[0047] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof. The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate, and may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0048] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically include computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0049] The above provides a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.

Claims

1. A video snow removal method, characterized in that, An application is made in a video desnowing system, the video desnowing system comprising an encoder, a first decoupling module, and a decoder, the first decoupling module being connected to the encoder and the decoder, and the video desnowing method comprising: Obtain a snow video sequence, input the snow video sequence into the encoder, and extract video frame feature maps; The video frame feature map is input to the first decoupling module to obtain the physical features constituting the snow video sequence, and the useless degradation layer in the physical features is removed to obtain the spatial feature map. A first network and a second network are constructed to obtain labeled data of the prediction results. The labeled data is input into the first network and the second network according to a preset supervision stage to perform a snow removal task on the video frame feature map, obtaining a first snow removal result and a second snow removal result. Semi-supervised processing is performed on the spatial feature map using the first network and the second network to obtain labeled data and first unlabeled data for the first network, and second unlabeled data for the second network. A first background feature and a first snow feature of the labeled data, a second background feature and a second snow feature of the first unlabeled data, and a third background feature and a third snow feature of the second unlabeled data are obtained. The second background feature and the first snow feature are combined into a positive sample, the third background feature and the second snow feature are combined into an anchor sample, and the first background feature and the third snow feature are combined into a negative sample. A contrast loss is calculated based on the positive sample, the anchor sample, and the negative sample. Based on the first snow removal result and the second snow removal result, positive samples, anchor samples and negative samples are constructed. The video frame feature map is replaced based on the positive samples, the anchor samples and the negative samples to obtain super positive samples. The superpositive sample synthesized snow layer is obtained, and the superpositive sample synthesized snow layer is input into the decoder to obtain the prediction result image.

2. The video desnow removal method according to claim 1, characterized in that, The video desnow removal system also includes a second converter, which has a multilayer sensor block. The decoupling module includes a physical converter block and a first processing unit. The acquisition of the physical features constituting the snow video sequence includes: Overlapping block embedding is performed on the video frame feature map to obtain linear blocks of the video frame feature map; Obtain the number of tags and the number of tags in the feature map of the video frame, and embed a first tag in the linear block according to the number of tags and the number of tags to obtain a tagged linear block; The marked linear block is input to the physical converter block for physical information separation, so as to integrate the first decoupling module into the second converter and replace the multilayer sensor block to obtain the second decoupling module.

3. The video desnow removal method according to claim 2, characterized in that, The video desnow removal system also includes a time-series decomposition router. After replacing the multilayer sensing module with the first decoupling module, the method further includes: Multiple input markers are obtained from the feature map of the video frame, wherein each input marker is set with a corresponding parameter vector; Obtain the temporal dimension of the temporal decomposition router, and obtain the adaptive weights of the video frame feature map based on the input label and the temporal dimension; The adaptive temporal information of the video frame feature map is obtained based on the adaptive weights and the input labels; Based on the first function preset in the second decoupling module and the adaptive timing information, the component labels of the video frame feature map are obtained; The time decomposition router dynamically decodes the physical features based on the component labels and the adaptive weights to obtain a first output label, and uses the first output label as a convex combination of the video frame feature map.

4. The video desnow removal method according to claim 1, characterized in that, The semi-supervised processing of the spatial feature map based on the first network and the second network includes: Complementary information is obtained based on the labeled data and the first unlabeled data; The first supervision loss and the second supervision loss of the first network, and the average mobility index of the second network are obtained based on the complementary information. The first network is updated based on the first supervision loss and the second supervision loss, and the second network is updated based on the average mobility index.

5. The video desnow removal method according to claim 1, characterized in that, The process of obtaining the synthetic snow layer from the hyperpositive sample includes: The hyperpositive samples are input into a preset Gaussian mixture model, and the hyperpositive samples are quantized to obtain the real snow layer distribution and the synthetic snow layer distribution. Calculate the divergence between the real snow layer distribution and the synthetic snow layer distribution to obtain the superpositive sample synthetic snow layer; The contrast loss of the spatial feature map is calculated based on the synthesized snow layer from the super-positive samples, the positive samples, and the negative samples.

6. The video desnow removal method according to claim 1, characterized in that, After calculating the contrast loss based on the positive sample, the anchor sample, and the negative sample, the method further includes: When the supervision stage is fully supervised training, the label data is input into the first network to obtain the restored frame and clean frame of the spatial feature map, and the pixel-wise supervised loss of the spatial feature map is calculated based on the restored frame and the clean frame. When the supervision stage is semi-supervised training, the first unlabeled data is input into the first network and the second network respectively, and the pixel-wise consistency loss, perceptual contrast loss and prior loss of the spatial feature map are calculated.

7. A video snow removal device, characterized in that, It includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, the instructions being executed by the at least one control processor to enable the at least one control processor to perform the video desnowing method as described in any one of claims 1 to 6.

8. An electronic device, characterized in that, Includes the video snow removal device as described in claim 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the video desnowing method as described in any one of claims 1 to 6.