Image snow removal method and system based on zero sample learning
Through the image snow removal method based on zero-sample learning, the preset zero-sample learning model is trained to process snow images, solving the problems of poor snow removal effect and complex network design in the existing technology, and achieving efficient image snow removal effect.
Patent Information
- Application Number
- CN202510585339.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-27
AI Technical Summary
The existing deep learning-based video image removal networks are not effective when processing images in real snow weather, and the network design is cumbersome, with large parameters and slow processing speed.
The image snow removal method based on zero-sample learning is adopted, and the preset zero-sample learning model is trained, including mask estimation network, snowflake foreground estimation network and clear image estimation network, to process the snow image to be processed, and the target mask image, snowflake foreground image and snow removal image are obtained.
The video image removal effect and processing efficiency are improved, and the problem of difficulty in obtaining data sets and the problem of threshold bias in the training network of synthetic data sets is avoided.
Smart Images

Figure CN120219206A_ABST
Abstract
Description
Background Art
[0002] In current deep learning-based video image snow removal networks, they are all end-to-end deep networks. By training the network with a large number of synthetic images, the designed network can obtain good results when processing synthetic images, but it is difficult to obtain good results for images captured in real snow weather. In addition, although the current convolutional neural network method can adaptively extract features during the snow removal process and also reduces the requirements for estimating the scene, snow grain size, and snow density in the image, the current deep learning-based method has a cumbersome network design and a large number of parameters, and there is still room for improvement in processing speed.
[0003] Therefore, there is an urgent need to provide a technical solution to solve the above problems. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides an image snow removal method and system based on zero-shot learning.
[0005] In a first aspect, the present invention provides an image snow removal method based on zero-shot learning. The technical solution of this method is as follows:
[0006] S1. Based on multiple original sample snow images and in combination with a snow image model, train a preset zero-shot learning model for image snow removal to obtain a target zero-shot learning model; wherein, the preset zero-shot learning model includes: a mask estimation network, a snowflake foreground estimation network, and a clear image estimation network; the snow image model is used to represent the correlation between the mask image extracted by the mask estimation network, the snowflake foreground image extracted by the snowflake foreground estimation network, the clear image extracted by the clear image estimation network, and the snow image;
[0007] S2. Input the snow image to be processed into the target zero-shot learning model to obtain the target mask image and target snowflake foreground image corresponding to the snow image to be processed, and input them into the snow image model to obtain the target snow removal image corresponding to the snow image to be processed.
[0008] The beneficial effects of an image snow removal method based on zero-shot learning of the present invention are as follows:
[0009] The method of the present invention performs image snow removal processing by constructing a lightweight zero-shot learning model, while avoiding the difficulty of obtaining a dataset and the problem of threshold bias caused by training the network with a synthetic dataset, and also improving the video image snow removal effect and processing efficiency.
[0010] On the basis of the above solution, an image snow removal method based on zero-shot learning of the present invention can also be improved as follows.
[0011] In an optional manner, S1 includes:
[0012] S11. Input any original snow image into the preset zero-shot learning model to obtain the first sample mask image, the first sample snow foreground image, and the first sample clear image corresponding to the any original snow image, and input them into the snow image model to obtain the reconstructed sample snow image corresponding to the any original snow image;
[0013] S12. Input the reconstructed sample snow image corresponding to the any original snow image into the preset zero-shot learning model to obtain the second sample mask image, the second sample snow foreground image, and the second sample clear image corresponding to the any original snow image;
[0014] S13. According to the first sample mask image, the first sample snow foreground image, the first sample clear image, the second sample mask image, the second sample snow foreground image, and the second sample clear image corresponding to the any original snow image, and in combination with the target loss function of the preset zero-shot learning model, obtain the loss value corresponding to the any original snow image until the loss values corresponding to each original snow image are obtained;
[0015] S14. Use all the loss values to optimize the network parameters of the preset zero-shot learning model to obtain an optimized zero-shot learning model. Take the optimized zero-shot learning model as the preset zero-shot learning model and return to execute S11 until the optimized zero-shot learning model meets the preset iterative training conditions, and then determine the optimized zero-shot learning model as the target zero-shot learning model.
[0016] In an alternative way, S11 includes:
[0017] S111. Input the any original snow image into the mask estimation network, the snow foreground estimation network, and the clear image estimation network respectively to obtain the first sample mask image, the first sample snow foreground image, and the first sample clear image corresponding to the any original snow image;
[0018] S112. Input the first sample mask image, the first sample snow foreground image, and the first sample clear image corresponding to the any original snow image into the snow image model to obtain the reconstructed sample snow image corresponding to the any original snow image; where the snow image model is: I2(x) = Z1(x) × S1(x) + (1 - Z1(x)) × J1(x); Z1(x) is the first sample mask image, S1(x) is the first sample snow foreground image, J1(x) is the first sample clear image, and I2(x) is the reconstructed sample snow image.
[0019] In an alternative approach, the step of obtaining the loss value corresponding to any one of the original sample snow images based on the first sample mask image, the first sample snow foreground image, the first sample clear image, the second sample mask image, the second sample snow foreground image, and the second sample clear image corresponding to the original sample snow image, and in combination with the objective loss function of the preset zero-shot learning model, includes:
[0020] Inputting the first sample mask image, the first sample snow foreground image, the first sample clear image, the second sample mask image, the second sample snow foreground image, and the second sample clear image corresponding to any one of the original sample snow images into the objective loss function to obtain the loss value corresponding to any one of the original sample snow images; wherein, the objective loss function is: is the loss value, is the image similarity loss, is the snow foreground image loss, is the image saturation penalty loss, is the mask image loss, and ω1, ω2, ω3, and ω4 are loss coefficients.
[0021] In an alternative approach, the mask image is: an image after secondary masking.
[0022] In a second aspect, the present invention provides an image desnowing system based on zero-shot learning, and the technical solution of the system is as follows:
[0023] Including: a training module and a running module;
[0024] The training module is configured to: train a preset zero-shot learning model for image desnowing based on a plurality of original sample snow images and in combination with a snow image model to obtain a target zero-shot learning model; wherein, the preset zero-shot learning model includes: a mask estimation network, a snow foreground estimation network, and a clear image estimation network; the snow image model is used to characterize the correlation between the mask image extracted by the mask estimation network, the snow foreground image extracted by the snow foreground estimation network, the clear image extracted by the clear image estimation network, and the snow image;
[0025] The running module is configured to: input the snow image to be processed into the target zero-shot learning model to obtain the target mask image and the target snow foreground image corresponding to the snow image to be processed and input them into the snow image model to obtain the target desnowed image corresponding to the snow image to be processed.
[0026] The beneficial effects of an image desnowing system based on zero-shot learning according to the present invention are as follows:
[0027] The system of the present invention performs image snow removal by constructing a lightweight zero - shot learning model, which not only avoids the difficulty in obtaining a dataset and the problem of threshold deviation caused by training a network with a synthetic dataset, but also improves the snow removal effect and processing efficiency of video images.
[0028] Based on the above - mentioned solution, an image snow - removal system based on zero - shot learning of the present invention can also be improved as follows.
[0029] In an alternative embodiment, the training module includes: a first training module, a second training module, a third training module, and an iterative training module;
[0030] The first training module is configured to: input any original sample snow image into the preset zero - shot learning model, obtain the first sample mask image, the first sample snowflake foreground image, and the first sample clear image corresponding to the any original sample snow image, and input them into the snow image model to obtain the reconstructed sample snow image corresponding to the any original sample snow image;
[0031] The second training module is configured to: input the reconstructed sample snow image corresponding to the any original sample snow image into the preset zero - shot learning model, obtain the second sample mask image, the second sample snowflake foreground image, and the second sample clear image corresponding to the any original sample snow image;
[0032] The third training module is configured to: according to the first sample mask image, the first sample snowflake foreground image, the first sample clear image, the second sample mask image, the second sample snowflake foreground image, and the second sample clear image corresponding to the any original sample snow image, and in combination with the objective loss function of the preset zero - shot learning model, obtain the loss value corresponding to the any original sample snow image until the loss values corresponding to each original sample snow image are obtained;
[0033] The iterative training module is configured to: use all the loss values to optimize the network parameters of the preset zero - shot learning model, obtain the optimized zero - shot learning model, use the optimized zero - shot learning model as the preset zero - shot learning model, and return to call the first training module until the optimized zero - shot learning model meets the preset iterative training conditions, and then determine the optimized zero - shot learning model as the target zero - shot learning model.
[0034] In an alternative embodiment, the first training module is specifically configured to:
[0035] Input the any original sample snow image into the mask estimation network, the snowflake foreground estimation network, and the clear image estimation network respectively to obtain the first sample mask image, the first sample snowflake foreground image, and the first sample clear image corresponding to the any original sample snow image;
[0036] Input the first sample mask image, the first sample snowflake foreground image, and the first sample clear image corresponding to any one of the original sample snow images into the snow image model to obtain the reconstructed sample snow image corresponding to any one of the original sample snow images; wherein, the snow image model is: I2(x) = Z1(x) × S1(x) + (1 - Z1(x)) × J1(x); Z1(x) is the first sample mask image, S1(x) is the first sample snowflake foreground image, J1(x) is the first sample clear image, and I2(x) is the reconstructed sample snow image.
[0037] In a third aspect, the technical solution of an electronic device according to the present invention is as follows:
[0038] It includes a memory, a processor, and a program stored on the memory and running on the processor. When the processor executes the program, it implements the steps of the image snow removal method based on zero-shot learning according to the present invention.
[0039] In a fourth aspect, the technical solution of a computer-readable storage medium provided by the present invention is as follows:
[0040] Instructions are stored in the computer-readable storage medium. When the computer-readable storage medium reads the instructions, it causes the computer-readable storage medium to execute the steps of the image snow removal method based on zero-shot learning according to the present invention.
[0041] The above description is only an overview of the technical solution of the present invention. In order to be able to more clearly understand the technical means of the present invention, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically listed below. Description of the Drawings
[0042] The drawings are only used to illustrate the embodiments and are not considered to limit the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0043] Figure 1 is a schematic flow chart of an embodiment of an image snow removal method based on zero-shot learning according to the present invention;
[0044] Figure 2 is a schematic structural diagram of a preset zero-shot learning model;
[0045] Figure 3 is a schematic structural diagram of the overall network;
[0046] Figure 4 is a schematic diagram of the training principle;
[0047] Figure 5Schematic diagram of the principle of the snow image model;
[0048] Figure 6 Schematic diagram of the structure of an embodiment of an image snow removal system based on zero - shot learning according to the present invention;
[0049] Figure 7 Schematic diagram of the structure of an embodiment of an electronic device according to the present invention. Detailed implementation manners
[0050] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein.
[0051] Figure 1 The flowchart of an embodiment of an image snow removal method based on zero - shot learning provided by the present invention is shown. This image snow removal method based on zero - shot learning can be executed by an electronic device such as a terminal device or a server. Among them, the terminal device can be any fixed or mobile terminal such as a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle - mounted device, a wearable device, etc. The server can be a single server or a server cluster composed of multiple servers. Any electronic device can implement the image snow removal method based on zero - shot learning by a processor calling computer - readable instructions stored in a memory. As Figure 1 shown, it includes the following steps:
[0052] S1. Based on a plurality of original sample snow images and in combination with a snow image model, train a preset zero - shot learning model for image snow removal to obtain a target zero - shot learning model.
[0053] Among them, as Figure 2 shown, the preset zero - shot learning model includes: a mask estimation network, a snowflake foreground estimation network, and a clear image estimation network. The snow image model is used to characterize the correlation relationship between the mask image extracted by the mask estimation network, the snowflake foreground image extracted by the snowflake foreground estimation network, the clear image extracted by the clear image estimation network, and the snow image. The mask image in this embodiment is an image after secondary masking.
[0054] Among them, the original sample snow images are randomly selected snow images for model training, and the snow image refers to an image containing snowflakes or / and snow scenes. The target zero - shot learning model is a trained zero - shot learning model.
[0055] It should be noted that zero-shot learning means that the model can directly predict or classify new and unseen categories without having seen any labeled examples of the target categories.
[0056] S2. Input the snow image to be processed into the target zero-shot learning model to obtain the target mask image and the target snow foreground image corresponding to the snow image to be processed, and input them into the snow image model to obtain the target snow-removed image corresponding to the snow image to be processed.
[0057] Among them, the snow image to be processed is the snow image that needs to be snow-removed in this embodiment. The target snow-removed image is the clear image obtained after snow-removing the snow image to be processed. The target mask image is the mask image obtained after processing the image to be processed by the target zero-shot learning model. The target snow foreground image is the snow foreground image obtained after processing the image to be processed by the target zero-shot learning model.
[0058] In an alternative manner, S1 includes:
[0059] S11. Input any original sample snow image into the preset zero-shot learning model to obtain the first sample mask image, the first sample snow foreground image, and the first sample clear image corresponding to the any original sample snow image, and input them into the snow image model to obtain the reconstructed sample snow image corresponding to the any original sample snow image.
[0060] Figure 3 Shows the overall network structure diagram of this embodiment. As Figure 3As shown in the figure: ① The clear image estimation network (J-Net) includes: eleven convolutional layers connected in sequence. Among them, the first ten convolutional layers are all 3*3 convolutional layers, and the last convolutional layer is a 1*1 convolutional layer. The third convolutional layer is also connected to the eighth convolutional layer, the fourth convolutional layer is also connected to the seventh convolutional layer, and the fifth convolutional layer is also connected to the sixth convolutional layer. The first convolutional layer receives the original sample image, and the eleventh convolutional layer outputs the sample clear image. ② The snow foreground estimation network (S-Net) includes: eleven convolutional layers connected in sequence. Among them, the first ten convolutional layers are all 3*3 convolutional layers, and the last convolutional layer is a 1*1 convolutional layer. The third convolutional layer is also connected to the eighth convolutional layer, the fourth convolutional layer is also connected to the seventh convolutional layer, and the fifth convolutional layer is also connected to the sixth convolutional layer. The first convolutional layer receives the original sample image, and the eleventh convolutional layer outputs the sample snow foreground image. ③ The mask estimation network (z-Net) includes: the first four convolutional layers, a feature representation layer, and the last four convolutional layers connected in sequence. Among them, the first four convolutional layers and the first three of the last four convolutional layers are all 5*5 convolutional layers, and the last convolutional layer is a 1*1 convolutional layer. The feature representation layer first extracts useful features from the original data. These features can be numerical values, text, images, or other forms of data, depending on the type of data and the requirements of the task. The process of feature extraction usually involves various algorithms and techniques, such as principal component analysis (PCA), factor analysis (FA), convolutional neural network (CNN), etc. These algorithms can identify and extract the most important features from the original data.
[0061] It should be noted that the mask estimation network, the snow foreground estimation network, and the clear image estimation network are not completely isolated, and some of the intermediate feature data is shared. The specific structures of the mask estimation network, the snow foreground estimation network, and the clear image estimation network can also be adjusted according to the actual situation, and no restrictions are set here.
[0062] S12. Input the reconstructed sample snow image corresponding to any one of the original sample snow images into the preset zero-shot learning model to obtain the second sample mask image, the second sample snow foreground image, and the second sample clear image corresponding to any one of the original sample snow images.
[0063] Among them, the first sample mask image, the first sample snow foreground image, and the first sample clear image are images generated based on the sample snow image before reconstruction, and the second sample mask image, the second sample snow foreground image, and the second sample clear image are images generated based on the sample snow image after reconstruction. The generation methods before and after reconstruction are the same, and will not be elaborated here.
[0064] S13. According to the first sample mask image, the first sample snowflake foreground image, the first sample clear image, the second sample mask image, the second sample snowflake foreground image, and the second sample clear image corresponding to any one of the original sample snow images, and in combination with the objective loss function of the preset zero-shot learning model, obtain the loss value corresponding to any one of the original sample snow images until the loss value corresponding to each original sample snow image is obtained.
[0065] S14. Use all the loss values to optimize the network parameters of the preset zero-shot learning model to obtain an optimized zero-shot learning model. Take the optimized zero-shot learning model as the preset zero-shot learning model and return to execute S11 until the optimized zero-shot learning model meets the preset iterative training conditions, and then determine the optimized zero-shot learning model as the target zero-shot learning model.
[0066] Among them, the preset iterative training conditions are default set to: reaching the maximum number of iterations or loss convergence, and can also be set according to the actual situation, without limitation here.
[0067] In an alternative way, S11 includes:
[0068] S111. Input any one of the original sample snow images into the mask estimation network, the snowflake foreground estimation network, and the clear image estimation network respectively to obtain the first sample mask image, the first sample snowflake foreground image, and the first sample clear image corresponding to any one of the original sample snow images.
[0069] S112. Input the first sample mask image, the first sample snowflake foreground image, and the first sample clear image corresponding to any one of the original sample snow images into the snow image model to obtain the reconstructed sample snow image corresponding to any one of the original sample snow images.
[0070] Among them, the snow image model is: I2(x) = Z1(x) × S1(x) + (1 - Z1(x)) × J1(x); Z1(x) is the first sample mask image, S1(x) is the first sample snowflake foreground image, J1(x) is the first sample clear image, and I2(x) is the reconstructed sample snow image.
[0071] In an alternative way, the step of obtaining the loss value corresponding to any one of the original sample snow images according to the first sample mask image, the first sample snowflake foreground image, the first sample clear image, the second sample mask image, the second sample snowflake foreground image, and the second sample clear image corresponding to any one of the original sample snow images, and in combination with the objective loss function of the preset zero-shot learning model, includes:
[0072] Input the first sample mask image, the first sample snowflake foreground image, the first sample clear image, the second sample mask image, the second sample snowflake foreground image, and the second sample clear image corresponding to any one of the original sample snow images into the target loss function to obtain the loss value corresponding to any one of the original sample snow images.
[0073] Among them, the target loss function is: is the loss value, is the image similarity loss, is the snowflake foreground image loss, is the image saturation penalty loss, is the mask image loss, and ω1, ω2, ω3, and ω4 are loss coefficients.
[0074] It should be noted that, as Figure 4 shown, specifically represents the similarity loss between the original sample snow image I1(x) and the reconstructed sample snow image I2(x), as well as between the first sample clear image J1(x) and the second sample clear image J2(x). specifically represents the loss between the first sample snowflake foreground image S1(x) and the second sample snowflake foreground image S2(x), specifically represents the loss between the first sample mask image Z1(x) and the second sample mask image Z2(x). specifically represents the loss of the difference in image saturation between the original sample snow image I1(x) and the reconstructed sample snow image I2(x). As Figure 5 shown, the snow image model can also be deformed as: In the inference (detection) stage, the required clear image can be restored according to the mask image, the snowflake foreground image, and the snow image.
[0075] The technical solution of this embodiment performs image desnowing by constructing a lightweight zero-shot learning model, which avoids the problems of difficult-to-obtain data sets and threshold bias problems in training networks with synthetic data sets, and at the same time improves the video image desnowing effect and processing efficiency.
[0076] Figure 6 Fig. shows a structural schematic diagram of an embodiment of an image desnowing system 200 based on zero-shot learning provided by the present invention. As Figure 6 shown, the system 200 includes: a training module 210 and a running module 220;
[0077] The training module 210 is configured to: train a preset zero-shot learning model for image snow removal based on a plurality of original sample snow images in combination with a snow image model, to obtain a target zero-shot learning model; wherein, the preset zero-shot learning model includes: a mask estimation network, a snowflake foreground estimation network, and a clear image estimation network; the snow image model is used to represent the correlation between the mask image extracted by the mask estimation network, the snowflake foreground image extracted by the snowflake foreground estimation network, the clear image extracted by the clear image estimation network, and the snow image.
[0078] The running module 220 is configured to: input the snow image to be processed into the target zero-shot learning model, obtain the target mask image and the target snowflake foreground image corresponding to the snow image to be processed, and input them into the snow image model to obtain the target snow-removed image corresponding to the snow image to be processed.
[0079] In an optional manner, the training module 210 includes: a first training module, a second training module, a third training module, and an iterative training module;
[0080] The first training module is configured to: input any one of the original sample snow images into the preset zero-shot learning model, obtain the first sample mask image, the first sample snowflake foreground image, and the first sample clear image corresponding to the any one of the original sample snow images, and input them into the snow image model to obtain the reconstructed sample snow image corresponding to the any one of the original sample snow images;
[0081] The second training module is configured to: input the reconstructed sample snow image corresponding to the any one of the original sample snow images into the preset zero-shot learning model, obtain the second sample mask image, the second sample snowflake foreground image, and the second sample clear image corresponding to the any one of the original sample snow images;
[0082] The third training module is configured to: according to the first sample mask image, the first sample snowflake foreground image, the first sample clear image, the second sample mask image, the second sample snowflake foreground image, and the second sample clear image corresponding to the any one of the original sample snow images, and in combination with the target loss function of the preset zero-shot learning model, obtain the loss value corresponding to the any one of the original sample snow images until the loss values corresponding to each original sample snow image are obtained;
[0083] The iterative training module is configured to: use all the loss values to optimize the network parameters of the preset zero-shot learning model, obtain the optimized zero-shot learning model, take the optimized zero-shot learning model as the preset zero-shot learning model, and return to call the first training module until the optimized zero-shot learning model meets the preset iterative training condition, and determine the optimized zero-shot learning model as the target zero-shot learning model.
[0084] In an alternative manner, the first training module is specifically configured to:
[0085] Input any one of the original snow images into the mask estimation network, the snowflake foreground estimation network, and the clear image estimation network respectively, to obtain a first sample mask image, a first sample snowflake foreground image, and a first sample clear image corresponding to any one of the original snow images;
[0086] Input the first sample mask image, the first sample snowflake foreground image, and the first sample clear image corresponding to any one of the original snow images into the snow image model to obtain a reconstructed sample snow image corresponding to any one of the original snow images; wherein, the snow image model is: I2(x) = Z1(x) × S1(x) + (1 - Z1(x)) × J1(x); Z1(x) is the first sample mask image, S1(x) is the first sample snowflake foreground image, J1(x) is the first sample clear image, and I2(x) is the reconstructed sample snow image.
[0087] In an alternative manner, the third training module is specifically configured to:
[0088] Input the first sample mask image, the first sample snowflake foreground image, the first sample clear image, the second sample mask image, the second sample snowflake foreground image, and the second sample clear image corresponding to any one of the original snow images into the target loss function to obtain a loss value corresponding to any one of the original snow images; wherein, the target loss function is: is the loss value, is the image similarity loss, is the snowflake foreground image loss, is the image saturation penalty loss, is the mask image loss, ω1, ω2, ω3, and ω4 are loss coefficients.
[0089] In an alternative manner, the mask image is: an image after secondary masking.
[0090] It should be noted that the beneficial effects of the image desnowing system based on zero-shot learning provided in the above embodiments are the same as those of the image desnowing method based on zero-shot learning, and will not be elaborated here. In addition, when the system provided in the above embodiments realizes its functions, only the division of the above function modules is used as an example for illustration. In practical applications, the above functions can be allocated to different function modules according to needs, that is, the system can be divided into different function modules according to the actual situation to complete all or part of the functions described above. In addition, the system provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiments, and will not be elaborated here.
[0091] Among them, the snow removal system for images based on zero-shot learning of the present invention can be a computer program (including program code) running on a computer device. For example, the snow removal system for images based on zero-shot learning of the present invention is an application software, which can be used to execute the corresponding steps in the snow removal method for images based on zero-shot learning of the present invention.
[0092] In some embodiments, the snow removal system for images based on zero-shot learning of the present invention can be implemented in a combination of software and hardware. As an example, the snow removal system for images based on zero-shot learning of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the snow removal method for images based on zero-shot learning of the present invention. For example, the processor in the form of a hardware decoding processor can employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0093] Among them, the modules involved in the embodiments of the present invention can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases.
[0094] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the snow removal method for images based on zero-shot learning as described in any one of the above is implemented. That is to say, an electronic device according to an embodiment of the present invention may include, but is not limited to: a processor and a memory; the memory is used to store the computer program; the processor is used to execute the snow removal method for images based on zero-shot learning shown in any embodiment of the present invention by calling the computer program.
[0095] In an alternative embodiment, an electronic device is provided, as Figure 7 shown, Figure 7The electronic device 4000 shown includes a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as being connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between this electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present invention.
[0096] The processor 4001 can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in connection with the disclosure of the present invention. The processor 4001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0097] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard structure) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only a thick line is used to represent the bus 4002 in the figure, but it does not mean that there is only one bus or one type of bus.
[0098] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0099] The memory 4003 is used to store the application program code (computer program) for executing the solution of the present invention and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the application program code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.
[0100] Among them, the electronic device can also be a terminal device. The terminal device can be any terminal device that can install an application and access a web page through the application, including at least one of a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart TV, and a smart vehicle device.
[0101] It should be noted that Figure 7 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0102] A computer-readable storage medium according to an embodiment of the present invention has a computer program stored thereon. When the computer program is executed by a processor, it implements any one of the above-described zero-shot learning-based image desnowing methods.
[0103] Optionally, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0104] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the above image de-snowing method based on zero-shot learning.
[0105] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0106] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0107] The computer-readable storage medium provided by the embodiments of the present invention may be, but is not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared ray, or semiconductor, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device or component.
[0108] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.
[0109] The above description is only a preferred embodiment of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solution formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present invention.
[0110] It should be noted that the terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, and represent a limitation on a specific order or sequence. Under appropriate circumstances, the order of use of similar objects may be interchanged so that the embodiments of this application described here can be implemented in an order other than the illustrated or described order.
[0111] Those skilled in the art know that the present invention can be implemented as a system, method or computer program product. Therefore, the present invention can be specifically implemented in the following forms: it can be completely hardware, can also be completely software (including firmware, resident software, microcode, etc.), and can also be in the form of a combination of hardware and software, generally referred to as "circuit", "module" or "system" in this article. In addition, in some embodiments, the present invention can also be implemented in the form of a computer program product in one or more computer-readable media, and the computer-readable media contains computer-readable program code.
[0112] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for removing snow from an image based on zero-shot learning, characterized in that: include: S1. Based on multiple original sample snow images and in combination with a snow image model, a preset zero-sample learning model for image snow removal is trained to obtain a target zero-sample learning model; wherein the preset zero-sample learning model includes: a mask estimation network, a snowflake foreground estimation network and a clear image estimation network; the snow image model is used to characterize the association between the mask image extracted by the mask estimation network, the snowflake foreground image extracted by the snowflake foreground estimation network, the clear image extracted by the clear image estimation network and the snow image; S2. Input the snow image to be processed into the target zero-sample learning model, obtain the target mask image and target snowflake foreground image corresponding to the snow image to be processed, and input them into the snow image model to obtain the target snow-removed image corresponding to the snow image to be processed.
2. The image snow removal method based on zero-shot learning according to claim 1, characterized in that: S1 includes: S11, inputting any original sample snow image into the preset zero-sample learning model, obtaining a first sample mask image, a first sample snowflake foreground image and a first sample clear image corresponding to any original sample snow image, and inputting them into the snow image model, to obtain a reconstructed sample snow image corresponding to any original sample snow image; S12, inputting the reconstructed sample snow image corresponding to any one of the original sample snow images into the preset zero-sample learning model to obtain a second sample mask image, a second sample snowflake foreground image and a second sample clear image corresponding to any one of the original sample snow images; S13, according to the first sample mask image, the first sample snowflake foreground image, the first sample clear image, the second sample mask image, the second sample snowflake foreground image and the second sample clear image corresponding to any of the original sample snow images, and in combination with the target loss function of the preset zero-sample learning model, obtaining the loss value corresponding to any of the original sample snow images, until the loss value corresponding to each original sample snow image is obtained; S14. Utilize all loss values to optimize the network parameters of the preset zero-sample learning model to obtain an optimized zero-sample learning model, use the optimized zero-sample learning model as the preset zero-sample learning model, and return to execute S11 until the optimized zero-sample learning model meets the preset iterative training conditions, and then determine the optimized zero-sample learning model as the target zero-sample learning model.
3. The image snow removal method based on zero-shot learning according to claim 2 is characterized in that: S11 includes: S111, inputting any one of the original sample snow images into the mask estimation network, the snowflake foreground estimation network and the clear image estimation network respectively, to obtain a first sample mask image, a first sample snowflake foreground image and a first sample clear image corresponding to any one of the original sample snow images; S112. Input the first sample mask image, the first sample snowflake foreground image and the first sample clear image corresponding to any of the original sample snow images into the snow image model to obtain a reconstructed sample snow image corresponding to any of the original sample snow images; wherein the snow image model is: I2(x)=Z1(x)×S1(x)+(1-Z1(x))×J1(x); Z1(x) is the first sample mask image, S1(x) is the first sample snowflake foreground image, J1(x) is the first sample clear image, and I2(x) is the reconstructed sample snow image.
4. The image snow removal method based on zero-shot learning according to claim 2, characterized in that: The step of obtaining a loss value corresponding to any original sample snow image according to the first sample mask image, the first sample snowflake foreground image, the first sample clear image, the second sample mask image, the second sample snowflake foreground image and the second sample clear image corresponding to any original sample snow image, and in combination with the target loss function of the preset zero-sample learning model, comprises: The first sample mask image, the first sample snowflake foreground image, the first sample clear image, the second sample mask image, the second sample snowflake foreground image and the second sample clear image corresponding to any of the original sample snow images are input into the target loss function to obtain the loss value corresponding to any of the original sample snow images; wherein the target loss function is: is the loss value, is the image similarity loss, is the snow foreground image loss, Penalize loss for image saturation, is the mask image loss, ω1, ω2, ω3 and ω4 are the loss coefficients.
5. The image snow removal method based on zero-shot learning according to any one of claims 1 to 4, characterized in that: The mask image is: an image after secondary masking.
6. An image snow removal system based on zero-shot learning, characterized in that: include: Training module and running module; The training module is used to: train a preset zero-sample learning model for image snow removal based on multiple original sample snow images and in combination with a snow image model to obtain a target zero-sample learning model; wherein the preset zero-sample learning model includes: a mask estimation network, a snowflake foreground estimation network and a clear image estimation network; the snow image model is used to characterize the association between the mask image extracted by the mask estimation network, the snowflake foreground image extracted by the snowflake foreground estimation network, the clear image extracted by the clear image estimation network and the snow image; The operation module is used to: input the snow image to be processed into the target zero-sample learning model, obtain the target mask image and target snowflake foreground image corresponding to the snow image to be processed and input them into the snow image model to obtain the target snow-removed image corresponding to the snow image to be processed.
7. The image snow removal system based on zero-shot learning according to claim 6, characterized in that: The training modules include: a first training module, a second training module, a third training module and an iterative training module; The first training module is used to: input any original sample snow image into the preset zero-sample learning model, obtain a first sample mask image, a first sample snowflake foreground image and a first sample clear image corresponding to the any original sample snow image, and input them into the snow image model to obtain a reconstructed sample snow image corresponding to the any original sample snow image; The second training module is used to: input the reconstructed sample snow image corresponding to any of the original sample snow images into the preset zero-sample learning model to obtain a second sample mask image, a second sample snowflake foreground image and a second sample clear image corresponding to any of the original sample snow images; The third training module is used to obtain the loss value corresponding to any original sample snow image according to the first sample mask image, the first sample snowflake foreground image, the first sample clear image, the second sample mask image, the second sample snowflake foreground image and the second sample clear image corresponding to any original sample snow image, and in combination with the target loss function of the preset zero-sample learning model, until the loss value corresponding to each original sample snow image is obtained; The iterative training module is used to: use all loss values to optimize the network parameters of the preset zero-sample learning model to obtain an optimized zero-sample learning model, use the optimized zero-sample learning model as the preset zero-sample learning model, and return to call the first training module until the optimized zero-sample learning model meets the preset iterative training conditions, and then determine the optimized zero-sample learning model as the target zero-sample learning model.
8. The image snow removal system based on zero-shot learning according to claim 7, characterized in that: The first training module is specifically used for: Inputting any one of the original sample snow images into the mask estimation network, the snowflake foreground estimation network and the clear image estimation network respectively, to obtain a first sample mask image, a first sample snowflake foreground image and a first sample clear image corresponding to any one of the original sample snow images; The first sample mask image, the first sample snowflake foreground image and the first sample clear image corresponding to any of the original sample snow images are input into the snow image model to obtain a reconstructed sample snow image corresponding to any of the original sample snow images; wherein the snow image model is: I2(x)=Z1(x)×S1(x)+(1-Z1(x))×J1(x); Z1(x) is the first sample mask image, S1(x) is the first sample snowflake foreground image, J1(x) is the first sample clear image, and I2(x) is the reconstructed sample snow image.
9. An electronic device, characterized in that: The electronic device includes a processor, the processor is coupled to a memory, at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor so that the electronic device implements the image snow removal method based on zero-sample learning as described in any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor so that the computer-readable storage medium implements the image snow removal method based on zero-sample learning as described in any one of claims 1 to 5.