Marine rainy day image processing method and device
By augmenting the images on the sea rainy day using pre-trained augmentation model, the problem of difficult to obtain ocean image acquisition in severe weather is solved, and a high-accurate image recognition effect is achieved.
Patent Information
- Application Number
- CN202411884753.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-23
AI Technical Summary
In severe weather, it is difficult to obtain ocean image collection, resulting in insufficient accuracy of image content recognition.
The pre-trained augmentation model is used to augment the sea rainy images through the pix2pix algorithm to generate the target image corresponding to the original image, thereby supplementing the lost data and improving the accuracy of image recognition.
Through image augmentation processing, the lack of image data in bad weather is supplemented and the accuracy of sea image content recognition is improved.
Smart Images

Figure CN120032201A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image processing, and in particular relates to a method and a device for processing images on rainy days at sea. Background Art
[0002] At present, with the continuous development of network technology, when ships are operating at sea, they need to collect images under different weather conditions at sea. However, in some severe weather conditions, the information of the collected marine images cannot be obtained. Therefore, how to improve the accuracy of content recognition of marine images is an urgent problem to be solved. Summary of the invention
[0003] The purpose of the present invention is to provide a method and device for processing rainy day images at sea.
[0004] According to a first aspect of the present invention, there is provided a method for processing rainy sea images, comprising:
[0005] Acquire rainy day images at sea;
[0006] The rainy sea image is augmented according to a pre-trained augmentation model to obtain an augmented target image corresponding to the rainy sea image, wherein the pre-trained augmentation model is obtained by training and processing pix2pix using sample data.
[0007] Optionally, the pre-trained augmented model is a generative adversarial network model, including a generator and a discriminator; the generator is used to generate a target image, the discriminator is used to determine whether the generated image is real, Z is the input noise signal, and Fake-B is a fake target category image generated by the generator.
[0008] Optionally, the method further comprises:
[0009] When generating the second image from the first image, the first image is first passed through the generator G to obtain the generated image, the second fake image, and the generator parameters are adjusted according to the discrimination result of the discriminator to improve the generation effect of the generator;
[0010] The first image is used as a label and is spliced with the target second image and the second fake image as positive and negative samples to train the discriminator, thereby improving the discriminative effect of the discriminator;
[0011] After multiple trainings, both improve together, so that a second fake picture corresponding to the second picture is generated by the first picture.
[0012] Optionally, the method further comprises:
[0013] For the generator part, the pix2pix algorithm adopts the U-Net structure, extracts the features of the image through multiple convolutional layers, and downsamples to reduce the size of the feature map and increase the receptive field. In the decoding process, in addition to upsampling to restore the size of the feature map, the corresponding feature map in the downsampling process is also copied and spliced, so that the network reduces the loss of image detail information caused by the downsampling process.
[0014] According to a second aspect of the present invention, there is provided a device for processing rainy images at sea, comprising:
[0015] An acquisition module is used to acquire images of rainy days at sea;
[0016] The processing module is used to perform augmentation processing on the rainy sea image according to a pre-trained augmentation model to obtain a target image corresponding to the rainy sea image after augmentation processing, wherein the pre-trained augmentation model is obtained by training and processing pix2pix using sample data.
[0017] Optionally, the pre-trained augmented model is a generative adversarial network model, including a generator and a discriminator; the generator is used to generate a target image, the discriminator is used to determine whether the generated image is real, Z is the input noise signal, and Fake-B is a fake target category image generated by the generator.
[0018] Optionally, the processing module is used to:
[0019] When generating the second image from the first image, the first image is first passed through the generator G to obtain the generated image, the second fake image, and the generator parameters are adjusted according to the discrimination result of the discriminator to improve the generation effect of the generator;
[0020] The first image is used as a label and is spliced with the target second image and the second fake image as positive and negative samples to train the discriminator, thereby improving the discriminative effect of the discriminator;
[0021] After multiple trainings, both improve together, so that a second fake picture corresponding to the second picture is generated by the first picture.
[0022] Optionally, the processing module is used to:
[0023] For the generator part, the pix2pix algorithm adopts the U-Net structure, extracts the features of the image through multiple convolutional layers, and downsamples to reduce the size of the feature map and increase the receptive field. In the decoding process, in addition to upsampling to restore the size of the feature map, the corresponding feature map in the downsampling process is also copied and spliced, so that the network reduces the loss of image detail information caused by the downsampling process.
[0024] In a third aspect, the present application shows an electronic device, which includes: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in any of the above aspects.
[0025] In a fourth aspect, the present application illustrates a non-temporary computer-readable storage medium, which, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute a method as described in any of the above aspects.
[0026] In a fifth aspect, the present application illustrates a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in any of the above aspects.
[0027] The beneficial effects brought by the present invention are as follows:
[0028] It can be seen from the above scheme that the embodiment of the present invention provides a method and device for processing rainy sea images, by acquiring a rainy sea image; performing augmentation processing on the rainy sea image according to a pre-trained augmentation model, and obtaining a target image corresponding to the rainy sea image after augmentation processing, wherein the pre-trained augmentation model is obtained by training and processing pix2pix using sample data, so that image augmentation is achieved through pix2pix, the lost data is supplemented, and high image recognition accuracy is provided. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A schematic diagram of a flow chart of a method for processing rainy sea images according to an embodiment;
[0030] Figure 2 A schematic flow chart of another method for processing rainy sea images according to an embodiment;
[0031] Figure 3 It is a structural block diagram of a rainy day image processing device at sea of the present application;
[0032] Figure 4 is a block diagram of an electronic device of the present application;
[0033] Figure 5 It is a block diagram of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0035] Reference Figure 1 , shows a flowchart of a method for processing rainy sea images of the present application, which can be applied to electronic devices, wherein the method may specifically include the following steps:
[0036] S101, acquiring an image of a rainy day at sea;
[0037] S102. Perform augmentation processing on the rainy sea image according to a pre-trained augmentation model to obtain an augmented target image corresponding to the rainy sea image, wherein the pre-trained augmentation model is obtained by training and processing pix2pix using sample data.
[0038] Generative adversarial network (GAN) is an image generation algorithm proposed by Ian Goodfellow et al. in 2014. It is also the most mainstream image generation algorithm at present. The introduction of this method makes it possible for computers to create by themselves. Rather than saying that the generative adversarial network is a network structure, it is better to say that it is a new training method. Unlike the traditional neural network, which directly calculates the loss between the network result and the target value, and then uses the loss to guide the neural network to perform gradient descent, the generative adversarial network contains two parts of the network structure, called the generator and the discriminator. The two adopt a zero-sum game mode, compete with each other, and jointly improve the accuracy of the network.
[0039] Pix2pix is an image translation algorithm based on conditional generative adversarial networks (CGAN), which can directly generate images from one domain to another domain by training paired dataset images.
[0040] Another embodiment of the present application further supplements the method for processing rainy sea images provided in the above embodiment.
[0041] Optionally, the pre-trained augmented model is a generative adversarial network model, including a generator and a discriminator; the generator is used to generate a target image, the discriminator is used to determine whether the generated image is real, Z is an input noise signal, and Fake-B is a fake target category image generated by the generator.
[0042] Optionally, the method further comprises:
[0043] When generating the second image from the first image, the first image is first passed through the generator G to obtain the generated image, the second fake image, and the generator parameters are adjusted according to the discrimination result of the discriminator to improve the generation effect of the generator;
[0044] The first image is used as a label and is spliced with the target second image and the second fake image as positive and negative samples to train the discriminator, thereby improving the discriminative effect of the discriminator;
[0045] After multiple trainings, both improve together, so that a second fake picture corresponding to the second picture is generated by the first picture.
[0046] Specifically, the generative adversarial network is generally used to generate images of the desired category from noise signals. G is the generator, which is used to generate the target image. D is the discriminator, which is used to determine whether the generated image is real. Z is the input noise signal, and Fake-B is a fake target category image generated by the generator. Although the generator and discriminator in Ian Goodfellow's original paper use a fully connected structure, the generative adversarial network is more concerned about its training process. In fact, the generator and discriminator can use any structure. With the development of the generative adversarial network series of algorithms in recent years and the superior performance of convolutional neural networks in image processing, almost all generators and discriminators in this series of algorithms for image generation use convolutional neural networks.
[0047] like Figure 2 As shown in the figure, before the training of the generative adversarial network, the generator and the discriminator do not have good generation and discrimination capabilities. Random noise Z is input into the generator to try to generate a fake sample B, namely Fake-B. Then Fake-B is input into the discriminator to let the discriminator determine whether it is real, and update the parameters of the generator according to the loss of the discrimination result. Since the parameters of the generator and the discriminator are randomly initialized at the beginning, there is almost no training effect in the first step, which is only to obtain the image Fake-B; next, the generated fake image Fake-B and the real image B in the training set are input into the discriminator respectively, and the discriminator is allowed to perform a binary discrimination of "true" and "false" to distinguish whether it is a real image in the data set or a generated fake image, and the parameters of the discriminator are updated according to the discrimination results to optimize the performance of the discriminator so that it has a certain discrimination ability;
[0048] At this time, starting from the second cycle, the discriminator will have a certain ability to distinguish between real images and fake images. The Fake-B generated by the generator is input into the discriminator, and its binary classification result is compared with the "real" to calculate the loss and let the loss result guide the training of the generator. The parameters of the generator can be optimized to make its generation result closer to the "real" direction and improve its generation effect. In the subsequent training process, the improvement of the generator's generation effect can improve its discrimination effect during the training of the discriminator, and the improvement of the discriminator's discrimination effect can improve its generation effect during the training of the generator. The two are trained alternately and make progress together. After multiple cycles, a generator that can generate target category images can be obtained.
[0049] Similar to some common data distributions, such as normal distribution and uniform distribution, the storage form of images in computers is a matrix, which can also be regarded as data that obeys a certain distribution. Images of the same category often have similar features, so they can be regarded as data that obey the same distribution. The purpose of the generative adversarial network is to enable computers to learn to create, that is, to autonomously generate images of a certain category that are not in the data set. It can be regarded as generating data that obeys the same distribution as the images of this category in the data set. If the traditional deep learning training method is used to directly update the parameters of the generator by comparing the generator's generation results with the images in the data set to calculate the loss, then the generated images will be compared with each pixel of each image in the data set. Due to the diversity of samples in the data set, the content learned by the generator will be different each time, and effective gradient descent cannot be performed; while in the generative adversarial network, the discriminator is used to determine whether the output of the generator conforms to the target distribution, so that the gradient descent direction of the generator is always in the direction where the output is more in line with the target distribution, so that the network parameters can be continuously optimized, and finally the images of the same category as the samples are generated.
[0050] Optionally, the method further comprises:
[0051] For the generator part, the pix2pix algorithm adopts the U-Net structure, extracts the features of the image through multiple convolutional layers, and downsamples to reduce the size of the feature map and increase the receptive field. In the decoding process, in addition to upsampling to restore the size of the feature map, the corresponding feature map in the downsampling process is also copied and spliced, so that the network reduces the loss of image detail information caused by the downsampling process.
[0052] In the earliest generative adversarial networks, random noise signals can be used to randomly generate images of a certain category, but the generated images are random and uncontrolled. For example, when generating images of handwritten digital datasets, various numbers are included in the output, and it is impossible to control it to generate the desired number. Based on this, the conditional generative adversarial network (CGAN) was proposed, which achieves the purpose of controlling the output category of the generator by adding constraints to the input of the generator and the discriminator.
[0053] Where Z is the noise signal, B is the image in the dataset, Fake-B is the image generated by the generator, and A is the label of image B. Label A can be in any form, such as onehot encoding, image, etc. During each training, the label is concatenated with the input noise signal, the input real image, and the fake image generated by the generator, and then passed to the generator or discriminator. At this time, its loss function becomes as follows.
[0054]
[0055] The probability expressions of both the generator and the discriminator have become conditional probability formulas, that is, they are only valid when specific labels are input. In this way, after the training is completed, the generator G is equivalent to a collection of generators of multiple categories. By controlling the input label, it can switch to generators of different categories, thereby generating pictures of different categories.
[0056] pix2pix is an image translation algorithm based on conditional generative adversarial network (CGAN). It can directly generate images from one domain to another domain by training paired dataset images. And get clearer results. Its training process is as follows Figure 2 As shown:
[0057] A (first picture) and B (second picture) are pictures belonging to two domains in the paired data set, respectively, and Fake-B (second fake picture) is a fake picture belonging to domain B generated from a picture in domain A. When the pix2pix algorithm generates picture B from picture A, picture A first passes through generator G to obtain the generated picture Fake-B, and adjusts the generator parameters according to the discriminant's discrimination result to improve the generator's generation effect; then picture A is used as a label to be spliced with target picture B and Fake-B respectively as positive and negative samples to train the discriminator and improve the discriminant's discrimination effect. After multiple trainings, the two make progress together, so that finally picture A can generate picture Fake-B that is realistic enough with picture B.
[0058] The loss function of the Pix2pix algorithm combines the traditional L1-Loss and CGAN_loss. L1-Loss is used to reconstruct the low-frequency part of the image, and CGAN is used to reconstruct the high-frequency part of the image. The specific loss function is shown below.
[0059]
[0060] L CGAN (G,D)=E x,y [log D(x,y)]+E x,z [log(1-D(x,G(x,z)))]
[0061] L L1 (G) = E x,y,z [||yG(x,z)|| 1 ]
[0062] In the loss function formula, x represents the input image of the generator, that is, image A, y represents the target image to be generated, that is, image B, and z represents noise. Noise can be added to the input of the generator and concatenated with x as the input of the generator to increase the diversity of the output, but not adding it will not affect the experimental results.
[0063] For the generator part, the pix2pix algorithm does not use ordinary encoders and decoders, but instead uses the U-Net structure, as shown in the figure below. The encoding process is a simple convolution downsampling process, which extracts the features of the image through multiple convolutional layers, and downsampling to reduce the size of the feature map and increase the receptive field. In the decoding process, in addition to upsampling to restore the size of the feature map, the corresponding feature map in the downsampling process is also copied and spliced, so that the network reduces the loss of image detail information caused by the downsampling process.
[0064] Since only GAN is used to reconstruct the high-frequency part in the algorithm, the original image can be divided into N×N small blocks, each of which represents a receptive field, and the loss is calculated for each receptive field. That is, the algorithm's discriminator uses Patch-D, and the network structure is four layers of convolution downsampling with a step size of 2 and a convolution layer with a step size of 1. The last layer is no longer fully connected, but remains as a convolution layer, outputting an N×N feature map, and calculating the loss with the label matrix of the same size. This method can make the output image have higher clarity and better detail preservation. In addition, since the algorithm is based on the conditional generative adversarial network (CGAN), the images A and B are spliced in the first layer, which doubles the number of input channels.
[0065] The embodiment of the present invention provides a method for processing rainy sea images, which includes acquiring a rainy sea image; performing augmentation processing on the rainy sea image according to a pre-trained augmentation model, and obtaining a target image corresponding to the rainy sea image after the augmentation processing, wherein the pre-trained augmentation model is obtained by training and processing pix2pix using sample data, so that image augmentation is achieved through pix2pix, the lost data is supplemented, and high image recognition accuracy is provided.
[0066] It should be noted that each implementable method in this embodiment may be implemented separately, or may be implemented in combination in any combination without conflict, and this application is not limited thereto.
[0067] Another embodiment of the present application provides a rainy day image processing device at sea, which is used to execute the rainy day image processing method at sea provided by the above embodiment.
[0068] like Figure 3 FIG. 3 is a schematic diagram of the structure of a rainy day image processing device for sea provided in an embodiment of the present application. The device includes an acquisition module 301 and a processing module 302, wherein:
[0069] The acquisition module 301 is used to acquire rainy day images at sea;
[0070] The processing module 302 is used to perform augmentation processing on the rainy sea image according to a pre-trained augmentation model to obtain an augmented target image corresponding to the rainy sea image, wherein the pre-trained augmentation model is obtained by training and processing pix2pix using sample data.
[0071] Regarding the device in this embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0072] Another embodiment of the present application further supplements the description of the rainy day image processing device at sea provided in the above embodiment.
[0073] Optionally, the pre-trained augmented model is a generative adversarial network model, including a generator and a discriminator; the generator is used to generate a target image, the discriminator is used to determine whether the generated image is real, Z is an input noise signal, and Fake-B is a fake target category image generated by the generator.
[0074] Optionally, the processing module is used to:
[0075] When generating the second image from the first image, the first image is first passed through the generator G to obtain the generated image, the second fake image, and the generator parameters are adjusted according to the discrimination result of the discriminator to improve the generation effect of the generator;
[0076] The first image is used as a label and is spliced with the target second image and the second fake image as positive and negative samples to train the discriminator, thereby improving the discriminative effect of the discriminator;
[0077] After multiple trainings, both improve together, so that a second fake picture corresponding to the second picture is generated by the first picture.
[0078] Optionally, the processing module is used to:
[0079] For the generator part, the pix2pix algorithm adopts the U-Net structure, extracts the features of the image through multiple convolutional layers, and downsamples to reduce the size of the feature map and increase the receptive field. In the decoding process, in addition to upsampling to restore the size of the feature map, the corresponding feature map in the downsampling process is also copied and spliced, so that the network reduces the loss of image detail information caused by the downsampling process.
[0080] The embodiment of the present invention provides a method and device for processing rainy sea images, which obtain rainy sea images; perform augmentation processing on the rainy sea images according to a pre-trained augmentation model to obtain a target image corresponding to the rainy sea image after the augmentation processing, wherein the pre-trained augmentation model is obtained by training and processing pix2pix using sample data, so that image augmentation is achieved through pix2pix, the lost data is supplemented, and high image recognition accuracy is provided.
[0081] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0082] Optionally, an embodiment of the present application further provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0083] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0084] Figure 4 800 is a block diagram of an electronic device 800 shown in the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0085] Reference Figure 4, the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .
[0086] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0087] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0088] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.
[0089] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0090] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0091] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0092] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800, and the sensor assembly 814 can also detect the position change of the electronic device 800 or a component of the electronic device 800, the presence or absence of contact between the user and the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0093] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0094] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0095] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by a processor 820 of an electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0096] Figure 5 19 is a block diagram of a computer-readable storage medium 1900 shown in the present application. For example, the computer-readable storage medium 1900 may be provided as a server.
[0097] Reference Figure 5 , the computer-readable storage medium 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0098] The computer readable storage medium 1900 may also include a power supply component 1926 configured to perform power management of the computer readable storage medium 1900, a wired or wireless network interface 1950 configured to connect the computer readable storage medium 1900 to a network, and an input / output (I / O) interface 1958. The computer readable storage medium 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™ or the like.
[0099] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0100] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0101] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
[0102] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0103] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0104] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0105] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0106] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0107] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks, or optical disks.
[0108] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0109] The above are preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for processing rainy sea images, characterized in that: include: Acquire rainy day images at sea; The rainy sea image is augmented according to a pre-trained augmentation model to obtain an augmented target image corresponding to the rainy sea image, wherein the pre-trained augmentation model is obtained by training and processing pix2pix using sample data.
2. The method for processing rainy sea images according to claim 1, characterized in that: The pre-trained augmented model is a generative adversarial network model, including a generator and a discriminator; the generator is used to generate a target image, and the discriminator is used to determine whether the generated image is real. Z is the input noise signal, and Fake-B is a fake target category image generated by the generator.
3. The method for processing rainy sea images according to claim 2, characterized in that: The method further comprises: When generating the second image from the first image, the first image is first passed through the generator G to obtain the generated image, the second fake image, and the generator parameters are adjusted according to the discrimination result of the discriminator to improve the generation effect of the generator; The first image is used as a label and is spliced with the target second image and the second fake image as positive and negative samples to train the discriminator, thereby improving the discriminative effect of the discriminator; After multiple trainings, both improve together, so that a second fake picture corresponding to the second picture is generated by the first picture.
4. The method for processing rainy sea images according to claim 3, characterized in that: The method further comprises: For the generator part, the pix2pix algorithm adopts the U-Net structure, extracts the features of the image through multiple convolutional layers, and downsamples to reduce the size of the feature map and increase the receptive field. In the decoding process, in addition to upsampling to restore the size of the feature map, the corresponding feature map in the downsampling process is also copied and spliced, so that the network reduces the loss of image detail information caused by the downsampling process.
5. A rainy day image processing device at sea, characterized in that: include: An acquisition module is used to acquire images of rainy days at sea; The processing module is used to perform augmentation processing on the rainy sea image according to a pre-trained augmentation model to obtain a target image corresponding to the rainy sea image after augmentation processing, wherein the pre-trained augmentation model is obtained by training and processing pix2pix using sample data.
6. The rainy day image processing device at sea according to claim 5, characterized in that: The pre-trained augmented model is a generative adversarial network model, including a generator and a discriminator; the generator is used to generate a target image, and the discriminator is used to determine whether the generated image is real. Z is the input noise signal, and Fake-B is a fake target category image generated by the generator.
7. The rainy day image processing device at sea according to claim 6, characterized in that: The processing module is used for: When generating the second image from the first image, the first image is first passed through the generator G to obtain the generated image, the second fake image, and the generator parameters are adjusted according to the discrimination result of the discriminator to improve the generation effect of the generator; The first image is used as a label and is spliced with the target second image and the second fake image as positive and negative samples to train the discriminator, thereby improving the discriminative effect of the discriminator; After multiple trainings, both improve together, so that a second fake picture corresponding to the second picture is generated by the first picture.
8. The rainy day image processing device at sea according to claim 7, characterized in that: The processing module is used for: For the generator part, the pix2pix algorithm adopts the U-Net structure, extracts the features of the image through multiple convolutional layers, and downsamples to reduce the size of the feature map and increase the receptive field. In the decoding process, in addition to upsampling to restore the size of the feature map, the corresponding feature map in the downsampling process is also copied and spliced, so that the network reduces the loss of image detail information caused by the downsampling process.
9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 4 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.