Image expansion method based on deep learning
By combining VAE and CGAN deep learning methods in border and coastal defense monitoring, realistic images are generated, and the problem of target tracking is solved, and the problem of discontinuity and low accuracy is achieved, which is higher accuracy and real-time, and is suitable for real-time data processing in border and coastal defense field.
Patent Information
- Application Number
- CN202411766769.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-05-13
AI Technical Summary
Existing target tracking has problems of discontinuity and low accuracy in border and coastal defense monitoring.
A deep learning-based image augmentation method is adopted to form a hybrid model of generated images by constructing a variational autoencoder (VAE) and a conditional generation adversarial network (CGAN), which is trained and optimized to generate realistic images.
It improves the accuracy and real-timeness of target tracking, enhances the quality and stability of images, has good scalability and flexibility, and is suitable for real-time data processing in the border and coastal defense field.
Smart Images

Figure CN119991457A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of border and coastal defense monitoring target tracking, and in particular to the application and optimization of an image expansion algorithm in this field. Background Art
[0002] Border and coastal defense monitoring monitors and manages coastlines, borders, and sea areas, aiming to detect and respond to security threats such as illegal border crossings, smuggling, and terrorist activities, and to safeguard national border security and maritime rights and interests. Modern border and coastal defense monitoring technology involves multiple fields, including radar monitoring, satellite monitoring, drone patrols, and video monitoring. The amount of data obtained by the border and coastal defense monitoring system is limited, and the data lacks diversity, which is not conducive to establishing an accurate and reliable monitoring model.
[0003] Image augmentation is a key technology. By transforming and processing existing image data, the diversity of training data sets can be increased, thereby improving the generalization ability of machine learning models. By improving the accuracy and efficiency of the monitoring system, the investment cost of human and material resources can be reduced, and the efficiency and economy of border and coastal defense monitoring can be improved. Image augmentation technology has important problem-solving and significance in the field of border and coastal defense, which helps to improve the capacity and efficiency of border and coastal defense monitoring systems and ensure the security of national borders and seas. Summary of the invention
[0004] The present invention proposes an image expansion method based on deep learning to solve the problems of discontinuity and low accuracy of existing target tracking.
[0005] The present invention is achieved through the following technical solutions.
[0006] A deep learning-based image expansion method comprises the following steps:
[0007] Step 1: Form a hybrid model for generating images by constructing a VAE model and a CGan model; specifically:
[0008] 2.1 Building a VAE model: The encoder receives the input image and maps it to the mean and variance vectors in the latent space. The decoder receives samples from the latent space and decodes them into the generated image.
[0009] 2.2 Constructing the CGan model: The generator receives random noise vectors and condition information as input, and converts the input into an image that matches the condition through a deconvolutional neural network; the discriminator receives the image and condition information as input, and classifies the input into a real image / fake image through a convolutional neural network, and matches the condition; the loss function The loss of the generator is determined by the probability that the discriminator misjudges the generated image, that is, the generator hopes that the generated image will be regarded as a real image by the discriminator;
[0010] 2.3 Combine the VAE model and the CGAN model to form a hybrid model for generating images, where the encoder generates latent vectors by learning the latent representation of the image; the generator receives these latent vectors and generates realistic images, and the discriminator tries to distinguish between real images and generated images;
[0011] Step 2: training and optimizing the hybrid model, specifically including:
[0012] 3.1 Initialize the parameters of the generator and discriminator;
[0013] 3.2 Iterative training: Randomly extract batches of image samples from the dataset, the generator generates fake images that match the batches of image samples, the discriminator receives real image samples and fake image samples generated by the generator, and classifies them, calculates the loss functions of the generator and the discriminator, and updates the parameters to minimize the loss; repeat the above steps until the predetermined number of training rounds is reached or the loss converges to a stable state;
[0014] 3.3 Optimize the trained model: Use the optimizer to update the model parameters and adjust the learning rate and other hyperparameters to improve the efficiency and stability of training.
[0015] Beneficial effects of the present invention:
[0016] 1. Accuracy: The present invention integrates VAE and cGAN so that the generated ship images can accurately reflect the input condition information, thereby meeting actual needs and improving recognition accuracy;
[0017] 2. Real-time: The efficient training and generation speed of the present invention ensures a rapid response to real-time data, and is suitable for scenarios requiring immediate processing in the field of border and coastal defense;
[0018] 3. Multi-method fusion: This paper takes advantage of the respective advantages of VAE and CGAN, such as VAE's learning of data potential representation and cGAN's realism in generating images, thus improving the quality and stability of generated images;
[0019] 4. Flexibility: The present invention can flexibly adjust and combine different deep learning modules, and customize and optimize according to specific scenarios and needs. By integrating multiple algorithms, it can make up for the shortcomings of a single algorithm, improve the robustness and stability of the overall algorithm, and enable it to have a good processing effect on data in various situations;
[0020] 5. Scalability: The present invention has good scalability and can adjust the network structure and increase training data as needed to cope with different goals and environments. The trained model can be applied to new image expansion tasks through transfer learning to accelerate the deployment and application of the model;
[0021] 6. Practical applicability: To ensure the feasibility and performance of the algorithm in real-time applications, the present invention optimizes the algorithm and selects a suitable hardware platform, designs an effective data flow pipeline and real-time feedback mechanism, integrates the model into specific scenarios and customizes it, establishes a real-time monitoring and maintenance mechanism, and continuously performs performance evaluation and optimization iterations, thereby improving the system's real-time processing capabilities and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a flow chart of an image expansion method based on deep learning in the present invention. DETAILED DESCRIPTION
[0023] The exemplary embodiments of the present invention are described in detail below with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the accompanying drawings are only exemplary and are intended to illustrate the principles and spirit of the present invention, rather than to limit the scope of the present invention.
[0024] like Figure 1 As shown, a deep learning-based image expansion method of the present invention specifically includes the following steps:
[0025] Step 1: Data collection and preprocessing; specifically:
[0026] Use a dataset or collect image data for target tracking, resize the image to a uniform size, use a denoising algorithm to reduce noise in the image, and convert the image from RGB to other color spaces (in specific implementation, the color space can be YUV or LAB, which can simplify the complexity of model input), reduce the impact of raindrops according to the characteristics of different channels, and finally standardize the image. This method can ensure the consistency of data in model training;
[0027] In a specific implementation, the image data may be a video stream or a separate image frame, and the data should cover various scenes and possible target objects to ensure the generalization ability of the model.
[0028] In this embodiment, the denoising algorithm adopts a mean filtering algorithm or a median filtering algorithm, which can improve the quality of subsequent feature extraction and model training.
[0029] Step 2: Form a hybrid model for generating images by constructing a VAE model and a CGan model; specifically:
[0030] 2.1 Building a VAE model: The encoder receives the input image and maps it to the mean and variance vectors in the latent space. The decoder receives samples from the latent space and decodes them into the generated image.
[0031] Among them, VAE is a generative model that reconstructs the input image by introducing distribution in the latent space. The core of VAE is that the encoder outputs not specific latent features but feature distribution, thus having generative capabilities. The model ensures the continuity and integrity of the latent space by regularizing the latent space. Based on this, the VAE model consists of two parts: the encoder and the decoder:
[0032] 1. Encoder: compresses the input data into a representation of a latent space (this representation is a probability distribution rather than a specific point), and the encoder outputs the mean and variance of the distribution; in this embodiment, a convolutional neural network (CNN) or a fully connected neural network is used as an encoder to learn the representation of the image.
[0033] 2. Decoder: Sample from the probability distribution output by the encoder to generate new data instances, and finally generate realistic images similar to the input images. In this embodiment, a deconvolutional neural network or a fully connected neural network is used as a decoder to generate images.
[0034] 2.2 Constructing the CGan model: The generator receives random noise vectors and condition information as input, and converts the input into an image that matches the condition through a deconvolutional neural network; the discriminator receives the image and condition information as input, and classifies the input into a real image / fake image through a convolutional neural network, and matches the condition; the loss function (Loss Function) The loss of the generator is determined by the probability that the discriminator misjudges the generated image, that is, the generator hopes that the generated image will be regarded as a real image by the discriminator;
[0035] In this embodiment, the loss of the discriminator is determined by the probability of correctly classifying real images and generated images, that is, the discriminator hopes to correctly distinguish between real images and generated images.
[0036] In this embodiment, the loss function includes reconstruction loss and KL divergence loss, wherein the reconstruction loss is used to measure the difference between the generated image and the original image, and the KL divergence loss is used to penalize the difference between the latent space distribution and the standard normal distribution; this approach helps to learn continuous and smooth latent representations.
[0037] In this embodiment, the condition information is a category label or a text description, which is used to guide the generator to generate images of a specific category.
[0038] 2.3 Combine the VAE model and the CGAN model to form a hybrid model for generating images, in which the encoder generates latent vectors by learning the potential representation of the image; the generator receives these latent vectors and generates realistic images, and the discriminator tries to distinguish between real images and generated images.
[0039] In terms of architecture, the VAE encoder outputs a latent vector, which is directly used as the input of the CGAN generator. In this way, the generator can use the latent features extracted by the VAE to generate images. If it is a conditional model, the conditional information can be processed together with the input image in the encoding stage so that the latent vector contains the conditional features. The generator receives the latent vector and the same conditional information to generate images that meet the conditions.
[0040] During the training process, a loss function combining the two is designed. The generator loss includes the reconstruction loss of VAE and the adversarial loss of CGAN, so that the generated image can reconstruct the original image and deceive the discriminator. The discriminator loss is based on binary cross entropy to distinguish between real and generated images. An alternating training strategy is adopted, first fixing the discriminator to train the generator, and then fixing the generator to train the discriminator, and the model converges after multiple iterations.
[0041] Step 3: training and optimizing the hybrid model; specifically:
[0042] 3.1 Initialize the parameters of the generator and discriminator;
[0043] 3.2 Iterative training: Randomly extract batches of image samples from the dataset, the generator generates fake images that match the batches of image samples, the discriminator receives real image samples and fake image samples generated by the generator, and classifies them, calculates the loss functions of the generator and the discriminator, and updates the parameters to minimize the loss; repeat the above steps until the predetermined number of training rounds is reached or the loss converges to a stable state;
[0044] 3.3 Optimize the trained model: Use an optimizer to update model parameters, adjust the learning rate and other hyperparameters to improve the efficiency and stability of training. In this embodiment, learning rate decay, momentum adjustment and other techniques are used to improve the convergence performance of the model; in specific implementation, the optimizer uses Adam or SGD. Specifically:
[0045] 1) Learning rate decay adjustment: Gradually reducing the learning rate as training progresses will help the model converge better. You can use step decay, such as multiplying the learning rate by a decay coefficient less than 1 after a certain number of training steps. You can also use exponential decay to make the learning rate decrease exponentially with the number of training steps. After each training cycle, call the step() method of the corresponding scheduler to update the learning rate.
[0046] 2) Momentum adjustment: Momentum can accelerate convergence, especially when dealing with high curvature or small batch data. In SGD, the momentum parameter is set, usually around 0.9; for Adam, there is also a similar momentum mechanism inside, and the beta1 and beta2 parameters can be adjusted to affect the momentum effect.
[0047] 3) Monitor and adjust hyperparameters:
[0048] Monitor the training process: observe the changes in indicators such as loss function value and accuracy to determine the convergence of the model. You can use tools such as TensorBoard to visualize the training process to better understand the model behavior;
[0049] Adjust hyperparameters: Adjust hyperparameters such as learning rate, momentum, and decay coefficient according to training performance. You can use grid search, random search, and other methods to find the optimal hyperparameter combination.
[0050] Step 4: post-processing the generated expanded image to improve image quality and fidelity; the post-processing specifically includes:
[0051] Denoising: Denoising is performed on the generated image to reduce noise points and artifacts in the image; the denoising methods in this embodiment include median filtering and Gaussian filtering.
[0052] Sharpening: sharpening the generated image to enhance the edges and details of the image; in this embodiment, a sharpening filter or an enhancement algorithm is used to achieve image sharpening.
[0053] Color correction: Perform color correction on the generated image to adjust the hue, contrast and saturation of the image; in this embodiment, a color correction algorithm or histogram equalization method is used to achieve color correction.
[0054] Sizing: adjusting the size of the generated image to meet specific size requirements or application scenarios; in this embodiment, an interpolation algorithm (such as nearest neighbor interpolation, bilinear interpolation, etc.) is used to perform image size adjustment.
[0055] Style transfer: applying different artistic styles to the generated image to increase the artistic sense or visual appeal of the image; in this embodiment, a style transfer algorithm or a deep learning model is used to achieve image style transfer.
[0056] Image synthesis: synthesizing multiple image elements together to generate a new image scene or a combined image effect; in this embodiment, an image synthesis algorithm or fusion technology is used to achieve image synthesis.
[0057] Enhance details: Enhance the details of specific areas or objects in the generated image to highlight the key points or improve the visual effect of the image; in this embodiment, methods such as local contrast enhancement (CLAHE) and detail enhancement filters are used to enhance the details of the image.
[0058] In the specific implementation, after step 4, the model is further tested and evaluated, specifically: the trained model is verified and evaluated using the test set, where the evaluation indicators include evaluating the trained image expansion model, calculating the accuracy and recall rate, and evaluating the performance and generalization ability of the model; after the evaluation, the model is tuned and improved to improve the performance and effect of the algorithm.
[0059] In the specific implementation, after step 4, further real-time application and system deployment are carried out, specifically: applying the optimized image expansion method to the actual scene, performing real-time monitoring and control, and ensuring that the model can run in real-time applications.
[0060] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
[0061] It is obvious to those skilled in the art that the embodiments of the present invention are not limited to the details of the above exemplary embodiments, and that the embodiments of the present invention can be implemented in other specific forms without departing from the spirit or basic features of the embodiments of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the embodiments of the present invention is limited by the attached claims rather than the above description, so it is intended to include all changes that fall within the meaning and scope of the equivalent elements of the claims in the embodiments of the present invention. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units, modules or devices stated in the system, device or terminal claims can also be implemented by the same unit, module or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
[0062] Finally, it should be noted that the above implementation modes are only used to illustrate the technical solutions of the embodiments of the present invention and are not intended to limit them. Although the embodiments of the present invention have been described in detail with reference to the above preferred implementation modes, those skilled in the art should understand that the technical solutions of the embodiments of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A deep learning-based image expansion method, characterized in that: The following steps are involved: Step 1: Form a hybrid model for generating images by constructing a VAE model and a CGan model; specifically: 2.1 Building a VAE model: The encoder receives the input image and maps it to the mean and variance vectors in the latent space. The decoder receives samples from the latent space and decodes them into the generated image. 2.2 Constructing the CGan model: The generator receives random noise vectors and condition information as input, and converts the input into an image that matches the condition through a deconvolutional neural network; the discriminator receives the image and condition information as input, and classifies the input into real images / fake images and matches the condition through a convolutional neural network; The loss function of the generator is determined by the probability that the discriminator misjudges the generated image, that is, the generator hopes that the generated image will be regarded as a real image by the discriminator; 2.3 Combine the VAE model and the CGAN model to form a hybrid model for generating images, where the encoder generates latent vectors by learning the latent representation of the image; the generator receives these latent vectors and generates realistic images, and the discriminator tries to distinguish between real images and generated images; Step 2: training and optimizing the hybrid model.
2. The image expansion method based on deep learning according to claim 1, characterized in that: Step 2 specifically includes: 3.1 Initialize the parameters of the generator and discriminator; 3.2 Iterative training: Randomly extract batches of image samples from the dataset, the generator generates fake images that match the batches of image samples, the discriminator receives real image samples and fake image samples generated by the generator, and classifies them, calculates the loss functions of the generator and the discriminator, and updates the parameters to minimize the loss; repeat the above steps until the predetermined number of training rounds is reached or the loss converges to a stable state; 3.3 Optimize the trained model: Use the optimizer to update the model parameters and adjust the learning rate and other hyperparameters to improve the efficiency and stability of training.
3. The image expansion method based on deep learning according to claim 1 or 2, characterized in that: After step 2, the following further includes: post-processing the generated expanded image, specifically including: Denoising: Denoising is performed on the generated image to reduce noise and artifacts in the image; Sharpening: sharpen the generated image to enhance the edges and details of the image; Color Correction: Perform color correction on the generated image to adjust the image's hue, contrast, and saturation; Resizing: Adjust the size of the generated image to meet specific size requirements or application scenarios; Style transfer: applying different artistic styles to the generated images to increase the artistic feel or visual appeal of the images; Image synthesis: synthesizing multiple image elements together to generate new image scenes or combined image effects; Enhance details: Enhance the details of specific areas or objects in the generated image to highlight the key points or improve the visual effect of the image.
4. The image expansion method based on deep learning according to claim 1 or 2, characterized in that: Before step one, data collection and preprocessing are further included; specifically, using a data set or collecting image data for target tracking, and adjusting the image to a uniform size, and then using a denoising algorithm to reduce noise in the image, and converting the image from RGB to other color spaces, reducing the impact of raindrops according to the characteristics of different channels, and finally standardizing the image.
5. The image expansion method based on deep learning according to claim 1 or 2, characterized in that: The loss of the discriminator is determined by the probability of correctly classifying real images and generated images, that is, the discriminator hopes to correctly distinguish between real images and generated images.
6. The image expansion method based on deep learning according to claim 1 or 2, characterized in that: The loss function includes reconstruction loss and KL divergence loss, wherein the reconstruction loss is used to measure the difference between the generated image and the original image, and the KL divergence loss is used to penalize the difference between the latent space distribution and the standard normal distribution.
7. The image expansion method based on deep learning according to claim 1 or 2, characterized in that: The conditional information is a category label or a text description, which is used to guide the generator to generate images of a specific category.