A data processing method and system based on an image acquisition device
By building a multimodal image processing system based on CNN-VAE and GAN, combined with U-Net and ResNet50 feature extraction, the problems of insufficient fusion and loss of details in multimodal image processing are solved, and high-precision object detection and recognition are achieved, improving the robustness and accuracy of image processing.
Patent Information
- Application Number
- CN202510297710.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-13
AI Technical Summary
In the multimodal image processing technology, existing image processing techniques have problems such as insufficient image fusion, loss of details and inaccurate object detection. Especially in low-light and shadow areas, existing methods are difficult to effectively combine features of different modes, resulting in limited accuracy and robustness of object recognition.
The training data is expanded by using CNN-VAE and inverse attention GAN models, and the image details are enhanced by dual-branch U-Net and local variance adaptive Gaussian filtering. Edge detection is optimized by support vector machine, and the data processing system is built by combining ResNet50 feature extraction and RPN network with non-maximum suppression positioning candidate boxes.
It significantly improves the accuracy and robustness of image processing, can accurately detect and identify objects under low light conditions, and enhances the image's detail retention and edge detection capabilities.
Smart Images

Figure CN119810394B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a data processing method and system based on an image acquisition device. Background Art
[0002] With the wide application of image acquisition devices, image processing technology has been greatly developed in multiple fields. Especially in the fields of medical imaging, autonomous driving, security monitoring, and remote sensing image analysis, the precise processing and analysis of image data have become core technologies. In recent years, image processing methods based on deep learning have gradually taken the leading position. In particular, the wide application of technologies such as convolutional neural networks (CNNs) and generative adversarial networks (GANs) has enabled remarkable progress in tasks such as image classification, detection, and enhancement. As a powerful feature extraction tool, CNNs have been widely applied to image recognition and segmentation tasks, while GANs have shown great potential in image generation and enhancement due to their generation ability. At the same time, variational autoencoders (VAEs) have also been increasingly applied to the field of image processing due to their advantages in generative models and data reconstruction. The combination of these technologies provides a more flexible solution for multi-modal image processing.
[0003] However, existing image processing technologies still face many challenges when dealing with multi-modal image processing. Multi-modal images usually come from different sensor devices, such as RGB images and infrared images, etc. The lighting conditions, resolutions, and detail performances of the two are quite different. How to effectively fuse these images and ensure the accuracy of processing remains a difficult problem. Existing image enhancement methods, such as those based on simple convolution operations, often struggle to fully extract and retain the detail information of images. Especially in low-light and shadow areas, detail loss often occurs. Although generative adversarial networks (GANs) can effectively enhance image quality, the generated images still have certain quality problems under complex conditions such as uneven lighting and noise, especially the insufficient accuracy in edge detection and object localization. Although existing deep learning-based object detection methods have achieved good results in many fields, in the joint processing and analysis of multi-modal data, existing technologies often fail to effectively combine the features of different modalities, resulting in limitations in the accuracy and robustness of object recognition. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a data processing method and system based on an image acquisition device to solve the problems of insufficient image fusion, detail loss, and inaccurate object detection existing in existing image processing technologies when processing multi-modal images;
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a data processing method based on an image acquisition device, which includes:
[0008] Collect multi-modal images of the image acquisition device and perform preprocessing;
[0009] The multi-modal images include RGB images and infrared images;
[0010] Integrate the multi-modal images into a training data set, construct a CNN-VAE to generate low-light images of multi-modal images, adopt a GAN model with an anti-attention module and a dual discriminator structure, and combine parallel extended dilated convolution to expand the training data set;
[0011] Enhance the images in the training data set through a dual-branch U-Net, enhance the image details using local variance adaptive Gaussian filtering, classify the pixel gradients through a support vector machine and perform edge detection;
[0012] Adopt ResNet50 to extract features, combine RPN with non-maximum suppression to accurately locate the candidate boxes, perform object detection and recognition through classification regression and perform visual display;
[0013] Store all data in the database and manage it.
[0014] As a preferred solution of the data processing method based on the image acquisition device of the present invention, wherein: the construction of the CNN-VAE to generate low-light images of multi-modal images, adopting a GAN model with an anti-attention module and a dual discriminator structure, and combining parallel extended dilated convolution to expand the training data set includes:
[0015] Use random rotation to simulate different directions of the preprocessed images, adjust the brightness of the multi-modal images through histogram equalization and integrate them into a training data set, map the training data set to the latent space through a variational autoencoder, and sample from the latent space to obtain latent space data , by maximizing the variational lower bound, make the variational autoencoder learn the distribution of the latent space and generate low-light images by minimizing the loss function, and use the generated low-light images as new samples to expand the training data set;
[0016] Use a generative adversarial network to expand the training data set, define an anti-attention module, construct a dual discriminator structure to extract the global and local features of the images and add them to each layer of the generator, use global average pooling to encode the spatial information of the multi-modal images in the training data set, perform non-linear transformation through the Sigmoid activation function, generate an anti-attention map and suppress the unwanted regions through reverse operations to generate images;
[0017] Integrate a parallel expansion convolutional module in the generative adversarial network, expand the receptive field of the generative adversarial network and perform image enhancement by using convolutions with different dilation rates in parallel, and output the final enhanced dataset.
[0018] As a preferred solution of the data processing method based on an image acquisition device according to the present invention, wherein: the constructing a dual discriminator structure to extract the global features and local features of the image includes:
[0019] The anti-attention module extracts the global features and local features of the image, adjusts the weights of the feature channels according to the importance of the global features and local features, uses PatchGAN as the local discriminator, and combines with the global discriminator for training. The real image and the generated image are used to train the discriminator, the discriminator loss is calculated, and the Adam optimizer is used to train the model until convergence.
[0020] As a preferred solution of the data processing method based on an image acquisition device according to the present invention, wherein: the enhancing the images in the training dataset by a dual-branch U-Net includes:
[0021] Use the U-Net architecture as the joint learning framework and integrate the anti-attention module for each convolutional layer. Input the RGB image and the infrared image in the final enhanced dataset into two parallel U-Net branches respectively to generate enhanced images. Retain the local features of the images through skip connections. Use the joint loss function to update the parameters of the U-Net, define the optimizer, calculate the gradients and update, optimize the U-Net model, and obtain the final enhanced images.
[0022] As a preferred solution of the data processing method based on an image acquisition device according to the present invention, wherein: the enhancing the image details by using local variance adaptive Gaussian filtering and classifying pixel gradients by a support vector machine and performing edge detection includes:
[0023] Convert the enhanced image into a grayscale image, perform block processing of a fixed size on the grayscale image to obtain a plurality of local sub-regions, calculate the local grayscale variance values within each local sub-region and use the grayscale variance values of all local sub-regions as a one-dimensional distribution, and determine the optimal grayscale variance threshold by calculating the grayscale histogram of the grayscale image and the Otsu algorithm formula;
[0024] Determine the unique local Gaussian smoothing window size for each sub-region according to the optimal grayscale variance threshold, perform a deterministic Gaussian filtering operation on the grayscale value at each pixel position within each corresponding region in the original image, replace the original grayscale value with the filtered grayscale value, deterministically splice the filtered sub-regions back to the complete image of the original size according to the original image position, calculate the gradient value of each pixel of the image, calculate the gradient amplitude and direction through the Canny edge detection algorithm, and perform edge detection on each pixel.
[0025] Select the support vector machine algorithm, use the RBF kernel as the kernel function, extract the features of each pixel from the enhanced image as the input features of the training data set, mark the edge detection situation of each pixel in the input features, use the labeled training data set, and train the model by minimizing the loss function of the support vector machine. By using the gradient descent method, solve to obtain the optimal hyperplane and bias;
[0026] Input the features of each pixel of the image into the trained support vector machine model for classification, and adjust the edge probability of each pixel in the image based on the output of the support vector machine model.
[0027] As a preferred solution of the data processing method based on the image acquisition device according to the present invention, wherein: the extraction of features using ResNet50, the accurate positioning of candidate boxes by combining RPN and non-maximum suppression, object detection and recognition by classification regression, and visual display include:
[0028] Use a pre-trained ResNet50 model as a feature extractor;
[0029] Use global average pooling to pool the features of each channel to obtain a one-dimensional feature vector, use the Region Proposal Network of Faster R-CNN to generate object candidate boxes, use the non-maximum suppression algorithm to select the optimal candidate box from multiple object candidate boxes, use the Softmax classifier to classify each candidate box, judge the category of each candidate box by setting a threshold, perform a regression operation on each candidate box, predict the specific position and size of the object within the box, visualize the object detection and recognition results, and generate the final output.
[0030] As a preferred solution of the data processing method based on the image acquisition device according to the present invention, wherein: the storage of all data in a database and management include:
[0031] Select a relational database to manage data and relationship analysis results, design the database table structure to store different types of data, set up a regular backup task to back up all data in the database, perform permission management on database users, and encrypt and store static data.
[0032] In a second aspect, the present invention provides a data processing system based on an image acquisition device, including:
[0033] A data acquisition module for acquiring multi-modal images of the image acquisition device and performing preprocessing;
[0034] The model construction module is used to integrate multi-modal images into a training dataset, construct low-light images of CNN-VAE generated multi-modal images, adopt a GAN model with an anti-attention module and a dual discriminator structure, and combine parallel extended dilated convolution to expand the training dataset;
[0035] The image enhancement module is used to enhance the images in the training dataset through a dual-branch U-Net;
[0036] The edge detection module is used to enhance image details by using locally adaptive Gaussian filtering of variance, classify pixel gradients through a support vector machine, and perform edge detection;
[0037] The feature extraction module is used to extract features using ResNet50, and combine RPN with non-maximum suppression to accurately locate candidate boxes;
[0038] The object recognition module is used to perform object detection and recognition through classification regression and perform visual display;
[0039] The data storage module is used to store all data in a database and manage it.
[0040] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the data processing method based on an image acquisition device as described in the first aspect of the present invention is implemented.
[0041] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the data processing method based on an image acquisition device as described in the first aspect of the present invention is implemented.
[0042] The beneficial effects of the present invention are: expanding the training data through CNN-VAE and anti-attention GAN models, enhancing image details using a dual-branch U-Net and locally adaptive Gaussian filtering of variance, and optimizing edge detection through a support vector machine, extracting features using ResNet50, combining with the RPN network and non-maximum suppression to locate object candidate boxes, achieving high-precision object detection and recognition, and significantly improving the accuracy and robustness of image processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1Flow chart of the data processing method based on an image acquisition device in the present invention.
[0045] Figure 2 Schematic diagram of the data processing system based on an image acquisition device in the present invention.
[0046] Figure 3 Flow chart of expanding the training data set in the present invention.
[0047] Figure 4 Flow chart of the support vector machine classifying pixel gradients and performing edge detection in the present invention. Detailed implementation manners
[0048] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings of the specification.
[0049] In the following description, many specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0050] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that excludes other embodiments.
[0051] Embodiment, referring to Figures 1 to 4 , this embodiment provides a data processing method based on an image acquisition device, including the following steps:
[0052] S1. Collect multi-modal images of the image acquisition device and perform preprocessing;
[0053] Specifically, according to the task requirements, in low-light environments (such as at night or on cloudy days), use an image acquisition device (such as a low-light photography device and an infrared sensor) to synchronously collect multi-modal images of the target area. Assume the target scenario is remote sensing monitoring;
[0054] The RGB image is a three-channel image of red, green, and blue. The data of each channel is a matrix, and the value of each pixel represents the RGB color information;
[0055] The infrared image is a single-channel grayscale image, and the acquisition timestamps are aligned to ensure that the multi-modal images correspond to the same area;
[0056] The preprocessing includes spatially aligning and normalizing the resolution of the multimodal images, upsampling the infrared image to the resolution of the RGB image, calculating new pixel values using bilinear interpolation, applying non-local means denoising algorithms to the RGB image and the infrared image respectively to suppress noise during device acquisition, checking for missing regions in the multimodal images and filling them using low-rank approximation methods based on matrix completion, and normalizing the pixel values of the multimodal images.
[0057] Through spatial alignment and resolution normalization, this preprocessing process makes the infrared image have the same resolution as the RGB image, ensuring the spatial consistency of the images, and thus facilitating subsequent analysis and fusion. The use of non-local means denoising algorithms effectively suppresses the noise that may be generated during acquisition, improving the clarity of the images. Especially in low-light environments, the image quality is enhanced. By filling the missing regions in the multimodal images using low-rank approximation methods based on matrix completion, the integrity of the data is ensured. Pixel value normalization eliminates the brightness differences under different image sources and acquisition conditions, ensuring a unified scale for the images, which helps to improve the training effect of subsequent deep learning models, thereby enhancing the accuracy and robustness of remote sensing monitoring tasks.
[0058] S2. Integrate the multimodal images into a training dataset, construct a CNN-VAE to generate low-light images of multimodal images, adopt a GAN model with an inverse attention module and a dual discriminator structure, and combine parallel extended dilated convolutions to expand the training dataset;
[0059] Specifically, constructing a CNN-VAE to generate low-light images of multimodal images, adopting a GAN model with an inverse attention module and a dual discriminator structure, and combining parallel extended dilated convolutions to expand the training dataset includes:
[0060] Use random rotation to simulate different directions of the preprocessed images, adjust the brightness of the multimodal images through histogram equalization, and integrate them into a training dataset;
[0061] To increase the diversity of the training dataset and improve the robustness of the model, the training dataset is mapped to the latent space through a variational autoencoder (VAE). The structure of the variational autoencoder includes an encoder, a decoder, and a latent space. The encoder is designed as a 5-layer convolutional neural network (CNN), outputting the mean and variance of the latent space (with the dimension set to 128). The decoder is designed as a 5-layer transposed convolution network, sampling latent space data from the latent space , the formula is:
[0062] ,
[0063] where, represents the mean of the latent space, represents the variance of the latent space, represents standard normal distribution random noise;
[0064] By maximizing the variational lower bound (ELBO), the variational autoencoder learns the distribution of the latent space and generates low-light images by minimizing the loss function, providing more samples for the model and enhancing the generalization ability of the model. The loss function is:
[0065] ,
[0066] where, represents the total loss function value of the variational autoencoder, is the output of the encoder, representing the distribution of the given input image and the latent space data , and the distribution is obtained by variational inference, is the output of the decoder, representing the probability distribution of generating the input image under the condition of the given latent space data , represents the expected value, that is, the expectation operator, represents the KL divergence, measuring the difference between the encoder distribution and the prior distribution , is the expected value part, representing the log-likelihood of reconstructing the input image when the given latent space data is , measuring the possibility of the latent space data reconstructing the input image in the decoder. By maximizing this expected value, the simulation can generate output images similar to the input image, represents the prior distribution of the latent space data . The variational autoencoder usually uses the standard normal distribution as the prior distribution;
[0067] Use the generated low-light images as new samples to expand the training dataset;
[0068] To further optimize the image quality, use the generative adversarial network (GAN) to expand the training dataset, define the anti-attention module (AAB), construct a dual discriminator structure to extract the global and local features of the image and add them to each layer of the generator, use global average pooling to encode the spatial information of the multimodal images in the training dataset, perform non-linear transformation through the Sigmoid activation function, generate the anti-attention map and perform anti-operations to suppress the unwanted regions (suppress the light pollution regions through anti-attention operations and emphasize the dark details of the image), and generate images;
[0069] Integrate a parallel extended convolution module in the generative adversarial network. By using convolutions with different dilation rates in parallel (such as dilation rates of 1, 3, and 5), expand the receptive field of the generative adversarial network and perform image enhancement to capture more context information, and output the final enhanced dataset (including the original image, VAE-enhanced image, and high-quality images generated by GAN). The formula is:
[0070] ,
[0071] Among them, represents the enhanced image, represents the splicing operation of images, represents the dilated convolution operation on the image and represents the dilation rate, which can extract features from multiple scales.
[0072] Based on the above content, the proposed scheme combines multiple advanced technologies, significantly enhancing the diversity and image quality of the training dataset, thereby improving the generalization ability and robustness of the model. Random rotation and histogram equalization are used for preprocessing and brightness adjustment of multimodal images to expand the diversity of the training dataset. Then, the variational autoencoder (VAE) is used to map the data to the latent space and generate low-light images, which are used to expand the dataset and enhance the model's adaptability to low-light conditions. Through the generative adversarial network (GAN) and the anti-attention module (AAB), the generated images are further optimized to emphasize dark details and suppress unwanted regions such as light pollution, generating clearer and more natural low-light images. Integrate the parallel extended dilated convolution module to extract multi-scale context information through convolutions with different dilation rates, enhancing the receptive field and quality of the generated images. Through these steps, the generated enhanced images can not only improve the model's performance under low-light conditions but also, through the combination of multimodal data, help the model perform better under different environmental conditions, enhancing the robustness and accuracy of the model.
[0073] Furthermore, a dual discriminator structure is constructed to extract the global and local features of the image, including:
[0074] The anti-attention module extracts the global features of the image (i.e., information such as the overall brightness and contrast of the image) and local features (i.e., the texture of the shadows or dark parts in the image), ensuring that the generative adversarial network not only pays attention to the overall perception of the image but also can handle dark details. According to the importance of the global and local features, the weights of the feature channels are adjusted. In this process, the weights of the dark regions are particularly strengthened, so that during the image generation process, the dark details can be fully restored;
[0075] PatchGAN is used as the local discriminator and trained in combination with the global discriminator. The real images and the generated images are used to train the discriminator, and the discriminator loss is calculated. The Adam optimizer is used to train the model until convergence. The global discriminator is responsible for evaluating the authenticity of the entire generated image and outputs a binary classification result indicating whether the image is real. Its structure is generally based on a fully convolutional network (FCN), and its goal is to evaluate the large-scale features of the image (such as texture, hue, etc.). The local discriminator focuses on the local details of the image and often uses a smaller convolutional kernel to process multiple local regions of the image (such as a 32×32 region). It can refine the details of the local regions of the image during the generation process, such as edges, textures, etc., to ensure that the generated image can also achieve a sense of reality at the detail level;
[0076] By combining the global discriminator and the local discriminator, the dual discriminator structure can enhance the generator's attention to image details, making the generated image not only look real as a whole but also more refined in details. The detail processing ability of the local discriminator can avoid distortion in some local regions of the generated image (such as shadow or highlight parts).
[0077] By constructing a dual discriminator structure and combining an anti-attention module, the performance of the generative adversarial network (GAN) in the image generation process can be significantly improved. The anti-attention module extracts global features (such as brightness, contrast, etc.) and local features (such as the texture of shadows or dark parts), and dynamically adjusts the weights of the feature channels according to the importance of these features, especially strengthening the weights of the dark regions, so as to ensure that the details of the dark parts are fully restored. Using PatchGAN as the local discriminator and combining it with the global discriminator based on the fully convolutional network (FCN), the dual discriminator structure enables the generator to not only pay attention to the overall perception of the image (such as texture, hue), but also finely process local details (such as edges, textures, etc.). This structure effectively improves the sense of reality of the generated image at the detail level and avoids distortion in local regions, especially in the shadow and highlight parts. Finally, the generated image can achieve a high quality both in terms of overall and details, enhancing the sense of reality and visual effect of the image.
[0078] S3. Enhance the images in the training dataset through a dual-branch U-Net, use local variance adaptive Gaussian filtering to enhance image details, classify pixel gradients through a support vector machine, and perform edge detection;
[0079] Specifically, enhancing the images in the training dataset through a dual-branch U-Net includes:
[0080] Use the U-Net architecture as the joint learning framework and integrate an anti-attention module for each convolutional layer to enhance the dark and detailed areas of the image. With the help of the anti-attention module, U-Net can generate clearer images under low-light and shadow conditions. Input the RGB images and infrared images in the final enhanced dataset into two parallel U-Net branches respectively to generate enhanced images. Retain the local features of the image through skip connections. Use the joint loss function to update the parameters of U-Net, define the optimizer, calculate the gradients and update them to optimize the U-Net model and obtain the final enhanced image;
[0081] The joint loss function includes exposure consistency loss, color consistency loss, and spatial consistency loss. The exposure consistency loss ensures that the exposure of the image is within a reasonable range, avoiding the image being too dark or too bright. This loss function calculates the difference between the brightness of the local area of the image and the target exposure range , and the loss function is:
[0082] ,
[0083] where, represents the enhanced image in the training dataset the th local area brightness value, represents the ideal exposure range. Usually, we hope that the brightness of the image is within this range, indicating that the exposure is moderate, represents the total number of local areas in the image;
[0084] The color consistency loss is used to avoid color distortion in the generated image and ensure that the colors of the RGB channels of the image are similar. This loss measures the color difference between the red, green, and blue channels of the image , and the loss function is:
[0085] ,
[0086] where, , , respectively represent the average intensity values of the red, green, and blue channels after image enhancement;
[0087] The spatial consistency loss ensures the consistency of the spatial structure between the generated image and the original image, avoiding the loss of image details. This loss calculates the intensity difference of the local area of the image , reflecting the degree of retention of the image structure;
[0088] ,
[0089] where, and respectively represent the average intensity values of the enhanced image and the original image in the th local region;
[0090] Calculate the joint loss function , and the formula is:
[0091] ,
[0092] where , , represent the weights of the loss function.
[0093] Image enhancement on the training dataset through a dual-branch U-Net architecture can effectively improve the quality of images under low-light and shadow conditions. The U-Net architecture combined with an anti-attention module enables the model to focus on the dark and detailed areas in the image, thereby enhancing the clarity and brightness of the image. This method processes RGB images and infrared images in parallel, and uses skip connections to retain the local features of the image, ensuring that important details are not lost during the enhancement process. The joint loss function includes exposure consistency loss, color consistency loss, and spatial consistency loss, which respectively ensure the consistency of the exposure, color, and spatial structure of the image, thus avoiding over-darkening, color distortion, and detail loss of the image. By optimizing the joint loss function, the model can generate more realistic and clear images under different environmental conditions, improve the diversity and quality of the training dataset, and ultimately enhance the model's adaptability to low-light and complex environments.
[0094] Furthermore, local variance adaptive Gaussian filtering is used to enhance image details. Pixel gradients are classified by a support vector machine and edge detection includes:
[0095] Convert the enhanced image to a grayscale image, perform block processing of a fixed size on the grayscale image, and use a fixed square window of 16 × 16 pixels to scan and segment the grayscale image region by region to obtain multiple local sub-regions. Calculate the local grayscale variance value within each local sub-region and use the grayscale variance values of all local sub-regions as a one-dimensional distribution to measure the complexity of regional details. Set the current candidate grayscale variance threshold, and determine the optimal grayscale variance threshold by calculating the grayscale histogram of the grayscale image and the Otsu algorithm formula, so as to effectively detect edges under different lighting conditions. The formula is:
[0096] ,
[0097] where represents the optimal grayscale variance threshold, respectively represent the number of pixels with pixel grayscale variance less than and greater than or equal to the current candidate grayscale variance threshold, and respectively represent the average gray values of regions where the pixel gray variance is less than and greater than or equal to the current candidate gray variance threshold, represents the optimal gray variance threshold to be obtained, which is used to distinguish the richness of regional details;
[0098] Determine the unique local Gaussian smoothing window size for each sub-region according to the optimal gray variance threshold. The formula is:
[0099] ,
[0100] where, represents the size of the Gaussian filtering window corresponding to the region in the th row and th column of the grayscale image. If , it means that the details of this region are rich, and the Gaussian filtering window selected for this region should be 3 ×3 pixels to retain as many image details as possible. If , it means that this region is smooth, and the Gaussian filtering window selected for this region should be 7 ×7 pixels to achieve a better noise reduction effect;
[0101] Based on the unique local Gaussian smoothing window size of each sub-region, perform a deterministic Gaussian filtering operation on the gray value of each pixel position in each corresponding region of the original image, and replace the original gray value with the filtered gray value;
[0102] Stitch the filtered sub-regions back to the original size of the complete image according to the original image position deterministically;
[0103] By calculating the gradient value of each pixel in the image (i.e., how fast the image brightness changes), calculate the gradient magnitude and direction through the Canny edge detection algorithm, perform edge detection on each pixel, and determine whether each pixel is an edge or a non-edge;
[0104] Select the support vector machine (SVM) algorithm, use the RBF kernel (radial basis kernel) as the kernel function, extract the features of each pixel (such as gradient magnitude and direction) from the enhanced image as the input features of the training data set, mark the edge detection situation (edge or non-edge) of each pixel in the input features, and these labels will be used as the target output of the SVM classifier. Use the labeled training data set to train the model by minimizing the loss function of the support vector machine. By using the gradient descent method, solve for the optimal hyperplane and bias to minimize the classification error, thereby obtaining the best classification model;
[0105] The features of each pixel of the image are input into a trained support vector machine (SVM) model for classification. The SVM model will determine whether each pixel belongs to the edge region. Based on the output of the SVM model, the edge probability of each pixel in the image is adjusted to make the detected edges more accurate.
[0106] The output after SVM optimization will improve the accuracy of edge detection, reduce the phenomena of false detection and missed detection. Especially in areas where the edge transition is relatively blurred or there is more noise, SVM can help to more accurately identify the real edges.
[0107] This method combines local variance adaptive Gaussian filtering and support vector machine (SVM) classification for image detail enhancement and edge detection, significantly improving the accuracy of image edge recognition. During the detail enhancement process, local gray variance is used to measure the detail complexity of each sub-region, so as to allocate the most suitable Gaussian smoothing window for each region, which can not only retain details but also suppress noise. The SVM algorithm uses gradient magnitude and direction as features for edge classification to optimize the accuracy of edge detection. By minimizing the loss function of SVM, the model can accurately distinguish edge and non-edge pixels, reducing false detection and missed detection, especially in areas where the edge transition is blurred or there is more noise, ensuring more accurate and reliable edge detection. This method can still stably and effectively identify image edges under complex lighting conditions and has strong robustness.
[0108] S4: Use ResNet50 to extract features, combine RPN with non-maximum suppression to accurately locate candidate boxes, and perform object detection and recognition through classification regression and visualize the display.
[0109] Specifically, using ResNet50 to extract features, combining RPN with non-maximum suppression to accurately locate candidate boxes, and performing object detection and recognition through classification regression and visualize the display includes:
[0110] Use a pre-trained ResNet50 model as a feature extractor. ResNet50 is a convolutional neural network containing 50 layers, with strong image feature extraction capabilities. It can extract multi-level features from images through multiple layers of convolution, pooling and residual connections.
[0111] Pool the features of each channel using global average pooling to obtain a one-dimensional feature vector for describing the objects in the image. Use the Region Proposal Network (RPN) of Faster R-CNN to generate object candidate boxes. The RPN slides a small window over the convolutional feature map, predicts the candidate boxes at each position, and scores each candidate box to represent the probability that it contains an object. Use the non-maximum suppression (NMS) algorithm to select the optimal candidate boxes from multiple object candidate boxes. Use a Softmax classifier to classify each candidate box, and determine the category of each candidate box by setting a threshold. If the probability of a certain category exceeds the set threshold, it is considered that the object belongs to that category; otherwise, it is considered that the object does not exist or belongs to the "background" category. Perform a regression operation on each candidate box to predict the specific position and size of the object within the box. Visualize the object detection and recognition results and generate the final output.
[0112] By combining a pre-trained ResNet50 feature extractor, RPN, and non-maximum suppression algorithm, this method can accurately extract multi-level features from images, generate high-quality object candidate boxes, and perform accurate classification and regression. ResNet50 provides powerful image feature extraction capabilities. Using global average pooling effectively reduces the number of parameters and improves the model efficiency. The RPN can generate potential object regions by sliding a window and assigns a confidence score to each candidate box. Non-maximum suppression ensures the optimization of candidate boxes and avoids the interference of duplicate boxes. Through the Softmax classifier and regression operation, the model can accurately identify the object category and make a detailed prediction of the specific position of the object. Combining these methods for object detection and recognition not only improves the detection accuracy but also clearly shows the position and category of the object in the image, achieving efficient and accurate object detection and visualization.
[0113] S5. Store all data in a database and manage it;
[0114] Specifically, storing all data in a database and managing it includes:
[0115] Select a relational database to manage the data and relationship analysis results, design the database table structure to store different types of data, set up a regular backup task to back up all data in the database, manage the permissions of database users, and encrypt and store static data.
[0116] Storing all data in a database and managing it can bring multiple beneficial effects. Relational databases can effectively manage different types of data and the relationships between them, ensuring data consistency, integrity, and queryability. By reasonably designing the database table structure, data storage and retrieval performance can be optimized, improving the efficiency of data management. Regular backup tasks guarantee data security, enabling the system to be quickly restored in case of failures and preventing data loss. Database user privilege management ensures that only authorized users can access sensitive data, enhancing data security. Encrypted storage of static data can effectively prevent data leakage and ensure data confidentiality during storage. Adopting these measures can enhance data security, reliability, and management efficiency, while also strengthening the system's fault tolerance and maintainability.
[0117] This embodiment also provides a data processing system based on an image acquisition device, including:
[0118] A data acquisition module, configured to acquire multimodal images of the image acquisition device and perform preprocessing;
[0119] A model construction module, configured to integrate the multimodal images into a training data set, construct a low-light image of the multimodal image using a CNN-VAE, adopt a GAN model with an inverse attention module and a dual discriminator structure, and combine parallel extended dilated convolutions to expand the training data set;
[0120] An image enhancement module, configured to enhance the images in the training data set through a dual-branch U-Net;
[0121] An edge detection module, configured to enhance image details using locally adaptive Gaussian filtering of variance, classify pixel gradients through a support vector machine, and perform edge detection;
[0122] A feature extraction module, configured to extract features using ResNet50, and accurately locate candidate boxes in combination with RPN and non-maximum suppression;
[0123] An object recognition module, configured to perform object detection and recognition through classification regression and perform visual display;
[0124] A data storage module, configured to store all data in a database and manage it.
[0125] This embodiment also provides a computer device, applicable to the case of a data processing method based on an image acquisition device, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the data processing method based on an image acquisition device as proposed in the above embodiment.
[0126] The computer device may be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0127] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the data processing method based on an image acquisition device as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (Static Random Access Memory, abbreviated as SRAM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), an erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), a programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), a read-only memory (Read-Only Memory, abbreviated as ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk.
[0128] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A data processing method based on an image acquisition device, characterized in that: Including: Collecting multi-modal images of an image acquisition device and performing preprocessing; The multi-modal images include RGB images and infrared images; Integrating the multi-modal images into a training data set, constructing a CNN-VAE to generate low-light images of multi-modal images, using a GAN model with an anti-attention module and a dual discriminator structure, and combining parallel extended dilated convolutions to expand the training data set; Enhancing the images in the training data set through a dual-branch U-Net, enhancing image details using local variance adaptive Gaussian filtering, classifying pixel gradients through a support vector machine and performing edge detection; Extracting features using ResNet50, accurately positioning candidate boxes in combination with RPN and non-maximum suppression, performing object detection and recognition through classification regression and performing visual display; Storing all data in a database and managing it; The constructing a CNN-VAE to generate low-light images of multi-modal images, using a GAN model with an anti-attention module and a dual discriminator structure, and combining parallel extended dilated convolutions to enhance images and expand the training data set includes: Using random rotation to simulate different directions of the preprocessed images, adjusting the brightness of the multi-modal images through histogram equalization and integrating them into a training data set, mapping the training data set to the latent space through a variational autoencoder, sampling latent space data z from the latent space, maximizing the variational lower bound, enabling the variational autoencoder to learn the distribution of the latent space and generating low-light images by minimizing the loss function, and using the generated low-light images as new samples to expand the training data set; Using a generative adversarial network to expand the training data set, defining an anti-attention module, constructing a dual discriminator structure to extract global and local features of the images and adding them to each layer in the generator, using global average pooling to encode the spatial information of the multi-modal images in the training data set, performing non-linear transformation through a Sigmoid activation function, generating an anti-attention map and suppressing unnecessary regions through an inverse operation to generate images; Integrating a parallel extended convolution module into the generative adversarial network, expanding the receptive field of the generative adversarial network through parallel use of convolutions with different dilation rates and performing image enhancement, and outputting a final enhanced data set.
2. The data processing method based on an image acquisition device according to claim 1, characterized in that: The constructing a dual discriminator structure to extract global and local features of the images includes: The anti-attention module extracts global and local features of the images, adjusts the weights of the feature channels according to the importance of the global and local features, uses PatchGAN as the local discriminator, and combines it with the global discriminator for training, uses real images and generated images to train the discriminator, calculates the discriminator loss, and uses an Adam optimizer to train the model until convergence.
3. The data processing method based on an image acquisition device according to claim 1, characterized in that: The enhancing the images in the training data set through a dual-branch U-Net includes: Use the U-Net architecture as the joint learning framework and integrate the anti-attention module for each convolutional layer. Input the RGB image and the infrared image in the final enhanced dataset into two parallel U-Net branches respectively to generate enhanced images. Retain the local features of the images through skip connections. Use the joint loss function to update the parameters of the U-Net, define the optimizer, calculate the gradients and update them to optimize the U-Net model and obtain the final enhanced images.
4. The data processing method based on an image acquisition device according to claim 1, characterized in that: The enhancement of image details using local variance adaptive Gaussian filtering and the classification of pixel gradients and edge detection through a support vector machine include: Convert the enhanced image into a grayscale image, perform block processing of a fixed size on the grayscale image to obtain multiple local sub-regions, calculate the local grayscale variance value within each local sub-region and use the grayscale variance values of all local sub-regions as a one-dimensional distribution, and determine the optimal grayscale variance threshold by calculating the grayscale histogram of the grayscale image and the Otsu algorithm formula; Determine the unique local Gaussian smoothing window size for each sub-region according to the optimal grayscale variance threshold, perform deterministic Gaussian filtering operations on the grayscale values at each pixel position within each corresponding region in the original image, replace the original grayscale values with the filtered grayscale values, splice the filtered sub-regions back to the complete image of the original size according to the original image position deterministically, calculate the gradient value of each pixel of the image, calculate the gradient amplitude and direction through the Canny edge detection algorithm, and perform edge detection on each pixel; Select the support vector machine algorithm, use the RBF kernel as the kernel function, extract the features of each pixel from the enhanced image as the input features of the training dataset, mark the edge detection situation of each pixel in the input features, use the labeled training dataset, train the model by minimizing the loss function of the support vector machine, and solve for the optimal hyperplane and bias by using the gradient descent method; Input the features of each pixel of the image into the trained support vector machine model for classification, and adjust the edge probability of each pixel in the image based on the output of the support vector machine model.
5. The data processing method based on an image acquisition device according to claim 1, wherein: The extraction of features using ResNet50, the precise positioning of candidate boxes by combining RPN with non-maximum suppression, and object detection and recognition through classification and regression and visual display include: Use the pre-trained ResNet50 model as the feature extractor; Use global average pooling to pool the features of each channel to obtain a one-dimensional feature vector, use the Region Proposal Network of Faster R-CNN to generate object candidate boxes, use the non-maximum suppression algorithm to select the optimal candidate box from multiple object candidate boxes, use the Softmax classifier to classify each candidate box, judge the category of each candidate box by setting a threshold, perform regression operations on each candidate box, predict the specific position and size of the object within the box, visualize the object detection and recognition results, and generate the final output.
6. The data processing method based on an image acquisition device according to claim 1, characterized in that: The storage of all data in the database and management include: Select a relational database to manage data and relationship analysis results, design the database table structure to store different types of data, set up a regular backup task to back up all data in the database, manage the permissions of database users, and encrypt and store static data.
7. A data processing system for an image acquisition device based on the data processing method for an image acquisition device according to any one of claims 1-6, characterized in that: Including: A data acquisition module for acquiring and preprocessing multi-modal images of an image acquisition device; A model construction module for integrating multi-modal images into a training data set, constructing a CNN-VAE to generate low-light images of multi-modal images, adopting a GAN model with an anti-attention module and a dual discriminator structure, and combining parallel extended dilated convolution to expand the training data set; An image enhancement module for enhancing the images in the training data set through a dual-branch U-Net; An edge detection module for enhancing image details using locally adaptive Gaussian filtering of local variance, classifying pixel gradients through a support vector machine, and performing edge detection; A feature extraction module for extracting features using ResNet50 and accurately positioning candidate boxes in combination with RPN and non-maximum suppression; An object recognition module for performing object detection and recognition through classification regression and visual display; A data storage module for storing and managing all data in a database.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the data processing method based on an image acquisition device according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the data processing method based on an image acquisition device according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multispectral target detection blind guiding system
CN112418163A
Power transformation equipment defect image data expansion and data cleaning method
CN117079078A
Steel bridge defect fastener detection model construction method and system based on data enhancement
CN118781447A