Lightweight hyperspectral recognition method and system for sea surface oil spills based on smooth activation function

By constructing the SR-SqueezeNet lightweight hyperspectral recognition model based on smooth activation function, the problems of complex models and low recognition accuracy in UAV remote sensing technology are solved, and efficient and accurate sea surface oil spill detection is achieved.

CN120014429BActive Publication Date: 2025-09-23FIRST INSTITUTE OF OCEANOGRAPHY MNR
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411896212.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-09-23
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing drone remote sensing technology has problems in marine oil spill identification, such as complex models, many parameters, long training time, and unsuitability for real-time monitoring. In addition, lightweight models have low recognition accuracy in complex environments.

Method used

The SR-SqueezeNet lightweight hyperspectral recognition model based on smooth activation function is adopted. By constructing the smooth activation function Smooth-ReLU, the SqueezeNet network structure is optimized to reduce parameters and improve recognition accuracy.

Benefits of technology

It achieves significant reduction in computing resource consumption while maintaining high recognition performance, improves the efficiency and accuracy of sea surface oil spill detection, and is suitable for airborne environments with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014429B_ABST
    Figure CN120014429B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of marine oil spill information identification and discloses a lightweight hyperspectral identification method and system for marine oil spills based on a smooth activation function. The method acquires land-based and airborne oil spill data and real images at different times to construct a full-chain system for airborne oil spill identification and verification. The method also constructs a lightweight oil spill identification model, SR-SqueezeNet, analyzes and searches for the optimal parameters of the lightweight oil spill identification model, constructs a smooth activation function, Smooth-ReLU, and conducts a comparative analysis of different activation functions and their application locations. The method also verifies the comparative experimental results of the lightweight oil spill identification model, SR-SqueezeNet, and prior art models. The present invention improves recognition accuracy by 1.92%, reduces the number of parameters by 75.11%, and reduces the model size from 26.46MB to 12.15MB.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of marine oil spill information recognition, and in particular relates to a lightweight hyperspectral recognition method and system for marine oil spills based on a smoothed activation function. Background Art

[0002] The demand for real-time identification of oil spills in disaster response is very urgent. Drones have become an important means of monitoring sea oil spills due to their flexibility, speed and low cost. Therefore, it is crucial to develop a lightweight drone identification model.

[0003] The marine environment is closely related to human life, and marine oil spills are frequent. Difficult-to-clean oil spills and their emulsions cause various hazards to the marine and coastal environments, resulting in long-term negative impacts. Marine oil spills are characterized by strong suddenness, a large distribution range after occurrence, and high dynamics of drift and diffusion. Accurate detection of oil spills, precise estimation of the scope and amount of oil spills, and efficient tracking of their dynamic distribution are prerequisites for effective management of oil spill disasters. The emergence of drone remote sensing technology has brought revolutionary changes to marine oil spill emergency response. UAVs equipped with hyperspectral cameras, thermal infrared cameras, lidars and other equipment can achieve fast and accurate sea surface oil spill data collection and measurement. At the same time, it has the advantages of flexible operation, low cost, and high efficiency, and is suitable for oil spill detection in various complex sea conditions. [7] However, the application of drone remote sensing still faces some challenges, such as the limited number of airborne image datasets available for analysis and model training, the complexity of airborne models and airborne data processing, and the questionable airspace management and safety.

[0004] Among the commonly used remote sensing methods, thermal infrared remote sensing is an important means to study the emission characteristics of ground objects. [8] However, airborne thermal infrared remote sensing suffers from problems such as limited available data, traditional processing methods, limited inversion accuracy, and insufficient attention. Hyperspectral remote sensing is also an important means of monitoring marine oil spills in optical remote sensing. It can obtain continuous spectral characteristics at fine spectral scales, which helps to accurately identify ground objects and accurately invert ground and atmospheric characteristic parameters. Based on its principles, relevant scholars at home and abroad have conducted relevant research on airborne hyperspectral identification of oil spills. For example, Ren Guangbo et al. used drone hyperspectral to construct a marine oil spill detection model and obtained effective characteristic bands for oil spill identification. However, the common problem is that the model is relatively complex, has many parameters, and takes a long time to train, making it unsuitable for airborne real-time monitoring. Therefore, it is very necessary to develop a lightweight airborne oil spill identification model.

[0005] In the field of image processing, deep learning models, compared to traditional models, can extract higher-dimensional, more abstract, and more expressive information by building multi-layer networks and training them to reveal implicit internal relationships between data, thus improving oil spill detection. However, these models are often too complex to be suitable for airborne applications with limited computing resources. To address this issue, a common approach is to leverage existing neural network models and, through lossy compression, compress them into a model with fewer parameters while maintaining accuracy. Existing technologies focus on simplifying and compressing complex deep learning models. For example, Denton et al. applied singular value decomposition (SVD) to a pre-trained CNN model. Han et al. developed network pruning, which replaces parameters below a certain threshold in a pre-trained model with zeros to form a sparse matrix, followed by iterative training. In the field of oil spill detection, Hou et al. proposed an improved DeepLabv3+ model that reduces computational complexity and improves the accuracy of detecting small oil spills in complex environments. L. Chen et al. proposed a lightweight oil spill detection network based on YOLOv3. While this simplified the model complexity, the detection accuracy was lower. In summary, the main approach adopted by lightweight models is to reduce parameters and reduce model complexity, but this will also lead to a decrease in the accuracy of oil spill identification. Summary of the Invention

[0006] To overcome the problems existing in the related art, the disclosed embodiments of the present invention provide a lightweight hyperspectral recognition method and system for sea surface oil spills based on a smooth activation function, specifically relating to a SqueezeNet sea surface oil spill lightweight hyperspectral recognition model based on a smooth activation function.

[0007] The technical solution is as follows: a lightweight hyperspectral recognition method for sea surface oil spills based on a smooth activation function, comprising:

[0008] S1: Conduct hyperspectral and thermal infrared oil spill data acquisition experiments in ideal scenarios and simulated real scenarios, obtain land-based and airborne oil spill data and real images at different times, use land-based measured data to perform oil-water separability analysis at different bands, and build a full-chain system for airborne oil spill identification and verification;

[0009] S2: Based on the measured data, a lightweight oil spill recognition model based on the SR-SqueezeNet smooth activation function was constructed. The experimental results of SR-SqueezeNet were compared with those of the unimproved squeeze network and semantic segmentation network. The optimal parameters of the lightweight oil spill recognition model SR-SqueezeNet were analyzed and found. The smooth activation function Smooth-ReLU was constructed, and the lightweight performance of the lightweight oil spill recognition model SR-SqueezeNet was comprehensively evaluated.

[0010] S3, airborne thermal infrared and high-resolution RGB images, verifies the optimal parameters of the model and analyzes the applicability of the model from the perspective of light and heat combination. The spatiotemporal transferability of the lightweight oil spill identification model SR-SqueezeNet is verified through airborne hyperspectral images acquired at different times.

[0011] In step S1, a hyperspectral and thermal infrared oil spill data acquisition test is conducted, including:

[0012] S101, field experiments and data acquisition;

[0013] S102, hyperspectral reflectance data preprocessing and analysis;

[0014] S103, hyperspectral airborne image preprocessing and dataset construction. The constructed dataset includes: training set, validation set, and test set with a ratio of 6:2:2.

[0015] In step S2, a SR-SqueezeNet lightweight oil spill recognition model based on a smooth activation function is constructed, including:

[0016] Adding a Flatten layer to the basic SqueezeNet network structure to obtain a pixel-level feature map, determine the category of each pixel, and ultimately achieve semantic segmentation.

[0017] The SqueezeNet basic network includes the Fire module, which consists of two layers: the squeeze layer and the expand layer. The squeeze layer is a convolutional layer with a 1×1 convolution kernel, and the expand layer is a convolutional layer with 1×1 and 3×3 convolution kernels. In the expand layer, the feature maps obtained by 1×1 and 3×3 are fused; the Fire Module uses 1×1 convolution to replace some 3×3 convolutions, reducing parameters and the number of input channels; after reducing the number of channels, convolution kernels of multiple sizes are used for calculations.

[0018] Furthermore, convolution kernels of multiple sizes are used for calculations, including: the input active image is convolved through the previous layer to obtain a set of feature maps, and then for each size of the convolution kernel, a block of the corresponding size is taken from the input feature map, and the block is element-wise multiplied with the convolution kernel; the product results at each position are accumulated to form a new eigenvalue; the results of all convolution kernels of different sizes are superimposed, and the generated feature map contains information from different spatial scales.

[0019] In step S2, a smooth activation function Smooth-ReLU is constructed, including:

[0020] The Fire Module uses the ReLU activation function, as shown in formula (2):

[0021] ReLU={max(0,x)} (2)

[0022] Where max(0,x) is the larger value compared to 0, and x is the neuron input;

[0023] For the ReLU activation function, there is neuron necrosis and the derivative does not exist at zero. The smooth activation function Smooth-ReLU is used to make the curve smooth and continuous at zero. The negative saturation region is designed to improve the robustness to noise. The smooth activation function Smooth-ReLU is shown in formula (3):

[0024]

[0025] Where, e x Perform an exponential function operation on the input of the neuron with the natural constant e as the base.

[0026] Furthermore, after constructing the smooth activation function Smooth-ReLU, the categorical_crossentropy used by the ReLU activation function is used as the loss function of the model. The loss function Loss is shown in formula (4):

[0027]

[0028] Where y i is the i-th element of the input vector, m is the number of categories, i is the i-th category, is the true label (0 or 1) of the i-th category.

[0029] Furthermore, the Flatten layer is added to the SqueezeNet basic network structure, including:

[0030] Remove the maximum pooling layer after the convolutional layer and add a Flatten layer to flatten the output of the convolutional layer and output it through the softmax function to avoid spatial dimensionality reduction of the feature map. The calculation of the softmax function is shown in formula (5):

[0031]

[0032] Where P(y|x) is the softmax value of the output, x is the neuron input, and y is the output vector. is a natural exponential operation on the input, k is the dimension of the input vector, j is the summation element, n is the number of classes in the multi-class classifier, and y i is the i-th element of the input vector, is a normalization term to ensure that the sum of all output values ​​of the function is 1 and each output value is in the range of (0, 1), thus forming a probability distribution. Stochastic gradient descent SGD is used as the optimization algorithm with a learning rate of 0.001.

[0033] In step S2, after constructing the smooth activation function Smooth-ReLU, the lightweight oil spill recognition model SR-SqueezeNe is lightweight optimized for airborne oil spill detection through network pruning or unstructured pruning. A comparative analysis of different activation functions and their application locations is performed, including:

[0034] Evaluation index analysis, based on the specific experiment to build a confusion matrix, select the overall classification accuracy OA, Kappa coefficient and F1-score three accuracy standards:

[0035] The first metric: FLOPs, which represents the computational effort of the model and is used to measure the complexity of the algorithm and model. The unit is GB.

[0036] The second metric is the number of parameters in the network, which is related to the size of the model and is usually measured in M.

[0037] The third metric: the time for model training and classification, which objectively reflects the computing speed of the model.

[0038] In step S3, the optimal parameters of the model are verified, including:

[0039] S301, experiment on the effect of spatial neighborhood size;

[0040] S302, impact analysis experiment of different training rounds;

[0041] S303, experiment on the impact of different activation functions.

[0042] Another object of the present invention is to provide a lightweight hyperspectral identification system for sea surface oil spills based on a smooth activation function, which implements the lightweight hyperspectral identification method for sea surface oil spills based on a smooth activation function. The system comprises:

[0043] The full-chain system construction module for airborne oil spill identification and verification is used to conduct hyperspectral and thermal infrared oil spill data acquisition experiments in ideal scenarios and simulated real-world scenarios, obtain land-based and airborne oil spill data and real images at different times, use land-based measured data to perform oil-water separability analysis at different bands, and build a full-chain system for airborne oil spill identification and verification;

[0044] A lightweight oil spill recognition model and smooth activation function construction module are used to build the SR-SqueezeNet lightweight oil spill recognition model based on the measured data. The experimental results of SR-SqueezeNet are compared with those of the unimproved squeeze network and semantic segmentation network. The optimal parameters of the lightweight oil spill recognition model SR-SqueezeNet are analyzed and found. The smooth activation function Smooth-ReLU is constructed, and the lightweight performance of the lightweight oil spill recognition model SR-SqueezeNet is comprehensively evaluated.

[0045] The experimental verification module uses airborne thermal infrared and high-resolution RGB images to verify the optimal parameters of the model and analyze the applicability of the model from the perspective of light and heat combination. The airborne hyperspectral images acquired at different times are used to verify the spatiotemporal transferability of the lightweight oil spill identification model SR-SqueezeNet.

[0046] Combining all the above technical solutions, the present invention has the following beneficial effects: the present invention uses the designed smooth activation function Smooth-ReLU to construct the SR-SqueezeNet sea surface oil spill hyperspectral recognition model, and conducts a series of experiments based on multi-dimensional airborne oil spill images obtained from field experiments. The experiments show that SR-SqueezeNet performs best in terms of extraction accuracy and model lightweighting. Compared with the traditional SqueezeNet, the recognition accuracy is improved by 1.92%, the number of parameters is reduced by 75.11%, and the model size is reduced from 26.46MB to 12.15MB. Therefore, SR-SqueezeNet meets the actual needs of airborne lightweight detection models and has high practicality.

[0047] This paper constructs the SR-SqueezeNet lightweight oil spill recognition model using the Smooth-ReLU activation function. Its expected benefits primarily lie in improving the efficiency and accuracy of marine oil spill detection. The model's lightweight nature allows it to significantly reduce computing resource consumption while maintaining high recognition performance. This translates to lower operating costs and faster real-time response for commercial applications. For example, in the field of drone monitoring, this can achieve longer operating times and greater coverage, thereby improving the timeliness and economic efficiency of emergency response to oil spill disasters.

[0048] As an innovative approach, SR-SqueezeNet fills the gap in the industry's need for lightweight, high-performance models for identifying oil spills. Compared to traditional recognition models, it not only improves recognition performance but also reduces recognition time. This represents a significant technological breakthrough both domestically and internationally, driving technological innovation and development in the field of remote sensing hyperspectral data. SR-SqueezeNet addresses the long-standing challenges faced by lightweight models in complex environments, such as ocean hyperspectral data. Previous models often struggled to strike a balance between accuracy and resource usage. SR-SqueezeNet achieves both high performance and low resource consumption, meeting the recognition needs of real-world scenarios. By breaking through conventional design concepts and technical bottlenecks, SR-SqueezeNet successfully overcomes skepticism and technical prejudice regarding the performance of lightweight models. Its emergence demonstrates that efficient oil spill detection is possible even in resource-constrained environments, broadening the application of deep learning in disaster monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure;

[0050] Figure 1 This is a flow chart of a lightweight hyperspectral identification method for sea surface oil spills based on a smoothed activation function provided by an embodiment of the present invention;

[0051] Figure 2 is a graph showing the average remote sensing reflectivity of crude oil and seawater provided by an embodiment of the present invention;

[0052] Figure 3 is a processed airborne hyperspectral image provided by an embodiment of the present invention;

[0053] Figure 4 This is an oil-water distribution label diagram provided by an embodiment of the present invention;

[0054] Figure 5 This is a diagram of the SqueezeNet architecture used in the present invention;

[0055] Figure 6 A schematic diagram of the Fire module in the SqueezeNet architecture used in the present invention;

[0056] Figure 7 This is the ReLU activation function curve of the existing technology;

[0057] Figure 8 The improved Smooth-ReLU curve of the smooth activation function of the present invention;

[0058] Figure 9The present invention shows the influence of the size of the spatial neighborhood on the accuracy and time of the oil spill OA and Kappa coefficient;

[0059] Figure 10 The figure shows the time diagram of one epoch training and global classification of samples under six different spatial neighborhood sizes: 1×1, 3×3, 5×5, 7×7, 9×9, and 11×11;

[0060] Figure 11 The image shows the classification results under six different spatial neighborhood sizes: 1×1, 3×3, 5×5, 7×7, 9×9, and 11×11. (a) is the 1×1 classification result, (b) is the 3×3 classification result, (c) is the 5×5 classification result, (d) is the 7×7 classification result, (e) is the 9×9 classification result, and (f) is the 11×11 classification result.

[0061] Figure 12 The present invention shows the effect of increasing the number of rounds from 1 on the oil spill classification accuracy;

[0062] Figure 13 This is an experimental diagram showing the effect of different numbers of training rounds on the overall training time of the present invention;

[0063] Figure 14 The present invention shows the result images of global oil spill classification after different numbers of training rounds, where (a) is the 1-epoch image, (b) is the 10-epoch image, (c) is the 20-epoch image, (d) is the 30-epoch image, (e) is the 40-epoch image, and (f) is the 50-epoch image.

[0064] Figure 15 The oil spill classification images of the SqueezeNet model using different activation functions of the present invention, where (a) is the Smooth-ReLU activation function diagram, (b) is the ReLU activation function diagram, (c) is the ReLU6 activation function diagram, and (d) is the Leaky-ReLU activation function diagram;

[0065] Figure 16 The present invention shows the influence of applying four activation functions to the shallow Fire module, deep Fire module and all Fire module convolutional layers on the oil spill classification accuracy of the SqueezeNet model;

[0066] Figure 17 The present invention shows the effect of applying four activation functions to the convolutional layers of the Fire module at different depths on the single training time of the SqueezeNet model;

[0067] Figure 18 Figures 2 and 3 show the oil spill recognition results of different models of the present invention, where (a) is the SqueezeNet model, (b) is the SR-SqueezeNet model, (c) is the FCN model, (d) is the SegNet model, (e) is the U-Net model, and (f) is the MobileNet-V3 model.

[0068] Figure 19 This is the airborne 4K real image acquired synchronously by the present invention;

[0069] Figure 20 This is the airborne thermal infrared image of the present invention;

[0070] Figure 21 The following are the classification results of the application of different dimensional models of the present invention, where (a) is the airborne hyperspectral image recognition result, (b) is the airborne thermal infrared image recognition result, and (c) is the airborne 4K image recognition result;

[0071] Figure 22 This is the oil spill airborne hyperspectral image ROI-2 after image processing and enhancement of the present invention;

[0072] Figure 23 This is the spatial distribution diagram of ROI-2 samples of the present invention;

[0073] Figure 24 These are the recognition result diagrams of different models of the present invention, where (a) is the SqueezeNet model diagram, (b) is the SR-SqueezeNet model diagram, (c) is the FCN model diagram, (d) is the SegNet model diagram, (a) is the U-Net model diagram, and (a) is the MobileNet-V3 model diagram. DETAILED DESCRIPTION

[0074] To make the above-mentioned objects, features, and advantages of the present invention more readily apparent, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. The following description sets forth numerous specific details to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art may make similar modifications without departing from the scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0075] The innovation of the present invention lies in: the present invention adopts a smooth activation function Smooth-ReLU, which maintains the nonlinear characteristics of ReLU while improving the continuity and stability of the output, greatly improving the performance and generalization ability of the model. Through optimized design, the SR-SqueezeNet model achieves lightweight while maintaining high recognition accuracy, effectively reducing the number of parameters and model volume, and alleviating the computational burden of resource-limited equipment such as drones. Experimental results show that SR-SqueezeNet surpasses similar models in recognition accuracy and efficiency, demonstrating its high efficiency advantage in real-time sea surface oil spill monitoring. In addition, it also demonstrates extremely strong multi-dimensional adaptability, and can stably perform recognition on both thermal infrared imaging and 4K high-definition airborne imaging, as well as dynamic airborne hyperspectral data. Finally, it has been verified in practice that the model meets the actual needs of airborne and satellite-based oil spill monitoring, proving its practical value and broad prospects in real-world applications.

[0076] Example 1, in order to solve the problem that the current airborne hyperspectral image data set is small, the airborne hyperspectral oil spill recognition model has many parameters and the recognition time is long, and at the same time to achieve the multi-dimensional complementary advantages of oil spill recognition in the field of photothermal, such as Figure 1 As shown, the lightweight hyperspectral recognition method for sea surface oil spills based on smooth activation function provided by the embodiment of the present invention mainly includes:

[0077] S1: Conduct hyperspectral and thermal infrared oil spill data acquisition experiments in ideal scenarios and simulated real scenarios, obtain land-based and airborne oil spill data and real images at different times, use land-based measured data to perform oil-water separability analysis at different bands, and build a full-chain system for airborne oil spill identification and verification;

[0078] S2: Based on the measured data, a lightweight oil spill recognition model based on the SR-SqueezeNet smooth activation function was constructed. The experimental results of SR-SqueezeNet were compared with those of the unimproved squeeze network and semantic segmentation network. The optimal parameters of the lightweight oil spill recognition model SR-SqueezeNet were analyzed and found. The smooth activation function Smooth-ReLU was constructed, and the lightweight performance of the lightweight oil spill recognition model SR-SqueezeNet was comprehensively evaluated.

[0079] S3, airborne thermal infrared and high-resolution RGB images, verifies the optimal parameters of the model and analyzes the applicability of the model from the perspective of light and heat combination. The spatiotemporal transferability of the lightweight oil spill identification model SR-SqueezeNet is verified through airborne hyperspectral images acquired at different times.

[0080] Exemplarily, step S1 includes data acquisition and processing.

[0081] S101, field experiment and data acquisition. The present invention conducted a land-based and airborne hyperspectral oil spill data acquisition test under ideal scenarios and simulated real scenarios. The test lasted for two days and was located near the coast of a certain district in a certain city. The outdoor temperature was 23-30℃, the weather was clear, and there was a breeze of 2-3 levels. Considering that the oil leaked at sea was mainly crude oil, the experimental oil used in the test to simulate the oil spill at sea was crude oil produced in a certain oil field, which was black in color and had a density of 0.882 / (g·mL -1 In an ideal outdoor oil spill scenario, using sunlight as the natural light source, a black matte PVC pool was filled with coastal seawater from a certain city. 2L of crude oil was slowly poured into it, allowing it to spread evenly on the water surface, forming an oil film of approximately 1.5mm.

[0082] The experiment first acquired land-based hyperspectral data under ideal conditions. Using the ASD FieldSpec4 spectrometer and a standard plate, the hyperspectral remote sensing reflectance of seawater and crude oil was measured at different times and solar intensities. The ASD FieldSpec4 spectrometer was used to measure the spectral radiance of the oil and seawater vertically downward from approximately 10 cm above the water surface. The spectral radiance of the Lambertian standard plate and skylight was also measured simultaneously. During the measurements, the lighting was kept stable and free from shadows and strong reflectors.

[0083] Subsequently, a UAV-mounted Cubert S185 hyperspectral imager was used to conduct vertical observations above the oil spill detection pool, acquiring airborne hyperspectral images at an altitude of approximately 40 meters. The Cubert S185 sensor parameters are shown in Table 1. Simultaneous observations were conducted using a UAV equipped with a 4K and thermal infrared sensor. Six sets of UAV hyperspectral data (six time periods) were collected on-site, each containing hyperspectral images of the oil spill acquired at a different altitude. The data acquired by the imaging spectrometer includes raw measured spectral data and reference plate and dark current calibration data. The raw spectral data collected by the hyperspectral imager is DN value data, which needs to be converted into radiance data.

[0084] Table 1 Parameters of Cubert S185 sensor

[0085] parameter spectral range Number of bands Spectral resolution Spatial resolution Observation angle CubertS185 sensor 450-950 nm 126 4 nanometers 0.016 meters 90 degrees

[0086] S102, using land-based measured data to perform oil-water separability analysis at different bands, including: hyperspectral reflectance data preprocessing and analysis.

[0087] During the airborne hyperspectral image acquisition period, the experiment collected spectral data of background seawater, crude oil, skylight, and standard plates for response analysis. The oil-water remote sensing reflectance was calculated using formula (1) based on the water surface radiance, incident irradiance, and skylight:

[0088]

[0089] Where R rs (λ) is the oil-water remote sensing reflectivity, L w (λ) is the oil reflectivity, E s (λ) is the water reflectivity, L t (λ) is the sea surface radiance; ρ is the water-air interface transmittance, which does not change with the spectrum. Considering the stability of the water body, ρ is taken as 0.01 during data processing; L sky (λ) is the sky radiance, L p (λ) is the irradiance of the standard plate, ρ p (λ) is the reflectivity of the standard plate;

[0090] The function of formula (1) is to process the acquired radiance data into reflectance data that is easier to analyze and more intuitive after drawing.

[0091] For synchronously acquired ground feature spectral data, the spectral resolution is downsampled to obtain spectral reflectance data consistent with the S185 sensor band. Data affected by the strong absorption of water at 1400nm and 1900nm, and data with irregular oscillations at 2400nm due to the edge effect of the photosensitive device are removed. The average remote sensing reflectance curves of crude oil and seawater are obtained, as shown in Figure 2. Figure 2 As shown in the figure. The spectral responses of crude oil and seawater in different spectral ranges are different, which is related to their absorption and scattering properties. In the visible light band, the spectral curves of crude oil and seawater are very similar. The reflectivity of crude oil is significantly lower than that of seawater, with a clear reflection peak around 480nm. In the near-infrared and short-wave infrared bands, the incident light absorption is strong and the reflectivity is low. The spectral reflectivity of crude oil is low, especially in the near-infrared band, where the reflectivity of crude oil is higher than that of seawater. This is because pure seawater is almost a "blackbody" in the near-infrared band. Therefore, in the spectral range of 850-2500nm, the reflectivity of the two groups of seawater is very low, almost zero, and the reflectivity of the two groups of crude oil is higher than that of seawater. Based on the above feature analysis of land-based hyperspectral reflectivity data, it is feasible to conduct feature screening of airborne hyperspectral reflectivity data and use its spatial spectrum characteristics for oil spill identification.

[0092] For example, Figure 2These are two sets of hyperspectral reflectance data for crude oil and seawater in different water types, consistent with the S185 sensor band. These curves were obtained by excluding data from the strong absorption regions at 1400nm and 1900nm, as well as data from irregular oscillations at 2400nm due to the edge effect of the photosensitive device. Based on the optical properties of water, seawater can be divided into two categories: Class I water and Class II water. The optical properties of Class I water are primarily determined by phytoplankton and its appendages, while the content of other suspended matter is relatively low. Typical Class I waters are pelagic waters. Class II waters contain more suspended matter and soluble organic matter, and their spectra have reflection peaks in the visible light band. The impact of different background water bodies should be considered in the remote sensing identification of marine oil spills. To bridge the gap between the changes in oil spill characteristics in different water bodies, this paper analyzes the differences in positive and negative contrast in different water bodies. In the visible light band, the spectral curves of crude oil and seawater are similar, and the reflectance of crude oil is significantly lower than that of seawater. The seawater of Class II water bodies is affected by suspended matter and soluble organic matter, and has a clear reflection peak at about 480nm. In the near-infrared and short-wave infrared bands, the reflectivity of seawater is low (close to zero), and the reflectivity of crude oil is higher than that of seawater. Since the background seawater of the oil spill observation experiment was collected in the Qingdao waters, a typical Class II water body, the ASD data against the background of Class II water bodies was used to select characteristic bands for the S185 airborne data. Based on the description of land hyperspectral reflectance data, it is feasible to perform feature screening on airborne hyperspectral reflectance data and use its spatial and spectral characteristics for oil spill identification.

[0093] S103, building a full-chain system for airborne oil spill identification and verification, including: hyperspectral airborne image preprocessing and dataset construction.

[0094] The experimental data selected are airborne hyperspectral data obtained by the M600 Pro six-rotor drone equipped with a Cubert S185 frame-type hyperspectral imager in a field experiment. The flight altitude is about 15m, the image size is 1000×1000 pixels, the spectral resolution is 8nm, and the number of bands is 126. The output image is aligned and cropped, and the final image size used to identify the test oil spill is 332×330. The airborne hyperspectral image is processed and enhanced by changing the brightness value of the pixel to increase the contrast of the entire or local image and improve the image quality. Taking into account the reflection peak of crude oil at 480nm, the 6th, 16th, and 25th bands (corresponding to 470nm, 510nm, and 546nm, respectively) are selected as the reference bands for the color synthesis of airborne hyperspectral images. The processed airborne hyperspectral images are as follows Figure 3 shown.

[0095] For example, changing pixel brightness to increase overall or local contrast in an image involves first performing weighted averaging, stretching, and thresholding on the brightness of each pixel. A global contrast enhancement method is then used to apply a uniform gain to all pixels. Local contrast adjustments are then performed on low-saturation areas, such as by applying a local mean filter followed by brightness adjustments.

[0096] The pixels in the sample area shown in the figure are selected as training samples and divided into three categories: crude oil, fence and seawater. In order to ensure the accuracy and reliability of oil spill identification, the results of visual interpretation by experts at the test site are referred to when making oil and water distribution labels. The distribution status of crude oil and seawater in the pool is judged based on the 4K images and thermal infrared images taken simultaneously at the same time of the day. The verification data is outlined and the visually interpreted oil and water distribution labels are made. The labels can be supported by the measured ASD land-based reflectivity data. The oil and water distribution labels are the sample spatial distribution as shown in the figure. Figure 4 shown.

[0097] For example, using 4K and thermal infrared images captured simultaneously on the same day to determine the distribution of crude oil and seawater in a pool involves injecting seawater from Qingdao's coastal waters into a black colloidal pool, followed by the gradual addition of crude oil. Crude oil, being denser than seawater, initially sinks into the water, but after a while, some rises and spreads on the water surface, forming a visible oil film. Using the synchronized 4K live and thermal infrared images, field experts were able to clearly determine the shape and location of the oil film, perform visual interpretation, and create oil-water distribution labels.

[0098] Implementing pixel-by-pixel image classification requires collecting a large dataset. This paper constructs spatial blocks and their labels pixel by pixel, randomly dividing the sample area and the corresponding labels into pixel blocks of radius r. To obtain sufficient sample data and avoid overfitting during model training, horizontal, vertical, and diagonal flipping methods are also used to enhance the samples. Remote sensing images from different regions have relatively large differences in average brightness and pixel value distribution. Therefore, before deep learning training, the extracted samples are normalized using the maximum and minimum normalization method to ensure that the sample data has as similar a distribution as possible, accelerate the convergence of the training network, and improve the accuracy of model detection. The ratio of the training set, validation set, and test set is 6:2:2.

[0099] It can be understood that the present invention constructs spatial blocks and their labels pixel by pixel, including: performing semantic segmentation on the oil spill hyperspectral image, dividing the image into pixel blocks of fixed size or variable size, and for each pixel in each spatial block, using the extracted features for classification, and predicting its corresponding pixel label as the corresponding category. Smaller spatial blocks, such as 3x3 or 5x5, pay more attention to local features and help capture subtle structures and boundaries in the image. Therefore, in tasks that require precise boundary detection, small blocks may be more effective. This step is usually achieved through a fully connected layer. Finally, all pixel prediction results are merged to generate the final pixel-level semantic segmentation map. Spatial block size plays a key role in semantic segmentation, which directly affects the detail capture ability and efficiency of the model.

[0100] This sample augmentation operation is used to obtain sufficient training samples, improving the model's understanding of the input data through various transformation techniques, such as horizontal flipping, vertical flipping, and diagonal flipping. These transformations increase image diversity and enable the model to better recognize changes in oblique angles. The flipped samples are then used in deep learning training to simulate more possible scenarios, prevent the model from overfitting to a single angle or orientation, and improve its generalization ability.

[0101] The method of normalizing the extracted samples using the maximum and minimum normalization method before deep learning training includes: using the maximum and minimum normalization method to process the extracted samples, scaling the input features to the range of 0 to 1, improving the model convergence speed and optimizing weight initialization. Considering the lightweight design of the model, the present invention uses local normalization. Compared with traditional normalization methods, interval normalization only considers a small area near each pixel, which helps to retain local structural information.

[0102] For example, in step S2, constructing a lightweight oil spill identification model SR-SqueezeNet based on the measured data obtained in the experiment in step S1 includes:

[0103] SqueezeNet is composed of several Fire modules combined with convolutional layers, downsampling layers, and fully connected layers in the convolutional network. The SqueezeNet architecture used in this invention is as follows: Figure 5As shown. In the input layer, a 5×5 pixel block of the surrounding pixels is extracted for each pixel. Each pixel block has spectral feature information of 126 bands. After standardized preprocessing, the data is sent to other layers. The convolution layer can extract different features from the input data. The role of the downsampling layer in the convolutional neural network is mainly to retain the main information while reducing the amount of calculation and increasing the speed of the model. The fully connected layer maps the features to the label space of the sample and highly purifies the features. In order to realize the semantic segmentation of airborne hyperspectral images, the present invention adds a Flatten layer to the basic network structure of SqueezeNet to obtain a pixel-level feature map, judge the category of each pixel point, and finally realize the function of semantic segmentation.

[0104] The core of SqueezeNet is the Fire module, which consists of two layers: the squeeze layer and the expand layer. Figure 6 The Fire module diagram shown in the figure shows a squeeze layer with a 1×1 convolution kernel, and an expand layer with both 1×1 and 3×3 convolution kernels. In the expand layer, the 1×1 and 3×3 feature maps are fused. The Fire Module primarily optimizes the network structure, replacing some 3×3 convolutions with 1×1 convolutions, reducing the number of parameters to 1 / 9 of the original number while also reducing the number of input channels. After reducing the number of channels, convolution kernels of multiple sizes are used for calculations to retain more information, improve classification accuracy, reduce network parameters, and enhance network performance. These optimization strategies result in fewer SqueezeNet parameters and smaller network memory usage, thereby maximizing computational speed without significantly compromising model accuracy.

[0105] Exemplarily, calculations using convolution kernels of multiple sizes involve convolution of the input active image through the previous layer to produce a set of feature maps. For each convolution kernel size, a correspondingly sized block is taken from the input feature map and element-wise multiplied by the convolution kernel. The product of each position is accumulated to form a new feature value. Finally, the results of all convolution kernels of different sizes are superimposed. The resulting feature map contains information from different spatial scales, helping to capture richer features.

[0106] In addition, choosing a suitable activation function is very important for neural networks. The activation function is a function that runs on the neurons of the neural network. It is responsible for mapping the input of the neuron to the output end, with the purpose of helping the network learn complex patterns in the data. Using the activation function can introduce nonlinear factors to the neurons, so that the neural network can arbitrarily approximate any nonlinear function, making the deep neural network more expressive. Choosing a suitable activation function is crucial for neural networks because it affects the output of neurons and the learning performance of the entire network. The Fire Module uses the ReLU activation function by default. Its formula is shown in formula (2). The ReLU activation function curve is as follows: Figure 7 As shown, in Figure 7 and Figure 8 In the equation, the horizontal coordinate x is the neuron input x, and the vertical coordinate y is the output of the activation function f(x). Due to the problems of neuron necrosis and the non-existence of the derivative at zero, the present invention innovatively improves the ReLU activation function and proposes a smooth activation function Smooth-ReLU, which has a smooth and continuous curve at zero and designs a negative saturation region to make the model more robust to noise. The formula is shown in equation (3). The smooth activation function Smooth-ReLU curve is as follows: Figure 8 shown.

[0107] ReLU={max(0,x)} (2)

[0108] Where max(0,x) is the larger value compared to 0, and x is the neuron input;

[0109]

[0110] Where, e x Perform an exponential function operation on the input of the neuron with the natural constant e as the base.

[0111] As you can understand, the improved smooth activation function Smooth-ReLU solves the neuron necrosis problem of traditional ReLU, and the output average is close to 0 and centered on 0. It reduces the impact of bias offset, making the normal gradient closer to the unit natural gradient, thereby accelerating learning of the mean toward zero. At the same time, when the input is in the negative region, the model will reach saturation with smaller inputs, thereby reducing the information propagated forward.

[0112] The improved smooth activation function Smooth-ReLU solves the neuron necrosis problem of traditional ReLU. The output average is close to 0 and centered on 0. It reduces the impact of bias offset, making the normal gradient closer to the unit natural gradient, thereby accelerating learning of the mean toward zero. At the same time, when the input is in the negative region, the model will reach saturation with smaller inputs, thereby reducing the information propagated forward.

[0113] The present invention uses categorical_crossentropy, which is often used with the ReLU activation function, as the loss function of the model. The formula of the loss function is shown in the following formula (4):

[0114]

[0115] Where y i is the i-th element of the input vector, m is the number of categories, i is the i-th category, is the true label (0 or 1) of the i-th category.

[0116] The present invention innovatively proposes that in order to avoid spatial dimensionality reduction of the feature map, the maximum pooling layer after the convolutional layer is removed, and a Flatten layer is added to flatten the output of the convolutional layer and output it through the softmax function.

[0117] The calculation formula of the Softmax function is shown in the following formula (5):

[0118]

[0119] Where P(y|x) is the softmax value of the output, x is the neuron input, and y is the output vector. is a natural exponential operation on the input, k is the dimension of the input vector, j is the summation element, n is the number of classes in the multi-class classifier, and y i is the i-th element of the input vector, is a normalization term to ensure that the sum of all output values ​​of the function is 1 and each output value is in the range of (0, 1), thus forming a probability distribution. Stochastic gradient descent SGD is used as the optimization algorithm with a learning rate of 0.001.

[0120] At the same time, in order to face the airborne oil spill detection, the present invention is committed to making the model more lightweight, and optimizes the model from two perspectives, namely reducing the number of learnable parameters and reducing the computational complexity of the entire network. Network pruning is one of the main technologies for network compression. It is an important technology for reducing memory size and bandwidth, and can remove redundant parameters or neurons that do not significantly contribute to the accuracy of the results. The network pruning of SqueezeNet is completed in the following steps: First, the connectivity between layers is learned through normal network training. Next, connections with small weights are pruned, and all connections with weights below the threshold will be deleted from the network. Finally, the network is retrained to learn the final weights of the remaining sparse connections.

[0121] SqueezeNet employs unstructured pruning, using a splicing function to mask weights. Its formulas are shown in Equations (6) and (7). Its advantages include a simple pruning algorithm, high model compression ratio, no drastic changes in weight values, and the ability to integrate the pruning process with retraining. Backpropagation is written as Δw, where h(w) gradually reduces unnecessary weights to zero. Hyperparameters a and b control the strength of the threshold, with a = b corresponding to a typical binary mask.

[0122]

[0123] In the formula, μ is the weight coefficient, is the partial derivative of the input value, The pruned SqueezeNet network parameters are shown in Table 2. After pruning, the network parameters are reduced to approximately one-quarter of their original values, demonstrating a significant lightweighting effect. Pruning reduces the computational effort required for model training and testing, making each step faster. It also reduces the model file size, facilitating model storage and transfer. With fewer learnable parameters, the network consumes less video memory. Due to its lightweight nature, SqueezeNet can be widely applied to deep learning models for oil spill detection, promoting the development of on-orbit data processing.

[0124] Table 2 SqueezeNet network parameters (r = 2, Output size = 5 × 5 × 126)

[0125]

[0126]

[0127] Exemplarily, the comparative analysis of different activation functions and their application locations includes:

[0128] The comparison methods selected by the present invention are the classic semantic segmentation neural networks FCN, SegNet, U-Net and the lightweight neural network MobileNet. FCN (Fully Convolutional Networks) is a framework for image semantic segmentation proposed by Jonathan Long et al. in 2015, which can recover the category to which each pixel belongs from abstract features. SegNet was proposed by Cambridge in 2016. It adopts a symmetrical structure of left and right network layers of encoder-decoder. In the decoder, the reduced feature map is sampled and convolved to improve the geometric shape of objects in the image. U-Net was born to solve the problem of biomedical image segmentation. Its network structure was first proposed by Ronneberger et al. in 2015. The core idea of ​​this semantic segmentation model is to introduce jump connections, which greatly improves the accuracy of image segmentation. FCN, SegNet and U-Net are commonly used traditional semantic segmentation models. Performance testing under the same data set and parameters can effectively evaluate the semantic segmentation ability of the SR-SqueezeNet model. The MobileNet series is a lightweight neural convolutional network specifically designed for mobile or embedded devices. Its core concept is the use of depthwise separable convolution, which significantly reduces parameters and computational complexity while maintaining near-perfect accuracy. Comparing the proposed SR-SqueezeNet network with the MobileNet model allows us to assess the model's lightweightness.

[0129] Exemplarily, performing comparative analysis of different activation functions and their application locations further includes: evaluation indicators.

[0130] Selecting appropriate accuracy evaluation indicators to evaluate and analyze the classification results of hyperspectral remote sensing images is an important part of the result analysis. This paper constructs a confusion matrix based on specific experiments and selects three accuracy standards: overall classification accuracy (OA), Kappa coefficient, and F1-score. Overall classification accuracy refers to the probability that the classification results of the test sample are consistent with the label data. Its formula is shown in Equation (8).

[0131]

[0132] Where OA is the accuracy, TP is the number of correctly identified positive samples, TN is the number of correctly identified negative samples, FN is the number of incorrectly identified negative samples, FP is the number of incorrectly identified positive samples, and N is the total number of samples (pixels).

[0133] The Kappa coefficient is an indicator that can more comprehensively express classification accuracy. Its formula is shown in Equation (9). The overall classification accuracy can indicate the general accuracy of the classification results, but it does not take into account the situation of a specific category. The combination of the Kappa coefficient and OA can provide a more comprehensive and objective analysis of classification accuracy.

[0134]

[0135] In the formula, kappa is a statistic for measuring the consistency of classification problems, r is the number of rows in the confusion matrix, and x ii is the number of correctly classified samples of oil type, N is the total number of samples (pixels), x i+ is the sum of the observed samples in row i, x +i is the total number of observation samples in column i;

[0136] F1-score is the harmonic mean of precision and recall, which is used to comprehensively consider the performance of the classifier. Its formula is shown in the following equation (10).

[0137]

[0138] Where Precision is the precision rate and Recall is the recall rate;

[0139] For the evaluation of model lightweighting, the present invention selects three main evaluation indicators. The first metric is the number of floating point operations (FLOPs), which represents the computational amount of the model and can be used to measure the complexity of the algorithm and model. The unit is usually G. The second metric is the number of parameters in the network, which is related to the model size. The unit is usually M. The last one is the time for model training and classification, which can objectively reflect the calculation speed of the model. In actual model calculation, in addition to the above parameters, network architecture information and optimizer information are also included.

[0140] Exemplarily, in step S3, it specifically includes:

[0141] S301, the influence of spatial neighborhood size.

[0142] The experimental results were generated on a personal computer equipped with an Intel(R) Core(TM) i7 (180GHz) processor and an Nvidia GeForce MX250 graphics card. The image is a 1000×1000 pixel airborne hyperspectral image with 126 spectral bands covering the spectral range of 450–950 nm. To determine the optimal spatial neighborhood for the input data, we followed the principle of controlling variables and experimented with the following spatial neighborhood sizes: 1×1, 3×3, 5×5, 7×7, 9×9, and 11×11.

[0143] Through experiments, different spatial neighborhoods have a significant impact on the accuracy and time of oil spill classification. Figure 9 The following figure shows the effect of the spatial neighborhood size on the accuracy and time of the oil spill OA and Kappa coefficient. For the training samples, as the spatial scale increases from 1×1 to 5×5, the oil spill OA and Kappa coefficient show an upward trend, with a maximum accuracy of 98.2% and a maximum Kappa coefficient of 0.96. As the spatial scale increases from 5×5 to 11×11, the overall classification accuracy OA and Kappa coefficient show a downward trend. This indicates that appropriately increasing the spatial scale can enhance the spatial information of oil spill identification and play a positive role in improving classification accuracy. However, the spatial neighborhood should not be too large. When the spatial neighborhood is very large, such as 9×9 and 11×11, the excessive redundant spatial information will disrupt the effective classification of the model, resulting in a decrease in accuracy.

[0144] Figure 10 The time it takes to perform one epoch training and global classification on samples with six different spatial neighborhood sizes: 1×1, 3×3, 5×5, 7×7, 9×9, and 11×11. Figure 11 As can be seen from the figure, as the size of the spatial neighborhood gradually increases, both the model training time and the global classification time increase significantly. When the sample spatial neighborhood is 1×1, the training time per epoch is only 3 seconds, and the global classification time is 11 seconds. When the sample spatial neighborhood is increased to 5×5, the training time per epoch is 16 seconds, and the global classification time is 53 seconds. When the sample spatial neighborhood is 11×11, the training time per epoch is 195 seconds, and the global classification time reaches 360 seconds. Therefore, the size of the sample spatial neighborhood is directly proportional to the model training time and global classification time. The larger the sample spatial neighborhood, the higher the time cost.

[0145] Figure 11Classification results are presented for six different spatial neighborhood sizes: 1×1, 3×3, 5×5, 7×7, 9×9, and 11×11. Taking into account the relationship between the sample spatial neighborhood, overall classification accuracy, and time cost, the model was selected as the optimal spatial scale, considering that the 5×5 spatial neighborhood achieved the highest overall classification accuracy, the largest Kappa coefficient, and short single-epoch training time and global classification time, both within an acceptable range. This scale not only meets the requirements of airborne imagery for real-time and rapid oil spill detection, but also meets the requirements of a lightweight model for a reduced sample size.

[0146] S302, the impact of different training rounds.

[0147] Similarly, to ensure the best experimental results, this paper follows the principle of controlling variables. Under the premise that the sample space neighborhood is set to 5×5, different numbers of training rounds are designed, namely 1, 10, 20, 30, 40 and 50 times, to explore the impact of training rounds on the model classification accuracy.

[0148] Figure 12 The effect of increasing the number of rounds from 1 on the oil spill classification accuracy is shown. Figure 12 It can be seen that when the number of training rounds increases, the oil spill classification accuracy increases accordingly, with the lowest being 79.77%. When the number of training rounds increases to 50, the sample classification accuracy reaches the highest, with the highest being 97.64%. This shows that increasing the number of training rounds plays a positive role in improving the global classification accuracy. However, when the number of training rounds starts at 20, the global classification accuracy improves relatively slowly. Excessive training causes the learning ability of the model to tend to saturate, and the accuracy cannot be further significantly improved. The present invention also explores the impact of different numbers of training rounds on the overall training time. The experimental results are as follows: Figure 13 As shown. Figure 13 It can be clearly seen that as the number of training rounds increases, the training time gradually increases. When the spatial neighborhood is 5×5, the single epoch time is about 21 seconds, and the overall training time is approximately the product of the single epoch time and the number of training rounds. More training rounds will prolong the model training and classification time, which does not meet the requirements of the onboard model for fast and real-time performance.

[0149] Figure 14 The images show the global classification results of oil spills after different numbers of training rounds. Since increasing the number of training rounds positively improves global classification accuracy, global classification accuracy stabilized at around 97% with 20 and 30 training rounds, and did not improve significantly thereafter. While accuracy was slightly lower with 20 training rounds, the training time was reduced by nearly two minutes compared to 30 training rounds. Considering the need for rapid, lightweight airborne oil spill detection, 20 training rounds was ultimately selected as the standard number of training rounds.

[0150] S303, the impact of different activation functions.

[0151] Activation functions play an important role in the recognition accuracy and parameter iteration speed of deep neural networks. To determine the impact of different activation functions on model performance, this paper follows the principle of controlling variables and uses different activation functions for the model. Under the conditions of a 5×5 spatial neighborhood and 20 training rounds, the activation effects of the Smooth-ReLU activation function designed by the present invention are compared with those of the ReLU activation function, the ReLU6 activation function, and the Leaky ReLU activation function. The effects of using these functions in the Fire module convolutional layers of different depths in the SqueezeNet lightweight model are also explored.

[0152] ReLU (Rectified Linear Unit) is a commonly used neural network activation function, widely used in PyTorch. Its function characteristic is to return the value when the input value is greater than zero, and return zero when the input value is less than zero. Due to its simplicity and effectiveness, ReLU has become the default activation function used throughout deep learning. However, the output of the ReLU activation function is not zero-mean, and there is a Dead ReLU Problem (neuron necrosis phenomenon) in the negative region. That is, when x<0, the gradient of the neuron is judged to be 0, and the gradient of the related neuron is always 0, and it no longer responds to any data, resulting in the corresponding parameters never being updated. Therefore, the existing technology is committed to proposing new activation functions to replace ReLU and solve the neuron necrosis phenomenon. For example, the leaky rectified linear unit (Leaky ReLU) initializes the neuron with a value close to zero, so that the input value is more likely to be activated rather than dead in the negative region. Its formula is shown in formula (11). The ReLU6 activation function is a commonly used improved form of the ReLU function, as shown in formula (12). This function limits the input value to between 0 and 6. Values ​​less than 0 are truncated to 0, and values ​​greater than 6 are truncated to 6. The ReLU6 activation function can help improve the nonlinear expression ability and anti-saturation resistance of the model. However, the above-mentioned improved ReLU activation functions also have the problem that the derivative at zero point does not exist, or the activation effect is unstable between different models and data sets.

[0153]

[0154] ReLU6={max(x,0),6} (12)

[0155] Table 3 shows the oil spill detection results of the SqueezeNet model using different activation functions. As shown in the table, different activation functions have a significant impact on global classification accuracy and vary in single-epoch training time. The Smooth-ReLU activation function achieves higher global classification accuracy than the ReLU and its improved activation functions. This is because the Smooth-ReLU derivative exists at zero, resulting in a smoother and more continuous curve. This improves the model's activation performance in the negative region, leading to higher global oil spill classification accuracy. The experiment also compared single-epoch training times, showing that the training times of different activation functions are relatively similar. The ReLU6 activation function has the shortest training time. Because its output is truncated when it exceeds 6, it is lightweight. Smooth-ReLU is second. Because its output average is close to zero and centered around zero, it accelerates learning towards zero. Furthermore, when the input is in the negative region, the model reaches saturation with small inputs, reducing the forward propagation information and enabling faster training. The Leaky ReLU and ReLU activation functions, on the other hand, train more slowly and have lower accuracy. Different activation functions are applied to the oil spill classification images of the SqueezeNet model. Figure 15 As shown, from Figure 15 It can be seen that Smooth-ReLU and ReLU have better overall oil spill recognition effects, followed by Leaky-ReLU, while ReLU6 has the worst recognition effect.

[0156] Table 3 Effects of different activation functions (Output size = 5 × 5, Epochs = 20)

[0157] Activation function name Accuracy (%) Kappa coefficient <![CDATA[F1 score]]> Training time (s) Smooth-ReLU 97.83 0.97 95.39 14 ReLU 96.53 0.97 94.16 16 ReLU6 93.17 0.92 91.05 13 Leaky-ReLU 96.15 0.95 94.73 17

[0158] In order to achieve the best activation effect of the model, the present invention also compares the difference in global classification accuracy and single training time when different activation functions are applied to convolutional layers of different depths. The experimental results are as follows Figure 16-17 As shown, Figure 16This study demonstrates the impact of applying four activation functions to the shallow Fire module, deep Fire module, and all Fire module convolutional layers on the oil spill classification accuracy of the SqueezeNet model. The improved activation functions exhibited varying activation effects when applied to Fire modules of varying depths. Overall, the improved activation functions generally achieved better results when applied to deeper convolutional layers. Global classification accuracy reached its highest level of 98.74% when Smooth-ReLU was applied only to the deep Fire module. The lowest global classification accuracy was achieved when the ReLU6 activation function was applied to all Fire module convolutional layers, achieving only 93.18%. Leaky-ReLU achieved similar classification accuracy across convolutional layers of varying depths. Figure 17 This study demonstrates the impact of applying four activation functions to convolutional layers in the Fire module at varying depths on the single-shot training time of the SqueezeNet model. Overall, the Smooth-ReLU and ReLU6 activation functions have relatively short single-shot training times, followed by ReLU. Leaky-ReLU has the longest single-shot training time, which is not significantly correlated with the location of the activation function. Applying Smooth-ReLU to a deep Fire module achieves the shortest single-shot training time of just 12 seconds. Experiments demonstrate that Smooth-ReLU is smoother and more continuous at zero in deeper convolutional layers, resulting in better activation performance in negative regions, improved global classification accuracy, and shorter training time.

[0159] Simulation experiment.

[0160] (1) Performance comparison of different models.

[0161] To evaluate the model's lightweightness, this paper conducted oil spill identification experiments comparing SR-SqueezeNet with traditional semantic segmentation networks (FCN, SegNet, and U-Net) commonly used in oil spill detection in recent years, as well as the lightweight MobileNet-V3 network, under controlled variable conditions. Oil spill classification accuracy metrics, such as OA, Kappa coefficient, and F1-score, were compared, as were model lightweightness metrics, such as FLOPs, Parameters, and Model Size. To achieve optimal results, different models were paired with different optimizers and activation functions. The parameters for each model are shown in Table 4.

[0162] Table 4 Parameters of different models

[0163]

[0164]

[0165] The oil spill recognition results of different models are shown in Table 5. The recognition accuracy was evaluated using three metrics: OA, Kappa coefficient, and F1 score. As shown in Table 5, SR-SqueezeNet achieved the highest global classification accuracy, kappa coefficient, and F1 score, resulting in the best recognition performance, indicating that the majority of oil spill pixels could be successfully predicted. Among other methods, SegNet and U-Net achieved similar recognition accuracy to SR-SqueezeNet, achieving similarly good recognition results. FCN and MobileNet-V3 exhibited poor segmentation accuracy. Analysis indicates that 20 rounds of training were insufficient for these two models, resulting in suboptimal recognition performance. However, SR-SqueezeNet achieved good recognition accuracy after 20 rounds of training. This is because the Smooth-ReLU used by SR-SqueezeNet can fully explore the spatial background information of the 126 bands of hyperspectral images and learn more basic features, making each training more fitting to the real feature curve. At the same time, due to the existence of the negative saturation region, it has strong robustness and stability when processing noisy and blurred images, making the activation effect better, thereby achieving higher classification accuracy.

[0166] This paper evaluates the lightweight capabilities of models by calculating lightweight metrics for SR-SqueezeNet and different models. Table 5 also shows a quantitative evaluation of lightweight metrics for different models in the oil spill identification experiment. The results show that the proposed SR-SqueezeNet network achieves the best recognition performance compared to other models, with the fewest parameters, 2.22M, a 75.11% reduction compared to the unpruned SqueezeNet. Furthermore, the SR-SqueezeNet model size is the smallest among the models, at 12.15MB, a 17.62MB reduction compared to the unpruned SqueezeNet. This makes the algorithm advantageous for rapid extraction when deployed on resource platforms. Table 5 also shows that the FLOPs values ​​of the FCN, SegNet, and U-Net models are relatively large, with SegNet having the highest FLOPs value, reaching 156.73 GB. This indicates that traditional semantic segmentation models are computationally intensive, occupying a large amount of computer memory when loading the model, leading to memory overflows, lengthy computation times, and higher hardware memory requirements, which may limit their application in actual production. In contrast, the lightweight networks MobileNet-V3, SqueezeNet, and SR-SqueezeNet are computationally intensive, with SR-SqueezeNet in particular boasting a FLOPs value of only 28.72 GB, making it more suitable for airborne platforms with low computing power and airborne applications with limited computing resources. SR-SqueezeNet outperforms SqueezeNet in both oil spill identification accuracy and model lightweight design, demonstrating the effectiveness of model pruning and the Smooth-ReLU activation function.

[0167] When training epochs are fixed at 20, the training time varies across models. SR-SqueezeNet has the shortest training time, at just 240 seconds, followed by the lightweight MobileNet-V3. Other methods have longer training times. When training epochs are fixed at 20, the SR-SqueezeNet model can shorten training time by up to 2 minutes and 40 seconds compared to other models, significantly reducing training time. This is crucial for lightweight oil spill identification in practical applications.

[0168] Table 5 Performance evaluation of different models (Output size = 5×5, Epochs = 20)

[0169]

[0170] Figure 18 The oil spill identification results of different models are shown in Figure 2. Figure 18It can be seen that different models have different global classification results for oil spills. SR-SqueezeNet and traditional semantic segmentation networks SegNet and U-Net have better classification effects. U-Net has better extraction performance than SegNet, with only a small number of oil spill pixels predicted incorrectly. Both SegNet and U-Net are deep learning networks with good segmentation effects. They have strong robustness and stability when processing noisy and blurred images, which can improve accuracy. Compared with other deep learning models, the proposed SR-SqueezeNet has achieved good results in oil spill identification. Figure 18 It can be seen from the six images in that compared with the original airborne hyperspectral images and the ground truth, most of the oil spill pixels have been successfully predicted, and the misclassification of oil film and seawater has been avoided to a certain extent, and the detection effect on the edge area of ​​the oil film is good. However, the edge areas of the oil film detected by the unimproved SqueezeNet, FCN and MobileNet-V3 are relatively rough, and sometimes the oil film is misclassified with seawater and fences. The above conclusions show that the traditional semantic segmentation network has a good recognition effect, but the training time is long, which does not meet the requirements of lightweight models. The lightweight network MobileNet-V3 and the unimproved SqueezeNet have a short training time, but the classification accuracy is not high. The SR-SqueezeNet model proposed in the present invention can maintain good classification accuracy while shortening the training time, which meets the requirements of lightweight airborne oil spill identification.

[0171] (2) Analysis of the model’s multi-dimensional applicability.

[0172] Thermal infrared images and 4K real images are airborne observation images that are easier to obtain in oil spill identification. Thermal infrared can sense the temperature difference between the oil film and the background seawater, and perform oil spill detection around the clock. 4K real images can observe the oil spill more intuitively and clearly, and are one of the important reference standards for visual interpretation. By constructing a multi-dimensional optical remote sensing detection and identification method for oil spills based on multi-dimensional features and deep learning, accurate monitoring of the oil spill range can be achieved. In order to explore the applicability of the SR-SqueezeNet model in different observation dimensions, the present invention applies the model to airborne thermal infrared images and airborne 4K real images acquired at the same time, and analyzes the oil spill identification results from the perspective of optical and thermal multi-dimensionality. The airborne 4K real images and airborne thermal infrared images acquired simultaneously are shown as follows: Figure 19 、 Figure 20 The processed and enhanced thermal infrared images and 4K real images are also divided into three categories (oil spill, seawater, and fence background) to construct the dataset. The detailed information of the training data and test data is shown in Table 6.

[0173] Table 6 Training data and test data

[0174] Classification crude seawater background total Thermal Infrared-Training Set 2263 15841 3921 22025 Thermal Infrared-Test Set 11696 71284 26580 109560 Thermal Infrared-Total 13959 87125 30501 131585 4K-training set 1738 12166 3420 17324 4K-test set 8690 60830 17382 86902 4K-Total 10428 72996 20802 104226

[0175] The oil spill recognition accuracy indicators of SR-SqueezeNet on airborne images of different dimensions are shown in Table 7. Figure 21 The classification results of different dimensional model applications are shown. Among them, the recognition accuracy of hyperspectral images is the highest, which is 97.23%. From the recognition results, the recognition effect of hyperspectral images is also the best, and the oil and water edges are relatively clear and stable. The recognition accuracy of thermal infrared images is 93.62%, and the recognition accuracy of airborne 4K images is 94.57%. From the recognition results, most of the oil spill pixels have been identified, which proves that the SR-SqueezeNet model has good applicability on airborne images of different observation dimensions. Figure 21 It can also be seen that hyperspectral imagery clearly identifies the oil-water boundary, while thermal infrared imagery misclassifies it. The oil spill identification results from 4K real-world imagery are similar to those from thermal infrared imagery. This is due to the different locational features of oil slicks of varying thicknesses. The SR-SqueezeNet model only learns some features of oil spills on the sea surface and cannot clearly distinguish thinner oil slicks from seawater. This demonstrates that combining oil spill features learned from hyperspectral and thermal infrared imagery can achieve a certain degree of separability for oil slicks of varying thicknesses, providing insights into the future use of airborne photothermal imagery to identify oil slicks of varying thicknesses.

[0176] From the perspective of training time, it can be seen from Table 7 that the recognition time of hyperspectral images is longer, while the recognition time of thermal infrared and 4K images is much shorter. This is because hyperspectral images have 126 spectral channels onboard, while thermal infrared images and 4K images have only 3 channels. The number of features learned during model training is greatly reduced, so the training time is also shortened a lot.

[0177] Table 7 Oil spill recognition results at different observation dimensions (Output size = 5 × 5, Epochs = 20)

[0178]

[0179]

[0180] (3) Application verification of SR-SqueezeNet model.

[0181] The present invention also applies the SR-SqueezeNet model to the airborne hyperspectral image ROI-2 acquired at another time to verify the reliability and stability of the model in identifying oil spills on the sea surface. The experimental data also uses the airborne hyperspectral data acquired by the DJI M600 Pro six-rotor drone equipped with the Cubert S185 frame-type hyperspectral imager in the field experiment. The image size is 1000×1000 pixels and the number of bands is 126. The oil spill airborne hyperspectral image ROI-2 after image processing and enhancement is as follows Figure 22 As shown, ROI-2 sample space distribution Figure 23 shown.

[0182] This paper uses SR-SqueezeNet and five models compared with the existing technology to identify oil spills on the airborne hyperspectral image ROI-2, such as Figure 24 The recognition results of different models are shown below. Compared with the unmodified SqueezeNet, SR-SqueezeNet is able to more accurately identify oil spills on the sea surface. Compared with other models, SegNet's recognition results are similar to SR-SqueezeNet, but its recognition results for fences are poor, with some underclassification. The remaining FCN, U-Net, and MobileNet-V3 networks all exhibited a high number of oil-water underclassification and confusion. In summary, the SR-SqueezeNet model also achieved good recognition results in the test sample area ROI-2, demonstrating its ability to maintain high classification accuracy within a short training time.

[0183] Currently, airborne hyperspectral oil spill identification is limited by drone hardware, leaving much room for improvement in key parameter inversion methods and typical industry applications. Using hyperspectral drones for marine oil spill identification offers the advantages of flexibility, efficiency, and cost-effectiveness. With the advancement of onboard storage and processing hardware, the lightweight airborne model SR-SqueezeNet is expected to demonstrate its advantages in high recognition accuracy and rapid training speed in marine oil spill emergency disaster monitoring. This model, combined with the spectral characteristics of thermal infrared imagery, will enable further improvements in oil spill parameter inversion accuracy and wider application.

[0184] In short, drones are playing an increasingly important role in the field of oil spill monitoring for marine emergency disasters due to their flexibility, speed and low cost. The existing airborne model for oil spill identification has the problems of many parameters and large memory usage, which limits its application. The present invention designs a Smooth-ReLU activation function, which has a smooth and continuous curve at zero point, has a negative saturation region, and has a certain robustness to noise. Based on this activation function, the present invention proposes the SR-SqueezeNet lightweight model, and conducts a series of experiments using airborne hyperspectral and thermal infrared images obtained from field experiments. The present invention conducts parameter comparison experiments to determine the optimal parameters, and compares Smooth-ReLU with other commonly used activation functions. The experimental results show that the best activation effect can be achieved by using Smooth-ReLU in deeper convolutional layers.

[0185] This paper compares SR-SqueezeNet with five commonly used oil spill detection methods: SqueezeNet, FCN, SegNet, U-Net, and MobileNet-V3. The results show that SR-SqueezeNet achieves the best recognition accuracy, achieving a global classification accuracy of 97.23%, the highest kappa coefficient and F1 score, and the best recognition performance. In terms of model lightweight performance, SR-SqueezeNet achieves the smallest model size, at only 12.15MB, with 2.22M parameters and 28.72G FLOPs. Compared with other algorithms, SR-SqueezeNet achieves superior lightweight performance, making it more suitable for airborne platforms with low computing power and airborne applications with limited computing resources.

[0186] To verify the applicability of SR-SqueezeNet in observation scenarios of different dimensions, experiments were conducted using airborne thermal infrared images and airborne 4K high-definition images acquired simultaneously during field experiments, with recognition accuracies of 93.62% and 94.57%, respectively. The recognition results show that most of the oil spill pixels were identified, and the recognition time was short, which proves that SR-SqueezeNet has good applicability for airborne images of different observation dimensions. The present invention also applies SR-SqueezeNet to airborne hyperspectral images at different times, and the model also achieves good recognition results. In the future, airborne three-in-one data can be used to train the model to achieve real-time multi-dimensional oil spill remote sensing detection on board.

[0187] Embodiment 2, another object of the present invention is to provide a lightweight hyperspectral recognition system for sea surface oil spills based on a smooth activation function, comprising:

[0188] The full-chain system building module for airborne oil spill identification and verification is used to conduct hyperspectral and thermal infrared oil spill data acquisition experiments in ideal scenarios and simulated real-world scenarios, obtain land-based and airborne oil spill data and real images at different times, use land-based measured data to perform oil-water separability analysis at different bands, and build a full-chain system for airborne oil spill identification and verification;

[0189] A lightweight oil spill identification model and a smooth activation function construction module are used to build a lightweight oil spill identification model SR-SqueezeNet based on the measured data obtained in the experiment, analyze and find the optimal parameters of the lightweight oil spill identification model SR-SqueezeNet, construct a smooth activation function Smooth-ReLU, and conduct a comparative analysis of different activation functions and their application locations;

[0190] The experimental verification module is used to verify the comparative experimental results of the lightweight oil spill identification model SR-SqueezeNet and the existing technology models, comprehensively evaluate the lightweight performance of the lightweight oil spill identification model SR-SqueezeNet, and analyze the versatility of the model from the perspective of multi-dimensional light and heat in combination with airborne thermal infrared imagery. The applicability of the lightweight oil spill identification model SR-SqueezeNet is verified through airborne hyperspectral images acquired at different times.

[0191] The above description is only a preferred specific implementation method of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.

Claims

1. A lightweight hyperspectral recognition method for sea surface oil spills based on smooth activation function, characterized in that: The method includes: S1: Conduct hyperspectral and thermal infrared oil spill data acquisition experiments in ideal scenarios and simulated real scenarios, obtain land-based and airborne oil spill data and real images at different times, use land-based measured data to perform oil-water separability analysis at different bands, and build a full-chain system for airborne oil spill identification and verification; S2: Based on the measured data, a lightweight oil spill recognition model based on the SR-SqueezeNet smooth activation function was constructed. The experimental results of SR-SqueezeNet were compared with those of the unimproved squeeze network and semantic segmentation network. The optimal parameters of the lightweight oil spill recognition model SR-SqueezeNet were analyzed and found. The smooth activation function Smooth-ReLU was constructed, and the lightweight performance of the lightweight oil spill recognition model SR-SqueezeNet was comprehensively evaluated. The construction of the smooth activation function Smooth-ReLU includes: The Fire Module uses the ReLU activation function, as shown in formula (2): ReLU = {max(0, x)} (2) Where max(0, x) is the larger value compared to 0, and x is the neuron input; For the ReLU activation function, there is neuron necrosis and the derivative does not exist at zero. The smooth activation function Smooth-ReLU is used to make the curve smooth and continuous at zero. The negative saturation region is designed to improve the robustness to noise. The smooth activation function Smooth-ReLU is shown in formula (3): Where, e x To perform an exponential function operation on the input of the neuron with the natural constant e as the base; S3, airborne thermal infrared and high-resolution RGB images, verifies the optimal parameters of the model and analyzes the applicability of the model from the perspective of light and heat combination. The spatiotemporal transferability of the lightweight oil spill identification model SR-SqueezeNet is verified through airborne hyperspectral images acquired at different times.

2. The method for lightweight hyperspectral identification of sea oil spills based on smooth activation function according to claim 1 is characterized in that: In step S1, a hyperspectral and thermal infrared oil spill data acquisition test is conducted, including: S101, field experiments and data acquisition; S102, hyperspectral reflectance data preprocessing and analysis; S103, hyperspectral airborne image preprocessing and dataset construction. The constructed dataset includes: training set, validation set, and test set with a ratio of 6:2:

2.

3. The method for lightweight hyperspectral identification of sea oil spills based on smooth activation function according to claim 1 is characterized in that: In step S2, a SR-SqueezeNet lightweight oil spill recognition model based on a smooth activation function is constructed, including: Adding a Flatten layer to the basic SqueezeNet network structure to obtain a pixel-level feature map, determine the category of each pixel, and implement semantic segmentation. The SqueezeNet basic network includes the Fire module, which consists of two layers: the squeeze layer and the expand layer. The squeeze layer is a convolution layer with a 1×1 convolution kernel, and the expand layer is a convolution layer with 1×1 and 3×3 convolution kernels. In the expand layer, the feature maps obtained by 1×1 and 3×3 are fused; the Fire Module uses 1×1 convolution to replace part of the 3×3 convolution to reduce the parameters and the number of input channels; after reducing the number of channels, convolution kernels of multiple sizes are used for calculation.

4. The method for lightweight hyperspectral identification of sea oil spills based on smooth activation function according to claim 3 is characterized in that: The calculation is performed using convolution kernels of multiple sizes, including: the input active image is convolved through the previous layer to obtain a set of feature maps, and then for each size of the convolution kernel, a block of the corresponding size is taken from the input feature map and the block is element-wise multiplied with the convolution kernel; the product results at each position are accumulated to form a new eigenvalue; the results of all convolution kernels of different sizes are superimposed, and the generated feature map contains information from different spatial scales.

5. The method for lightweight hyperspectral identification of sea oil spills based on smooth activation function according to claim 1 is characterized in that: After constructing the smooth activation function Smooth-ReLU, the categorical_crossentropy used by the ReLU activation function is used as the loss function of the model. The loss function Loss is shown in formula (4): Where y i is the i-th element of the input vector, m is the number of categories, i is the i-th category, is the true label of the i-th category.

6. The method for lightweight hyperspectral identification of sea oil spills based on smooth activation function according to claim 3 is characterized in that: Add the Flatten layer to the SqueezeNet basic network structure, including: Remove the maximum pooling layer after the convolutional layer and add a Flatten layer to flatten the output of the convolutional layer and output it through the softmax function to avoid spatial dimensionality reduction of the feature map. The calculation of the softmax function is shown in formula (5): Where P(y|x) is the softmax value of the output, x is the neuron input, and y is the output vector. is a natural exponential operation on the input, k is the dimension of the input vector, j is the summation element, n is the number of classes in the multi-class classifier, and y i is the i-th element of the input vector, is a normalization term to ensure that the sum of all output values ​​of the function is 1 and each output value is in the range of (0, 1), thus forming a probability distribution. Stochastic gradient descent SGD is used as the optimization algorithm with a learning rate of 0.

001.

7. The method for lightweight hyperspectral identification of sea oil spills based on smooth activation function according to claim 1 is characterized in that: In step S2, after constructing the smooth activation function Smooth-ReLU, the lightweight oil spill recognition model SR-SqueezeNe is lightweight optimized for airborne oil spill detection through network pruning or unstructured pruning. A comparative analysis of different activation functions and their application locations is performed, including: Evaluation index analysis, based on the specific experiment to build a confusion matrix, select the overall classification accuracy OA, Kappa coefficient and F1-score three accuracy standards: The first metric: FLOPs, which represents the computational effort of the model and is used to measure the complexity of the algorithm and model. The unit is GB. The second metric is the number of parameters in the network, which is related to the model size and is measured in M. The third metric: the time for model training and classification, which objectively reflects the computing speed of the model.

8. The method for lightweight hyperspectral identification of sea oil spills based on smooth activation function according to claim 1 is characterized in that: In step S3, the optimal parameters of the model are verified, including: S301, experiment on the effect of spatial neighborhood size; S302, impact analysis experiment of different training rounds; S303, experiment on the impact of different activation functions.

9. A lightweight hyperspectral recognition system for sea surface oil spills based on smooth activation function, characterized in that: The system implements the lightweight hyperspectral identification method for sea surface oil spills based on smooth activation function according to any one of claims 1 to 8, and the system includes: The full-chain system construction module for airborne oil spill identification and verification is used to conduct hyperspectral and thermal infrared oil spill data acquisition experiments in ideal scenarios and simulated real-world scenarios, obtain land-based and airborne oil spill data and real images at different times, use land-based measured data to perform oil-water separability analysis at different bands, and build a full-chain system for airborne oil spill identification and verification; A lightweight oil spill recognition model and smooth activation function construction module are used to build the SR-SqueezeNet lightweight oil spill recognition model based on the measured data. The experimental results of SR-SqueezeNet are compared with those of the unimproved squeeze network and semantic segmentation network. The optimal parameters of the lightweight oil spill recognition model SR-SqueezeNet are analyzed and found. The smooth activation function Smooth-ReLU is constructed, and the lightweight performance of the lightweight oil spill recognition model SR-SqueezeNet is comprehensively evaluated. The experimental verification module uses airborne thermal infrared and high-resolution RGB images to verify the optimal parameters of the model and analyze the applicability of the model from the perspective of light and heat combination. The airborne hyperspectral images acquired at different times are used to verify the spatiotemporal transferability of the lightweight oil spill identification model SR-SqueezeNet.