Chlorophyll calculation method based on double-branch multi-scale fusion network model

By combining the dual branch multi-scale fusion network model and Sentinel-2 satellite image, the problem of chlorophyll a calculation relying on field sampling is solved, and efficient and accurate monitoring of chlorophyll a concentration is achieved.

CN119988966APending Publication Date: 2025-05-13SHANDONG UNIV OF SCI & TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202411949592.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13

Smart Images

  • Figure CN119988966A_ABST
    Figure CN119988966A_ABST
Patent Text Reader

Abstract

The invention discloses a chlorophyll calculation method based on a double-branch multi-scale fusion network model, which belongs to the technical field of chlorophyll calculation, is used for calculating a chlorophyll a concentration value, and comprises the following steps: taking the chlorophyll a concentration value of a measurement site as a label of a water body characteristic sequence data set, dividing the data set into a training set and a verification set, inputting the training set into the double-branch multi-scale fusion network model to obtain a trained double-branch multi-scale fusion network model; inputting the verification set into the double-branch multi-scale fusion network model to obtain a verified double-branch multi-scale fusion network model; and taking the neighborhood spectral feature tensor and the neighborhood inherent optical feature tensor as the input of the verified double-branch multi-scale fusion network model to obtain the chlorophyll a concentration value of the pixel at the corresponding position. According to the method, the calculation result of the chlorophyll a with relatively high precision is obtained through neural network processing, and the complex relationship among the concentration change of the chlorophyll a in the water body, the spectrum and other optical parameters can be captured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention discloses a chlorophyll calculation method based on a dual-branch multi-scale fusion network model, belonging to the technical field of chlorophyll calculation. Background Art

[0002] As an important indicator for measuring water eutrophication and ecological health, the monitoring and inversion of chlorophyll a concentration has become an important topic in the field of water quality research and environmental monitoring. Traditional chlorophyll a monitoring methods mainly rely on on-site sampling and laboratory analysis, which is not only time-consuming and labor-intensive, but also difficult to complete due to limited spatial coverage. In recent years, the rapid development of unmanned ship technology and the application of satellite inversion technology have provided new solutions for water body monitoring; in terms of data processing and inversion technology, the neural network based on satellite inversion technology has received widespread attention due to its excellent feature extraction and learning capabilities, and it can be tried to apply it to the field of chlorophyll a concentration inversion. Summary of the invention

[0003] The purpose of the present invention is to provide a chlorophyll calculation method based on a dual-branch multi-scale fusion network model to solve the problem in the prior art that the calculation of chlorophyll a relies on on-site sampling, which is time-consuming and labor-intensive.

[0004] A chlorophyll calculation method based on a dual-branch multi-scale fusion network model comprises: constructing a neighborhood spectral feature tensor and a neighborhood intrinsic optical feature tensor for each measurement site, using the neighborhood spectral feature tensor and the neighborhood intrinsic optical feature tensor to generate features of a water body feature sequence data set, using the chlorophyll a concentration value of the measurement site as a label of the water body feature sequence data set, dividing the data set into a training set and a validation set, inputting the training set into the dual-branch multi-scale fusion network model to obtain a trained dual-branch multi-scale fusion network model, and then inputting the validation set into the dual-branch multi-scale fusion network model to obtain a validated dual-branch multi-scale fusion network model; using the neighborhood spectral feature tensor and the neighborhood intrinsic optical feature tensor as inputs of the validated dual-branch multi-scale fusion network model to obtain the chlorophyll a concentration value of the pixel at the corresponding position.

[0005] The dual-branch multi-scale fusion network model includes an upper branch and a lower branch.

[0006] The upper branch output and the lower branch output are concatenated into feature tensors and then input into five series-connected fully connected layers in sequence.

[0007] The neighborhood intrinsic optical feature tensor is input into the upper branch, passes through four serially connected KAN layers in sequence, and finally passes through the global average pooling layer to obtain the upper branch output.

[0008] The neighborhood spectral feature tensor is input into the lower branch, passes through four serially connected convolutional layers in sequence, and obtains the lower branch output after feature fusion.

[0009] A branch is provided before each convolution layer, and the branch includes a CBAM attention mechanism module layer and an average pooling layer connected in series. The four branches perform the feature fusion together with the fourth convolution layer.

[0010] The first KAN layer of the upper branch includes a KAN convolution layer, a maximum pooling layer, and a batch normalization + SiLU activation function layer; The second KAN layer of the upper branch includes a KAN convolution layer, a batch normalization layer, and a SiLU activation function layer; The third KAN layer of the upper branch includes a KAN convolution layer, a batch normalization layer, and a CBAM attention mechanism + SiLU activation function layer; The fourth KAN layer of the upper branch includes a KAN convolution layer, a batch normalization layer, and a SiLU activation function layer.

[0011] The first convolutional layer of the lower branch includes a two-dimensional convolutional layer; The second convolutional layer of the lower branch includes a 2D convolutional layer and a maximum pooling layer; The third convolutional layer of the lower branch includes a 2D convolutional layer and a maximum pooling layer; The fourth convolutional layer of the lower branch includes a 2D convolutional layer and a maximum pooling layer.

[0012] Constructing the neighborhood spectral feature tensor includes selecting the satellite pixel point where the measurement site is located as the central unit, combining the surrounding adjacent pixels to form a 9*9 pixel set, and the neighborhood data of each spectral band constitutes a channel. Six spectral bands, including coastal atmospheric aerosol band B1, green band B3, red band B4, red edge band B5, near-infrared band B7, and near-infrared band B8A, are used to generate spectral feature tensors of six channels with a dimension of 9*9*6.

[0013] Constructing the neighborhood intrinsic optical characteristic tensor includes selecting the satellite pixel point where the measurement site is located as the central unit, using a neighborhood of 129*129 pixels, and the neighborhood data of each spectral band constitutes a channel. The intrinsic optical characteristic tensor of each channel is the weighted sum of the intrinsic optical quantity of the measurement site and the square of the intrinsic optical characteristic tensors of other coordinate pixels in the neighborhood of 129*129 pixels. The dimension of the generated intrinsic optical characteristic tensor is 129*129*8, and the 8 channels correspond to the absorption coefficient and backscattering coefficient of 4 bands respectively. The 4 bands include 443nm band, 490nm band, 560nm band and 665nm band.

[0014] Compared with the prior art, the present invention has the following beneficial effects: the present invention obtains a high-precision calculation result of chlorophyll a through neural network processing, and can accurately capture the complex relationship between the change of chlorophyll a concentration in water and the spectrum and other optical parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is the structural diagram of the upper branch; Figure 2 It is the structure diagram of the lower branch; Figure 3 It is the structure diagram of the dual-branch multi-scale fusion network model; Figure 4 The inversion results of chlorophyll a using the dual-branch multi-scale fusion network model; Figure 5 This is the distribution map of chlorophyll a. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention is described clearly and completely below. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0017] A chlorophyll calculation method based on a dual-branch multi-scale fusion network model comprises: constructing a neighborhood spectral feature tensor and a neighborhood intrinsic optical feature tensor for each measurement site, using the neighborhood spectral feature tensor and the neighborhood intrinsic optical feature tensor to generate features of a water body feature sequence data set, using the chlorophyll a concentration value of the measurement site as a label of the water body feature sequence data set, dividing the data set into a training set and a validation set, inputting the training set into the dual-branch multi-scale fusion network model to obtain a trained dual-branch multi-scale fusion network model, and then inputting the validation set into the dual-branch multi-scale fusion network model to obtain a validated dual-branch multi-scale fusion network model; using the neighborhood spectral feature tensor and the neighborhood intrinsic optical feature tensor as inputs of the validated dual-branch multi-scale fusion network model to obtain the chlorophyll a concentration value of the pixel at the corresponding position.

[0018] The dual-branch multi-scale fusion network model includes an upper branch and a lower branch.

[0019] The upper branch output and the lower branch output are concatenated into feature tensors and then input into five series-connected fully connected layers in sequence.

[0020] The neighborhood intrinsic optical feature tensor is input into the upper branch, passes through four serially connected KAN layers in sequence, and finally passes through the global average pooling layer to obtain the upper branch output.

[0021] The neighborhood spectral feature tensor is input into the lower branch, passes through four serially connected convolutional layers in sequence, and obtains the lower branch output after feature fusion.

[0022] A branch is provided before each convolution layer, and the branch includes a CBAM attention mechanism module layer and an average pooling layer connected in series. The four branches perform the feature fusion together with the fourth convolution layer.

[0023] The first KAN layer of the upper branch includes a KAN convolution layer, a maximum pooling layer, and a batch normalization + SiLU activation function layer; The second KAN layer of the upper branch includes a KAN convolution layer, a batch normalization layer, and a SiLU activation function layer; The third KAN layer of the upper branch includes a KAN convolution layer, a batch normalization layer, and a CBAM attention mechanism + SiLU activation function layer; The fourth KAN layer of the upper branch includes a KAN convolution layer, a batch normalization layer, and a SiLU activation function layer.

[0024] The first convolutional layer of the lower branch includes a two-dimensional convolutional layer; The second convolutional layer of the lower branch includes a 2D convolutional layer and a maximum pooling layer; The third convolutional layer of the lower branch includes a 2D convolutional layer and a maximum pooling layer; The fourth convolutional layer of the lower branch includes a 2D convolutional layer and a maximum pooling layer.

[0025] Constructing the neighborhood spectral feature tensor includes selecting the satellite pixel point where the measurement site is located as the central unit, combining the surrounding adjacent pixels to form a 9*9 pixel set, and the neighborhood data of each spectral band constitutes a channel. Six spectral bands, including coastal atmospheric aerosol band B1, green band B3, red band B4, red edge band B5, near-infrared band B7, and near-infrared band B8A, are used to generate spectral feature tensors of six channels with a dimension of 9*9*6.

[0026] Constructing the neighborhood intrinsic optical characteristic tensor includes selecting the satellite pixel point where the measurement site is located as the central unit, using a neighborhood of 129*129 pixels, and the neighborhood data of each spectral band constitutes a channel. The intrinsic optical characteristic tensor of each channel is the weighted sum of the intrinsic optical quantity of the measurement site and the square of the intrinsic optical characteristic tensors of other coordinate pixels in the neighborhood of 129*129 pixels. The dimension of the generated intrinsic optical characteristic tensor is 129*129*8, and the 8 channels correspond to the absorption coefficient and backscattering coefficient of 4 bands respectively. The 4 bands include 443nm band, 490nm band, 560nm band and 665nm band.

[0027] The present invention uses a spectrometer to collect data, and derives the observation data corresponding to the three spectrometers from the deck control unit according to the operation time. The data content records the time information, sensor type, integration time, and spectral information of each corresponding wavelength. Using regular expressions, the spectral data measured by the three groups of sensors at different times are classified and stored one by one. If at a certain timestamp, due to machine failure, one or two groups of sensor data are missing, the invalid timestamp is deleted. The processed spectral data is used to calculate the off-water radiance and remote sensing reflectance. If the spectral resolution of the spectral sensor used is different, the data needs to be interpolated so that the wavelengths of the three groups of data are consistent. This process often uses a linear interpolation formula or a cubic spline interpolation formula: S i (n 1 )=a i +b i (n 1 -n i )+c i (n 1 -n i ) 2 +d i (n 1 -n i ) 3 ; S is the interpolated value of the corresponding interpolation wavelength obtained by linear interpolation, (n0, S0) and (n1, S1) are the data points of two adjacent wavelengths, n 1 is the interpolated wavelength, n i is the i-th interpolation wavelength, a i , b i 、c i ,d i is the coefficient of each interval, S i (n 1 ) is the cubic spline interpolation function.

[0028] Then the interpolated spectral data is aligned with the inertial navigation data for further processing. Inertial navigation is a navigation system that measures the acceleration and angular velocity of a moving object through inertial sensors such as gyroscopes and accelerometers to infer its current position, velocity and attitude. Its core principle is to integrate the continuous measurement of acceleration and angular velocity to obtain the changes in velocity, position and attitude. The inertial navigation system does not rely on external signals, but autonomously calculates the state of the moving body in real time from the known initial position and attitude. The accelerometer is used to measure linear acceleration, and the velocity is obtained after one integration, and the displacement is obtained by further integration. The gyroscope is used to measure angular velocity and determine the attitude of the object through integration. Therefore, inertial navigation is widely used in aviation, aerospace, submarines and unmanned driving, especially in the case of limited signals. Based on the above characteristics, the task of water body spectrum acquisition is combined with unmanned ships and inertial navigation systems to realize automatic data acquisition and error correction.

[0029] The position and attitude information of the unmanned ship during the experiment was extracted from the inertial navigation system and solved to form a text file of data. The data content includes seconds in a week, sailing distance, speed, longitude, latitude, water depth, roll angle, pitch angle, heading angle and instantaneous standard deviation at various angles and directions. The experimental survey line map was drawn using longitude, latitude and heading angle combined with geographic information.

[0030] According to the requirements of ocean optical measurement, the angle between the observation plane of the spectrometer and the incident plane of the sun should be controlled within At the same time, the solar altitude angle H S ≥30°, the formula for calculating the solar altitude angle and azimuth angle is as follows: Among them, H S represents the solar altitude angle, A s represents the solar azimuth, represents latitude, δ represents solar declination, t s Indicates the hour angle. Delete H in the spectral data S For the part less than 30°, the calculated solar azimuth angle A s The angle between the heading angle ψ of the unmanned ship, that is, the angle between the spectrometer observation plane and the solar incidence plane The value range is controlled within between, The calculation formula is as follows: After calculation, the in-situ spectra that do not meet the experimental requirements are eliminated. This can greatly avoid direct sunlight reflection and reduce the impact of ship shadows.

[0031] The influence of the heading angle on the experiment has been resolved. The unmanned ship is constantly impacted by waves on the sea surface, which may cause the attitude of the spectrometer to change, thereby causing unnecessary measurement errors. The experiment requires that the inclination of the instrument should be controlled within ±7° as much as possible. Inertial navigation provides the attitude information of the unmanned ship and indirectly reflects the attitude change of the spectrometer. In order to solve this problem, it is also necessary to combine the roll angle and pitch angle information provided by the inertial system. The roll angle usually indicates the angle of rotation of the hull around its longitudinal axis (front axis-rear axis), and is generally positive when the bow tilts to the right; the pitch angle usually indicates the angle of rotation of the hull around its transverse axis (left axis-right axis), and is generally positive when the bow lifts its head.

[0032] Normally, the sampling interval of inertial navigation is very short, so that the attitude information of the ship can be obtained at any time. In order to obtain better observation results, the spectrometer usually requires a certain integration time. A longer integration time allows the detector to collect more photons, and the different light source intensities and the reflection characteristics of the target will affect the integration time of the spectrometer. During the experiment, the unmanned ship was affected by the movement of the seawater, and the attitude changes were irregular. In addition, the weather conditions had a great impact on the experiment. It was impossible to ensure that the instrument was always level like land remote sensing, which greatly increased the difficulty of marine optical measurement. For small angle inclination (inclination < 7°), the accuracy of the data was improved by trying to measure the same site multiple times and using the attitude angle information to correct the spectrum and filter the measured data; when the inclination is too large (inclination > 7°), a more complex optical model may be required for correction. In order to ensure the reliability of the data, the data was screened and eliminated based on the attitude angle provided by the inertial navigation.

[0033] Normally, the sampling interval of the inertial navigation is set to less than 1s, which is much smaller than the minimum sampling interval of the spectrometer (usually 8s), so as to better reflect the changes in the ship's attitude and indirectly display the experimental sea conditions. First, the inertial navigation information and the spectrometer data are matched based on the time reference, and the matched data is separated by the sampling time of the spectrometer. The tilt of the unmanned ship reflected by the roll angle and pitch angle is deleted. The data that does not meet the experimental conditions (tilt angle > 7°) is deleted. Assuming that the sampling interval of the spectrometer is T, the inertial navigation records n data during this period, and the sampling frequency is f, so During each spectrometer sampling time, the roll and pitch angles θ are extracted. roll and the pitch angle θ pitch )Eigenvalues ​​of the angle: maximum value and average value. The maximum value and average value are calculated as follows: θ max,roll =max(θ roll,i ), i=1, ..., n; θ max,pitch =max(θpitch,i ), i=1, ..., n; Among them, θ max,roll and θ avg,roll are the maximum and average values ​​of the rolling angle in each sampling interval respectively; θ max,pitch and θ avg,pitch are the maximum and average pitch angles in each sampling interval, respectively. For relatively small angles, each attitude angle can be considered to independently affect the overall radiation, θ max,roll and θ max,pitch After the calibration is completed, in order to reduce the random fluctuations in the data, the data measured by the radiance and irradiance sensors are smoothed using a moving average filter. The filtering formula is shown below: Where N is the size of the filter window; t represents time; L measured and E measured are the measured values ​​of radiance and irradiance sensors respectively; L filtered and E filtered Represents the value after smoothing filtering.

[0034] According to the measurement date of the unmanned boat experiment, the available images closest to the sampling time and with less cloud cover are selected. The data source selected for this experiment is the Sentinel-2 satellite image. The Sentinel-2 satellite provides multi-spectral high-spatial resolution images, and the multi-spectral instrument (MSI) on board can capture 13 bands of information from visible light to short-wave infrared. In addition to common bands such as blue, green and red light, it also includes multiple red edge bands and short-wave infrared bands, which are of great value for the analysis of water characteristics, especially the red and near-infrared bands, which are very suitable for water environment monitoring. Because chlorophyll a has a specific spectral response to these bands, it is an ideal choice for inverting the chlorophyll a concentration of water bodies. With its high spatial and temporal resolution and rich band coverage, Sentinel-2 provides reliable data support for large-scale water monitoring and has great application potential. The Sentinel-2 satellite image data product levels are divided into two types: L1C and L2A. L1C is an atmospheric apparent reflectance product that has been orthorectified and geometrically corrected, without atmospheric correction; L2A is a reflectance product output by L1C through the Sen2Cor atmospheric correction algorithm, and does not require its own atmospheric correction. First, the downloaded L1C and L2A satellite images are resampled and cropped, and the processed L1C data is atmospherically corrected through the C2RCC module in the SNAP software. The corrected images are band synthesized and spliced ​​to complete the coverage of the study area.

[0035] In order to ensure the accuracy of the subsequent inversion model and obtain accurate chlorophyll a prediction values, the atmospheric correction results need to have high accuracy. Usually, the shape of the processed satellite spectrum curve is compared with the actual reflectance measured in situ to see if it is consistent. Due to the many differences between satellite spectra and measured spectra in spatial resolution, spectral resolution, and atmospheric influence, it is inaccurate to directly compare satellite spectra with measured spectral data. Therefore, it is necessary to use a spectral response function to convert the in-situ water surface reflectance into a remote sensing reflectance equivalent to the satellite band to ensure the consistency and comparability of the data. The calculation of the equivalent reflectance is shown in the following formula: Among them, SRF(λ) is the spectral response function value of the satellite band, λ1 and λ2 are the band ranges of the i-th band, and R rs (λ i ) is the equivalent reflectivity of the satellite’s i-th band, R rs (λ) is the in-situ water remote sensing reflectance at the corresponding wavelength. By comparison, it is found that the L2A image spectrum after the Sen2Cor atmospheric correction algorithm is basically the same as the measured water remote sensing reflectance, and its effect is better than the C2RCC model. Subsequently, the spectral reflectance of the Sentinel-2L2A satellite image is used as the characteristic input basis for chlorophyll inversion, and d(λ) is the wavelength interval of the measured object reflectance.

[0036] Different chlorophyll a concentrations in water bodies can cause changes in reflectance in specific bands, especially in blue, red, green and near-infrared bands, which provides a basis for satellite remote sensing inversion of chlorophyll concentration. By selecting characteristic bands that are highly correlated with chlorophyll a concentrations, the performance of the model and the inversion effect can be significantly enhanced. Due to the influence of background noise, spectral resolution and other factors, the results are often not ideal when the correlation between remote sensing reflectance and chlorophyll is directly established. In the water spectrum, the blue light band, green light band, and the area from the red light band to the near-infrared band usually contain important information related to chlorophyll. First, Pearson correlation analysis is used to perform feature optimization on the spectral reflectance of the chlorophyll-sensitive area band and its transformation forms (first-order differential, high-order differential, normalization, etc.), and the ones with large correlation coefficients are selected for further analysis.

[0037] Since Sentinel-2L2A images have a large pixel area, each pixel may cover multiple types of land objects. Especially in areas with complex and dense vegetation in water bodies, the spectral information of a single pixel may not accurately reflect the chlorophyll concentration in the area, resulting in a decrease in the estimation accuracy of chlorophyll a concentration. In order to better deal with this mixed effect, the neighborhood concept is introduced to fully consider the spectral feature information of each in-situ site and its neighboring pixels. The chlorophyll a concentration in natural water bodies usually has a certain spatial continuity. Using the information of surrounding pixels, the spatial understanding of spectral features can be increased.

[0038] The quasi-analytical algorithm (QAA) is a semi-analytical bio-optical model that uses spectral reflectance to calculate the intrinsic optical quantities (IOP) such as the total absorption coefficient and backscattering coefficient of water at different wavelengths. After continuous development and verification, it has the characteristics of high accuracy and strong adaptability. In coastal areas close to cities, the water colors in different regions often differ greatly. Due to the significant differences in intrinsic optical quantities, it is not reliable to invert chlorophyll a concentration using only remote sensing reflectance. The QAA algorithm is particularly suitable for processing remote sensing data with multi-band characteristics, and provides richer water characteristics information for chlorophyll inversion through the QAA algorithm. The reference band center wavelengths used by QAA_V6 are 555nm and 665nm respectively. In order to better match the use of Sentinel-2 satellites, the reference band center wavelengths of 555nm and 665nm are approximated to the Sentinel-2 center wavelengths of 560nm and 665nm. The Sentinel-2 satellite images were resampled using SNAP software. The Sentinel-2 bands required by the QAA algorithm are 443nm, 490nm, 560nm and 665nm. The L2A-level image of Sentinel-2 is resampled by SNAP, and the spatial resolution is optimized from 10 meters to 1 meter. The total absorption coefficient a and backscattering coefficient b of water bodies at other wavelengths are obtained through the QAA_V6 algorithm. bp , as shown below: Where λ is the wavelength; λ0 is the reference wavelength. rs The size of (665) is generally 550nm or 665nm; R rs (λ0) is the remote sensing reflectance at wavelength λ0; Y is the wavelength index; w o, w1 is the adjustment variable; u(λ) is the ratio of the backscattering coefficient to the sum of the absorption coefficient and the backscattering coefficient at wavelength λ, u(λ0) is the ratio of the backscattering coefficient to the sum of the absorption coefficient and the backscattering coefficient at wavelength λ0; a(λ) is the total absorption coefficient at wavelength λ, a(λ0) is the total absorption coefficient at wavelength λ0; b bp (λ) is the backscattering coefficient of suspended particles at wavelength λ, b bp (λ0) is the backscattering coefficient of pure water at wavelength λ0; a w (λ0) is the absorption coefficient of pure water; b bw (λ0) is the backscattering coefficient of pure water at wavelength λ0, b bw (λ) is the backscattering coefficient of pure water at wavelength λ. The absorption coefficient a and backscattering coefficient b of the four bands are calculated using the QAA-V6 algorithm. bp , a total of 8 IOP feature tensors, the 8-channel IOP feature matrix obtained by inversion is stored in the Sentinel-2 image band channel and then superimposed to obtain the geographic information of the satellite image.

[0039] The dual-branch network structure is usually used for processing tasks with multiple inputs or outputs. Two independent branches process data of different types or sources, and then fuse these different features for further processing. This structure allows the model to learn different feature representations and better adapt to complex tasks.

[0040] The dual-branch network model structure of the present invention is a dual-branch multi-scale fusion network-DMRF-KAN. DMRF-KAN can accept two feature tensor inputs, and the structure is divided into three layers, an input feature capture layer, a feature fusion layer, and a regression prediction layer. The final output result is the predicted chlorophyll a concentration data. The loss function is calculated with the chlorophyll a concentration in the test set, and the initial weights and biases are continuously adjusted using the loss value and the Adam optimizer to optimize the prediction model. Finally, the complex relationship between the chlorophyll a concentration and the spectral information and the intrinsic optical quantity is learned, and a trained deep learning model is obtained.

[0041] The input layer of the upper branch U1 inputs the IOP feature tensor A1 with a dimension size of 129*129*8, and the input layer of the lower branch U2 inputs the spectral feature tensor B1 with a dimension size of 9*9*6.

[0042] The upper branch is based on the KAN convolution (Convolutional KANs) architecture. KAN convolution is an innovative alternative to traditional convolutional neural networks (CNNs). Traditional CNNs mainly rely on fixed linear transformations and activation functions. KAN convolution combines the characteristics of KANs by introducing learnable spline functions to replace traditional activation functions. Compared with the multi-layer perceptron (MLP), the basic building block of today's deep learning, KAN has higher expressive power and realizes complex mapping of input data by changing the nonlinear activation function to combine inputs and weights.

[0043] Different from the MLP network that places fixed activation functions on neurons, KAN places learnable activation functions on weights. Through the learnable activation function and the summation operation on the nodes, the KAN structure can achieve more flexible and powerful function approximation capabilities in regression problems.

[0044] like Figure 1 As shown, the upper branch U1 of the input feature A1 capture layer uses the KAN convolutional neural network to extract features from the neighborhood IOP feature tensor. All KAN convolutions use the BSGottlieb-KAN activation function and are divided into 4 layers: In the first layer Layer1, the 129*129*8 A1 is convolved with 16 3*3 convolution kernels K, the step size s1 is 1, and the padding p1 is 1. The image is convolved with Conv-KAN, and the number of channels becomes 16 after convolution. Then the maximum pooling process Maxpool2d is performed, and the feature map size is reduced to half of the original size. The output feature map is recorded as A2. The output feature map A2 is subjected to batch normalization operation BN to stabilize the feature distribution of the KAN convolution output. After normalization, the SiLU activation function is used to perform nonlinear transformation on the feature map.

[0045] The second layer consists of KAN convolution, batch normalization and residual blocks. The feature map from the upper layer passes through the second layer of KAN convolution, which also uses a 3*3 convolution kernel but increases to 32 channels, with a step size of 2 and a padding of 1. The output feature map is recorded as A3, and the number of A3 channels increases to 32, while the size is halved to 32*32. After batch normalization and SiLU activation, the input is directly added to the output after the second convolution through a residual connection, further enhancing the robustness and nonlinear representation of the features.

[0046] The third layer consists of KAN convolution, batch normalization and attention modules. This layer first uses a 3*3 convolution kernel with a step size of 2 and a padding of 1 to expand the number of channels of feature map A3 to 64 and reduce the size to 16*16, generating the output feature map A4. Then batch normalization is performed on A4 to stabilize the feature distribution, and then the CBAM attention mechanism module is introduced. This module extracts the important features of each channel by calculating the global average pooling and the maximum pooling, and then fuses the two to generate channel weighting to highlight the important features related to chlorophyll concentration. The feature map after the CBAM module continues to pass through the SiLU activation function and is input to the next layer.

[0047] The fourth layer is similar to the first layer. It continues to use the 3*3 KAN convolution kernel on the feature map from the upper layer. The number of channels is further increased to 128, so that the feature map size is reduced to 8*8. The output feature map is recorded as A5. After batch normalization, A5 uses the SiLU activation function to enhance the nonlinear response ability of the network, followed by global average pooling. After this layer of processing, the feature map is flattened into a one-dimensional vector A6 with a size of 1*1*8192, which is used as the final output of the upper branch and provides input for the subsequent feature fusion layer.

[0048] like Figure 2 As shown in the figure, the lower branch U2 of the input feature capture layer extracts features from the neighborhood spectral feature tensor B1. The activation functions of all convolutional layers of B1 are ReLU. The lower branch U2 has two connection modes: the main network connection and the collateral network connection. The main network connection first performs a convolution operation on the input feature map (using the Same filling method), and adjusts the number of channels while keeping the size of the feature map unchanged. Next, the features are further extracted through maximum pooling, reducing the spatial size of the feature map and achieving spatial compression of the features. The pooled feature map is then passed to the next convolutional layer to continue extracting high-dimensional features. The connection of the collateral network introduces the input feature map into the CBAM attention mechanism module. The feature map output by CBAM is processed by average pooling, and the feature map size is reduced to 1×1 while retaining the channel dimension space. The collateral feature map can be used as a summary of global semantic information to help capture global features.

[0049] First, U2 accepts the input spectral feature tensor B1 with a dimension of 9*9*6. In the first layer of the backbone network, 16 3*3 convolution kernels are used to perform two same-filled convolution operations. Through this operation, the feature map size remains unchanged, the number of channels increases, and the convolution output size of the first layer is 9*9*16, denoted as B2 and input into the second layer of the backbone network; the second layer of the backbone network uses 32 3*3 convolution kernels to perform two same-filled convolution operations, keeping the feature map size at 9*9 and increasing the number of channels to 32. Then, the feature map size is reduced to 4*4 through the maximum pooling with a pooling kernel size of 2*2 and a step size of 2, denoted as B3, further compressing the feature map space to reduce the amount of calculation and retaining the most representative features; next, in the third layer, 64 3*3 convolution kernels are used to repeat the same convolution and pooling operations, so that the feature map size becomes 2*2*64, denoted as B4; finally, the fourth layer of the backbone network uses 128 3*3 convolution kernels, and the feature map size is further reduced to 1*1*128, and the final output B5 of the network is generated through maximum pooling.

[0050] In the collateral network path of U2, the first layer also uses B1 as input, and the input feature map is weighted through the CBAM attention mechanism module to strengthen the response of the key area features. The size of the feature map after the CBAM module remains unchanged, and then the spatial dimension is reduced to 1*1 through average pooling. The number of channels remains unchanged, and the output feature map size is 1*1*6, denoted as b1. Similarly, the second to fourth layers process the feature maps B2, B3, and B4 through the CBAM module respectively, and average pool them in the spatial dimension to generate collateral outputs b2, b3, and b4, corresponding to 1*1*16, 1*1*32, and 1*1*64, respectively. Finally, these collateral outputs are concat-fused with the final output B5 of the backbone network, denoted as B6, to form a rich multi-scale global feature expression, providing input for the subsequent feature fusion layer.

[0051] The present invention uses the CBAM attention mechanism module in both the upper and lower branches. The attention mechanism module improves the perception ability of the model by introducing channel attention and spatial attention, thereby improving performance without increasing network complexity. The CBAM attention mechanism module is divided into three steps: 1. Channel attention mechanism: It compresses the feature map F in the spatial dimension to obtain a one-dimensional vector. The specific operation includes: firstly generating two one-dimensional vectors through global average pooling and maximum pooling, then inputting them into a shared network, merging them element by element, and finally generating the channel attention map M through the sigmoid activation function. C , the feature map M obtained by the channel attention mechanism C Multiplying it with the original feature map F yields the weighted feature map F′.

[0052] 2. Spatial attention mechanism: The weighted feature map F′ is globally averaged and stacked on the channel, and then the spatial attention feature map M is generated through convolution operation and sigmoid activation function. s .

[0053] 3. Fusion attention mechanism: The feature map M obtained by the channel attention mechanism C Multiply it with the original feature map F, and the obtained weighted feature map F′ is then multiplied with the spatial attention feature map M s Multiply them together and finally get the output feature map F″.

[0054] Channel attention helps to enhance the feature representation of different channels, while spatial attention helps to extract key information at different locations in space. By combining these two attention mechanisms, the model can more effectively capture features related to chlorophyll a concentration and achieve significant performance gains.

[0055] In the feature fusion layer, the high-dimensional features output by the two feature extractions are integrated with the low-dimensional features from the input data and the intermediate layer to form a feature map containing multi-scale information. Through the Concatenate operation, the feature fusion layer effectively integrates the inherent optical quantity features extracted by the upper branch U1 and the spectral features captured by the lower branch U2, enhancing the network's ability to perceive changes in chlorophyll concentration. These fused feature vectors enter the subsequent regression prediction layer to further extract effective information and predict chlorophyll a concentration.

[0056] like Figure 3 After the feature fusion layer, a series of fully connected layers are added to connect the output of the previous layer to each neuron in the current layer to achieve regression prediction in a step-by-step manner. This network model sets up a five-layer fully connected network, gradually reducing the number of nodes, and the number of nodes in the layer is 512, 128, 64, 32, and 1, respectively, to extract key features and prevent overfitting. After several step-by-step fully connected layers, the final chlorophyll a concentration prediction value is output.

[0057] During the training process, the initial learning rate is set to 0.01, the batch size is set to 64 for each iteration, the Adam optimizer is selected as the optimizer, and the loss function is the RMSE function. The model is trained by minimizing the loss function until the preset maximum number of iterations or accuracy requirements are reached to obtain the optimal model.

[0058] like Figure 4 , the inversion result of the present invention is: R 2The RMSE is 1.710 μg / L and the inversion accuracy is 0.87. It can be seen that the model inversion has a high inversion accuracy, which verifies the effectiveness of the method of the present invention. The chlorophyll a concentration in the experimental area was inverted using the trained dual-branch multi-scale fusion network model, and the chlorophyll a concentration distribution map was drawn using ArcGIS software. Figure 5 shown.

[0059] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features may be replaced by equivalents, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A chlorophyll calculation method based on a dual-branch multi-scale fusion network model, characterized in that: The method includes constructing a neighborhood spectral feature tensor and a neighborhood intrinsic optical feature tensor for each measurement site, using the neighborhood spectral feature tensor and the neighborhood intrinsic optical feature tensor to generate features of a water body feature sequence data set, using the chlorophyll a concentration value of the measurement site as a label of the water body feature sequence data set, dividing the data set into a training set and a validation set, inputting the training set into a dual-branch multi-scale fusion network model to obtain a trained dual-branch multi-scale fusion network model, and then inputting the validation set into the dual-branch multi-scale fusion network model to obtain a verified dual-branch multi-scale fusion network model; using the neighborhood spectral feature tensor and the neighborhood intrinsic optical feature tensor as inputs of the verified dual-branch multi-scale fusion network model to obtain the chlorophyll a concentration value of the pixel at the corresponding position.

2. The chlorophyll calculation method based on the dual-branch multi-scale fusion network model according to claim 1 is characterized in that: The dual-branch multi-scale fusion network model includes an upper branch and a lower branch.

3. The chlorophyll calculation method based on the dual-branch multi-scale fusion network model according to claim 2 is characterized in that: The upper branch output and the lower branch output are concatenated into feature tensors and then input into five series-connected fully connected layers in sequence.

4. The chlorophyll calculation method based on a dual-branch multi-scale fusion network model according to claim 2 is characterized in that: The neighborhood intrinsic optical feature tensor is input into the upper branch, passes through four serially connected KAN layers in sequence, and finally passes through the global average pooling layer to obtain the upper branch output.

5. The chlorophyll calculation method based on a dual-branch multi-scale fusion network model according to claim 2 is characterized in that: The neighborhood spectral feature tensor is input into the lower branch, passes through four serially connected convolutional layers in sequence, and obtains the lower branch output after feature fusion.

6. The chlorophyll calculation method based on a dual-branch multi-scale fusion network model according to claim 5 is characterized in that: A branch is provided before each convolution layer, and the branch includes a CBAM attention mechanism module layer and an average pooling layer connected in series. The four branches perform the feature fusion together with the fourth convolution layer.

7. The chlorophyll calculation method based on a dual-branch multi-scale fusion network model according to claim 4 is characterized in that: The first KAN layer of the upper branch includes a KAN convolution layer, a maximum pooling layer, and a batch normalization + SiLU activation function layer; The second KAN layer of the upper branch includes a KAN convolution layer, a batch normalization layer, and a SiLU activation function layer; The third KAN layer of the upper branch includes a KAN convolution layer, a batch normalization layer, and a CBAM attention mechanism + SiLU activation function layer; The fourth KAN layer of the upper branch includes a KAN convolution layer, a batch normalization layer, and a SiLU activation function layer.

8. The chlorophyll calculation method based on a dual-branch multi-scale fusion network model according to claim 5 is characterized in that: The first convolutional layer of the lower branch includes a two-dimensional convolutional layer; The second convolutional layer of the lower branch includes a 2D convolutional layer and a maximum pooling layer; The third convolutional layer of the lower branch includes a 2D convolutional layer and a maximum pooling layer; The fourth convolutional layer of the lower branch includes a 2D convolutional layer and a maximum pooling layer.

9. The chlorophyll calculation method based on a dual-branch multi-scale fusion network model according to claim 1 is characterized in that: Constructing the neighborhood spectral feature tensor includes selecting the satellite pixel point where the measurement site is located as the central unit, combining the surrounding adjacent pixels to form a 9*9 pixel set, and the neighborhood data of each spectral band constitutes a channel. Six spectral bands, including coastal atmospheric aerosol band B1, green band B3, red band B4, red edge band B5, near-infrared band B7, and near-infrared band B8A, are used to generate spectral feature tensors of six channels with a dimension of 9*9*6.

10. The chlorophyll calculation method based on a dual-branch multi-scale fusion network model according to claim 1, characterized in that: Constructing the neighborhood intrinsic optical characteristic tensor includes selecting the satellite pixel point where the measurement site is located as the central unit, using a neighborhood of 129*129 pixels, and the neighborhood data of each spectral band constitutes a channel. The intrinsic optical characteristic tensor of each channel is the weighted sum of the intrinsic optical quantity of the measurement site and the square of the intrinsic optical characteristic tensors of other coordinate pixels in the neighborhood of 129*129 pixels. The dimension of the generated intrinsic optical characteristic tensor is 129*129*8, and the 8 channels correspond to the absorption coefficient and backscattering coefficient of 4 bands respectively. The 4 bands include 443nm band, 490nm band, 560nm band and 665nm band.

Citation Information

Cited By

  • Non-destructive detection method for content of chlorophyll and carotenoid in tobacco leaves

    CN120510153A

  • Water chlorophyll prediction model generation method and device based on space-time transmission mechanism and proxy model, and electronic equipment

    CN121483414A

  • Chlorophyll concentration prediction method based on hybrid expert architecture

    CN121662202A

  • A method for predicting chlorophyll-a concentration based on a hybrid expert architecture

    CN121662202B