Lightweight underwater image dynamic enhancement method
Through multi-scale enhancement prior algorithm and frequency domain-based convolution attention mechanism and other technical means, the frequency domain feature information of underwater images is extracted and optimized, and the problem of underwater image blur is solved, and the efficient image enhancement effect is achieved. It is suitable for the application of underwater robot cameras.
Patent Information
- Application Number
- CN202510075675.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
Smart Images

Figure CN120013782A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater image processing based on deep learning, and more particularly to a lightweight underwater image dynamic enhancement method. Background Art
[0002] Underwater human-computer interaction (U-HRI) is an emerging technology that explores and optimizes the interaction between humans, computer systems, and intelligent devices in underwater environments. With the increase of underwater activities, such as marine resource development, environmental monitoring, and scientific research, the demand for efficient and intuitive underwater interaction technologies has grown significantly. The development of U-HRI has promoted ocean exploration and provided innovative solutions for marine resource research. However, image blur is a key factor that limits the performance and efficiency of U-HRI in underwater tasks. The particularity of the underwater environment makes it extremely challenging to capture clear images by autonomous underwater vehicles (AUVs). With increasing depth, the absorption and scattering of underwater light intensify, significantly reducing the contrast and brightness of the image. Different wavelengths of light decay at different rates underwater, with red light decaying the fastest and blue light decaying the slowest, resulting in images that appear mainly blue or green. The presence of suspended particles and microorganisms further exacerbates image degradation. In addition, the underwater environment is highly dynamic, with uneven lighting and relative motion caused by turbulence often resulting in blurred images between the camera and the target. Considering the above factors, the research on underwater image enhancement (UIE) technology is crucial to accurately understand the underwater world and improve the efficiency and safety of U-HRI. Summary of the invention
[0003] Purpose of the invention: To address the problem that underwater robot cameras have difficulty dealing with underwater image blur when processing complex underwater environments, which seriously affects the performance of the automation system. The present invention provides a lightweight underwater image dynamic enhancement method that uses multi-scale frequency analysis to address this challenge.
[0004] The technical solution adopted by the present invention is as follows:
[0005] The present invention provides a lightweight underwater image dynamic enhancement method, which is characterized by the following steps:
[0006] Step 1: Use a multi-scale enhancement prior algorithm to enhance the high-frequency and low-frequency feature information of the original underwater image to obtain multi-scale prior image information;
[0007] Step 2: Use the frequency domain-based convolutional attention mechanism to extract the frequency domain feature information of multi-scale prior image information to obtain multi-scale frequency domain information;
[0008] Step 3: Based on step 2, multi-scale frequency domain information is input into the information flow interaction algorithm to solve the layering and blocking problems of feature information;
[0009] Step 4: Use the multi-scale cascade loss algorithm to stimulate network dynamic optimization.
[0010] Step 5: Based on step 4, the original underwater image is input into the optimized network to obtain an enhanced underwater image;
[0011] Furthermore, the enhancement steps of the multi-scale enhancement prior algorithm in step 1 are:
[0012] Step 1-1: Use bilinear downsampling to continuously downsample the original underwater image twice to obtain a multi-scale image, where the size of the original underwater image I1 is H i ×W i ×N i ; where H i Represents the length of the image, W i Represents the width of the image, N i Represents the channel size of the image; after two downsamplings, the sizes are The image I2 has a size of Image I3;
[0013] Step 1-2: Obtain the adaptive image enhancement coefficient, input the multi-scale image [I1, I2, I3] into the multi-scale enhancement prior module, and obtain the image enhancement coefficient R (C1, C2, C3, C4, C5);
[0014] Step 1-3: Using the obtained image enhancement coefficient R(C1, C2, C3, C4, C5), the multi-scale image is adjusted for brightness, contrast, saturation, gamma transformation and sharpness to obtain a multi-scale prior image [E1, E2, E3];
[0015] Furthermore, in step 1-2, the multi-scale enhanced prior module is expressed as:
[0016]
[0017] In the formula, is the enhancement coefficient R′(C1,C2,C3,C4,C5), Pool max and Pool mean They are maximum pooling and average pooling respectively, MHSA stands for multi-head attention mechanism, is the input feature, which is expressed as:
[0018]
[0019] In the formula, F IN is the input multi-scale image [I1, I2, I3], Conv is the convolution operation, CN(*) is the conditional normalization function, which is expressed as:
[0020] CN(*)=BN(F IN )*(1+α)+β
[0021] Where BN(*) is batch normalization, α and β are the scaling factor and offset factor generated by the convolution operation respectively;
[0022] The obtained enhancement factor R ′ (C1, C2, C3, C4, C5) are limited to a fixed range: C1 = [0.3, 1.3], C2 = [0.5, 1.5], C3 = [1.0, 2.0], C4 = [0.5, 1.5], C5 = [1.0, 5.0], and the final enhancement coefficient R (C1, C2, C3, C4, C5) is obtained; its limiting function is expressed as follows:
[0023]
[0024] Where V max and V min are the maximum and minimum values of the enhancement coefficient, respectively, and x is the initial enhancement coefficient R ′ (C1,C2,C3,C4,C5);
[0025] Furthermore, in steps 1-3, brightness adjustment, contrast adjustment, saturation adjustment, gamma change and sharpness adjustment are expressed as follows:
[0026] Image brightness =F IN *C1
[0027] Image contrast =F IN *C3+(1-C3)*Mean gray (F IN )
[0028] Image saturation =F IN *C4+(1-C4)*gray(F IN )
[0029]
[0030] Image sharpness =F IN *C5+(1-C5)*Laplace(F IN )
[0031] In the formula, F IN For the input multi-scale image [I1, I2, I3], Mean gray (*), gray(FIN ) and Laplace(F IN ) are average grayscale conversion, grayscale conversion and Laplace high-pass filtering respectively;
[0032] Furthermore, the frequency-domain-based convolutional attention mechanism in step 2 is expressed as:
[0033] A=Conv(IFFT(S)+DSC v (F input ))+F input
[0034] Where IFFT(*) is the inverse Fourier transform, F input is the input multi-scale prior image [E1, E2, E3], DSC(*) is the depth-wise separable convolution, and S is expressed as:
[0035] S=Conv Gelu (Q⊙K)⊙V
[0036] Where ⊙ represents the element-by-element dot multiplication operation, and Q, K, and V are expressed as:
[0037] Q, K, V = FFT q,k,v (DSC q,k,v (Conv(F input )))
[0038] Where, FFT(*) means Fourier transform;
[0039] By inputting the multi-scale prior image [E1, E2, E3] into the frequency domain-based convolutional attention mechanism to extract frequency domain features, multi-scale frequency domain information [F1, F2, F3] is obtained;
[0040] Furthermore, the information flow interaction algorithm in step 3 is expressed as:
[0041]
[0042] In the formula, Respectively expressed as:
[0043]
[0044] Where C is the number of channels of multi-scale frequency domain information, represents the feature information of the i-th channel, It is expressed as:
[0045]
[0046] In the formula is the multi-scale frequency domain information [F1, F2, F3], Linear SIt is expressed as:
[0047] A=Linear(Pool max (F IN )+Pool mean (F IN )) sigmoid
[0048] Furthermore, the multi-scale cascade loss algorithm in step 4 is expressed as:
[0049]
[0050] In the formula, α, β are correlation coefficients, where α = 0.9, β = 0.1; is the multi-scale mean absolute error, is the multi-scale frequency domain error, which are expressed as follows:
[0051]
[0052]
[0053] Where N is the data batch for optimizing the model once, p i , are the true value and predicted value respectively, fft(*) is Fourier transform;
[0054] Beneficial effects:
[0055] The present invention is a lightweight underwater image dynamic enhancement method, which uses a multi-scale enhancement prior algorithm to enhance the high-frequency and low-frequency feature information of underwater images to obtain multi-scale prior image information; uses a frequency domain-based convolutional attention mechanism to extract the frequency domain feature information of multi-scale prior image information to obtain multi-scale frequency domain information; inputs the multi-scale frequency domain information into the information flow interaction algorithm to solve the layering and blocking problems of feature information; finally, uses a multi-scale cascade loss algorithm to stimulate network dynamic optimization, and finally obtains an enhanced underwater image. The present invention can be deployed in an underwater robot, and while consuming less computing resources, it can well solve the image degradation problem of the underwater robot camera in a complex underwater environment, and the enhanced image has a higher matching degree with human visual perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is the overall network model of the present invention;
[0057] Figure 2 It is the multi-scale enhanced prior algorithm of the present invention;
[0058] Figure 3 The multi-scale enhancement a priori algorithm of the present invention enhances the effect;
[0059] Figure 4 The frequency domain-based convolutional attention mechanism of the present invention;
[0060] Figure 5 The information flow interaction algorithm of the present invention;
[0061] Figure 6 is a flow chart of the present invention;
[0062] Figure 7 This is the test result of the present invention on the LSUI dataset;
[0063] Figure 8 This is the test result of the present invention on the UIEB dataset;
[0064] Fig. 9 This is the test result of the present invention on the Challenge-60 dataset;
[0065] Fig.10 The test results of the present invention in downstream tasks; DETAILED DESCRIPTION
[0066] The specific implementation of the present invention will be described in detail below with reference to the accompanying drawings.
[0067] The present invention is a lightweight underwater image dynamic enhancement method. The overall algorithm framework is as follows: Figure 1 As shown. Use the multi-scale enhancement prior algorithm to enhance the high-frequency and low-frequency feature information of the underwater image to obtain multi-scale prior image information; use the frequency domain-based convolutional attention mechanism to extract the frequency domain feature information of the multi-scale prior image information to obtain multi-scale frequency domain information; input the multi-scale frequency domain information into the information flow interaction algorithm to solve the layering and blocking problems of feature information; finally, use the multi-scale cascade loss algorithm to stimulate the dynamic optimization of the network, and finally obtain the enhanced underwater image. The whole process is as follows Figure 6 shown.
[0068] A lightweight underwater image dynamic enhancement method, the main steps are as follows:
[0069] Step 1: Use a multi-scale enhancement prior algorithm to enhance the high-frequency and low-frequency feature information of the original underwater image to obtain multi-scale prior image information;
[0070] Enhance the frequency domain information of the original underwater image. Figure 2 This is the framework diagram of the multi-scale enhanced prior algorithm. Figure 3 This is the enhancement effect of the multi-scale prior enhancement algorithm on the original underwater image. The enhancement process is as follows:
[0071] Step 1-1: Use bilinear downsampling to continuously downsample the original underwater image twice to obtain a multi-scale image; the size of the original underwater image I1 is H i×W i ×N i ; where H i Represents the length of the image, W i Represents the width of the image, N i Represents the channel size of the image; after two downsamplings, the sizes are The image I2 has a size of Image I3;
[0072] Step 1-2: Obtain the adaptive image enhancement coefficient, input the multi-scale image [I1, I2, I3] into the multi-scale enhancement prior module, and obtain the image enhancement coefficient R (C1, C2, C3, C4, C5);
[0073] The multi-scale enhanced prior module is expressed as:
[0074]
[0075] In the formula, is the enhancement coefficient R′(C1,C2,C3,C4,C5), Pool max and Pool mean They are maximum pooling and average pooling respectively, MHSA stands for multi-head attention mechanism, is the input feature, which is expressed as:
[0076]
[0077] In the formula, F IN is the input multi-scale image [I1, I2, I3], Conv is the convolution operation, CN(*) is the conditional normalization function, which is expressed as:
[0078] CN(*)=BN(F IN )*(1+α)+β
[0079] Where BN(*) is batch normalization, α and β are the scaling factor and offset factor generated by the convolution operation respectively;
[0080] The obtained enhancement coefficient 0′(C1, C2, C3, C4, C5) is limited to a fixed range: C1 = [0.3, 1.3], C2 = [0.5, 1.5], C3 = [1.0, 2.0], C4 = [0.5, 1.5], C5 = [1.0, 5.0], and the final enhancement coefficient R (C1, C2, C3, C4, C5) is obtained; its limiting function is expressed as follows:
[0081]
[0082] Where V max and Vmin are the maximum and minimum values of the enhancement coefficient, respectively, and x is the initial enhancement coefficient R ′ (C1,C2,C3,C4,C5);
[0083] Step 1-3: Using the obtained image enhancement coefficient R(C1, C2, C3, C4, C5), the multi-scale image is adjusted for brightness, contrast, saturation, gamma transformation and sharpness to obtain a multi-scale prior image [E1, E2, E3];
[0084] Brightness adjustment, contrast adjustment, saturation adjustment, gamma change and sharpness adjustment are expressed as follows:
[0085] Image brightness =F IN *C1
[0086] Image contrast =F IN *C3+(1-C3)*Mean gray (F IN )
[0087] Image saturation =F IN *C4+(1-C4)*gray(F IN )
[0088]
[0089] Image sharpness =F IN *C5+(1-C5)*Laplace(F IN )
[0090] In the formula, F IN For the input multi-scale image [I1, I2, I3], Mean gray (*), gray(F IN ) and Laplace(F IN ) are average grayscale conversion, grayscale conversion and Laplace high-pass filtering respectively;
[0091] Step 2: Use the frequency domain-based convolutional attention mechanism to extract the frequency domain feature information of multi-scale prior image information to obtain multi-scale frequency domain information;
[0092] The multi-scale prior image information obtained in step 1 is subjected to frequency domain feature extraction to obtain multi-scale frequency domain information. The convolutional attention mechanism framework based on the frequency domain is as follows: Figure 4 shown.
[0093] Among them, the convolutional attention mechanism based on frequency domain is expressed as:
[0094] A=Conv(IFFT(S)+DSC v (F input ))+F input
[0095] Where IFFT(*) is the inverse Fourier transform, F input is the input multi-scale prior image [E1, E2, E3], DSC(*) is the depth-wise separable convolution, and S is expressed as:
[0096] S=Conv Gelu (Q⊙K)⊙V
[0097] Where ⊙ represents the element-by-element dot multiplication operation, and Q, K, and V are expressed as:
[0098] Q, K, V = FFT q,k,v (DSC q,k,v (Conv(F input )))
[0099] Where, FFT(*) means Fourier transform;
[0100] By inputting the multi-scale prior image [E1, E2, E3] into the frequency domain-based convolutional attention mechanism to extract frequency domain features, multi-scale frequency domain information [F1, F2, F3] is obtained;
[0101] Step 3: Based on step 2, multi-scale frequency domain information is input into the information flow interaction algorithm to solve the layering and blocking problems of feature information;
[0102] The multi-scale frequency domain information obtained in step 2 is fused and accelerated. The framework of the information flow interaction algorithm is as follows: Figure 5 shown.
[0103] The information flow interaction algorithm is expressed as:
[0104]
[0105] In the formula, Respectively expressed as:
[0106]
[0107] Where C is the number of channels of multi-scale frequency domain information, represents the feature information of the i-th channel, It is expressed as:
[0108]
[0109] In the formula is the multi-scale frequency domain information [F1, F2, F3], Linear S It is expressed as:
[0110] A=Linear(Pool max (F IN )+Pool mean (F IN )) Sigmoid
[0111] Step 4: Use the multi-scale cascade loss algorithm to stimulate network dynamic optimization.
[0112] Based on the fused and accelerated feature information obtained in step 3, a multi-scale cascade loss algorithm is used for back propagation to optimize the overall network parameters.
[0113] The multi-scale cascade loss algorithm is expressed as:
[0114]
[0115] In the formula, α, β are correlation coefficients, where α = 0.9, β = 0.1; is the multi-scale mean absolute error, is the multi-scale frequency domain error, which are expressed as follows:
[0116]
[0117]
[0118] Where N is the data batch for optimizing the model once, p i , are the true value and predicted value respectively, fft(*) is Fourier transform;
[0119] Step 5: Based on step 4, the original underwater image is input into the optimized network to obtain the enhanced underwater image.
[0120] The original underwater image is input into the optimized network to obtain the enhanced underwater image. Figure 7 , 8, 9 are the comparison results of the optimized network with the traditional algorithm and network on the data sets LSUI, UIEB and Challenge-60. Fig.10 The optimized network is applied to the downstream tasks of SIFT feature point detection and Canny edge detection.
[0121] Experimental Results
[0122] The present invention has conducted the following comparative experiments: (1) The enhancement effect of a lightweight underwater image dynamic enhancement method provided by the present invention is compared with that of traditional methods on the datasets LSUI, UIEB and Challenge-60. (2) The performance of a lightweight underwater image dynamic enhancement method provided by the present invention is compared with that of traditional methods in the downstream tasks of SIFT feature point detection and Canny edge detection.
[0123] The results of experiment (1) are shown in Tables 1 and 2 below.
[0124] As shown in Table 2 and Figure 7 As shown, on the LSUI dataset, a lightweight underwater image dynamic enhancement method provided by the present invention performs superiorly in all indicators, while achieving excellent image quality and reducing resource consumption. Although GAN-based networks perform better in inference speed, they consume higher resources and do not significantly improve image quality. Due to the dynamic and complex characteristics of underwater scenes, the UIE method based on physical models cannot achieve the best balance between image quality and inference speed in different underwater environments. Although the Encoder-Decoder network U-Shape achieved scores of 24.514 and 0.912 in PSNR and SSIM, respectively, which are lower than 4.177 and 0.044 of the present invention, it consumes too many resources and is therefore not suitable for optimal deployment in underwater image enhancement.
[0125] Table 1 Comparative experimental results on the datasets LSUI and UIEB
[0126]
[0127] Table 2 Comparative experimental results on the Challenge-60 dataset
[0128]
[0129] As shown in Table 1 and Figure 8 As shown in Figure 2, the UIEB dataset contains underwater images of different depths, qualities, and lighting conditions, which poses a challenge to the model performance. Although the PSNR of the proposed method on the UIEB dataset is inferior to that of LSUI, it outperforms other algorithms in PSNR, SSIM, and PI scores, demonstrating its robustness in complex underwater environments.
[0130] The LSUI and UIEB datasets use reference images, which cannot fully reflect the effectiveness in practical applications. Therefore, the present invention uses the Challenge-60 non-reference image dataset for evaluation to obtain a more objective and reliable evaluation. Fig. 9 It shows that the performance of the present invention is comparable to that of traditional methods in some indicators, or even exceeds them.
[0131] The results of experiment (2) are as follows Fig.10 And as shown in Table 3.
[0132] The information-rich images generated by the present invention facilitate feature point extraction and edge detection, while effectively reducing artifacts and noise in the original image. This improvement improves the performance of downstream tasks. Although MMLE extracts the largest number of feature points, its feature matching accuracy is unsatisfactory. This is because MMLE amplifies irrelevant interference in the image, resulting in larger noise in the extracted features. Similarly, images processed by MMLE also show large interference after edge detection, resulting in reduced clarity and continuity of edge details. For the GAN series of networks, missing feature points and blurred edges were observed. These limitations stem from their inability to fully enhance the original image, resulting in the loss of key information. The U-Shape network may have amplified some noise in the image, resulting in poor performance in both feature point and edge detection. In contrast, the present invention exhibits superior performance by extracting rich features and achieving higher feature matching accuracy, effectively suppressing most of the noise, and focusing on extracting more meaningful information from the image.
[0133] Table 3 Performance comparison on SIFT feature extraction and Canny edge detection tasks
[0134]
[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or some technical features can be replaced by equivalents without departing from the spirit of the technical solution of the present invention, which should be included in the scope of the technical solution for protection of the present invention.
Claims
1. A lightweight underwater image dynamic enhancement method, characterized by: The steps are: Step 1: Use a multi-scale enhancement prior algorithm to enhance the high-frequency and low-frequency feature information of the original underwater image to obtain multi-scale prior image information; Step 2: Use the frequency domain-based convolutional attention mechanism to extract the frequency domain feature information of multi-scale prior image information to obtain multi-scale frequency domain information; Step 3: Based on step 2, multi-scale frequency domain information is input into the information flow interaction algorithm to solve the layering and blocking problems of feature information; Step 4: Use the multi-scale cascade loss algorithm to stimulate network dynamic optimization. Step 5: Based on step 4, the original underwater image is input into the optimized network to obtain an enhanced underwater image.
2. A lightweight underwater image dynamic enhancement method according to claim 1, characterized in that: The enhancement steps of the multi-scale enhancement prior algorithm in step 1 are: Step 1-1: Use bilinear downsampling to continuously downsample the original underwater image twice to obtain a multi-scale image, where the size of the original underwater image I1 is H i ×W i ×N i ; where H i Represents the length of the image, W i Represents the width of the image, N i Represents the channel size of the image; after two downsamplings, the sizes are The image i2 has a size of image i3; Step 1-2: Obtain the adaptive image enhancement coefficient, input the multi-scale image [i1, i2, I3] into the multi-scale enhancement prior module, and obtain the image enhancement coefficient R (C1, C2, C3, C4, C5); Step 1-3: Use the obtained image enhancement coefficient R (C1, C2, C3, C4, C5) to adjust the brightness, contrast, saturation, gamma transform and sharpness of the multi-scale image to obtain a multi-scale prior image [E1, E2, E3].
3. A lightweight underwater image dynamic enhancement method according to claim 2, characterized in that: In step 1-2, the multi-scale enhanced prior module is expressed as: In the formula, is the enhancement coefficient R′(C1,C2,C3,C4,C5), Pool max and Pool mean They are maximum pooling and average pooling respectively, MHSA stands for multi-head attention mechanism, is the input feature, which is expressed as: In the formula, F IN is the input multi-scale image [I1, I2, I3], Conv is the convolution operation, CN(*) is the conditional normalization function, which is expressed as: CN(*)=BN(F IN )*(1+α)+β Where BN(*) is batch normalization, α and β are the scaling factor and offset factor generated by the convolution operation respectively; The obtained enhancement coefficient R′(C1, C2, C3, C4, C5) is limited to a fixed range: C1 = [0.3, 1.3], C2 = [0.5, 1.5], C3 = [1.0, 2.0], C4 = [0.5, 1.5], C5 = [1.0, 5.0], and the final enhancement coefficient R(C1, C2, C3, C4, C5) is obtained; its limiting function is expressed as follows: Where V max and V min are the maximum and minimum values of the enhancement coefficient respectively, and x is the initial enhancement coefficient R′(C1, C2, C3, C4, C5).
4. A lightweight underwater image dynamic enhancement method according to claim 2, characterized in that: In steps 1-3, brightness adjustment, contrast adjustment, saturation adjustment, gamma change and sharpness adjustment are expressed as follows: Image brightness =F IN *C1 Image contrast =F IN *C3+(1-C3)*Mean gray (F IN ) Image saturation =F IN *C4+(1-C4)*gray(F IN ) Image sharpness =F IN *C5+(1-C5)*Laplace(F IN ) In the formula, F IN For the input multi-scale image [I1, I2, I3], Mean gray (*), gray(F IN ) and Laplace(F IN ) are average grayscale conversion, grayscale conversion and Laplace high-pass filtering respectively.
5. A lightweight underwater image dynamic enhancement method according to claim 1, characterized in that: The frequency-domain-based convolutional attention mechanism in step 2 is expressed as: A=Conv(IFFT(S)+DSC v (F input ))+F input Where IFFT(*) is the inverse Fourier transform, F input is the input multi-scale prior image [E1, E2, E3], DSC(*) is the depth-wise separable convolution, and S is expressed as: S=Conv Gelu (Q⊙K)⊙V Where ⊙ represents the element-by-element dot multiplication operation, and Q, K, and V are expressed as: Q,K,V=FFT q,k,v (DSC q,k,v (Conv(F input ))) Where, FFT(*) means Fourier transform; By inputting the multi-scale prior image [E1, E2, E3] into the frequency domain-based convolutional attention mechanism for frequency domain feature extraction, multi-scale frequency domain information [F1, F2, F3] is obtained.
6. A lightweight underwater image dynamic enhancement method according to claim 1, characterized in that: The information flow interaction algorithm in step 3 is expressed as: In the formula, Respectively expressed as: Where C is the number of channels of multi-scale frequency domain information, represents the feature information of the i-th channel, It is expressed as: In the formula is the multi-scale frequency domain information [F1, F2, F3], Linear s It is expressed as: A=Linear(Pool max (F IN )+Pool mean (F IN )) Sigmoid。 7. A lightweight underwater image dynamic enhancement method according to claim 1, characterized in that: The multi-scale cascade loss algorithm in step 4 is expressed as: In the formula, α, β are correlation coefficients, where α = 0.9, β = 0.1; is the multi-scale mean absolute error, is the multi-scale frequency domain error, which are expressed as follows: Where N is the data batch for optimizing the model once, p i , are the true value and predicted value respectively, and fft(*) is the Fourier transform.