Multi-feature marine environment sensing method and system based on cross-scale attention mechanism
Through the multi-character marine environment perception method of cross-scale attention mechanism, the problems of insufficient utilization of medium-scale phase information and limited feature recognition in the existing technology are solved, and high-precision wind and wave parameter inversion is achieved, providing multi-character information for marine environment monitoring, supporting the development of marine economy and security.
Patent Information
- Application Number
- CN202510436418.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-08-08
AI Technical Summary
The existing X-band phase-particle navigation radar environmental perception technology does not fully utilize the amplitude phase information, the image feature recognition is limited, the output features are single, and the actual measurement inversion accuracy is not high, making it difficult to achieve high-precision monitoring of multi-character marine environmental parameters.
A multi-featured marine environment perception method based on a cross-scale attention mechanism is adopted. By multi-case selection and data segmentation of X-band phase-participated navigation radar data, a multi-featured marine environment perception model of a pyramid structure is constructed, a cross-scale embedded layer and a dynamic distance self-attention layer are introduced, a multi-scale feature of radar echo is extracted, and a activation value cooling layer is used to improve model performance, and a mapping relationship between radar amplitude phase information and wind and wave parameters is established.
High-precision real-time perception of marine wind and wave environments is achieved. The inversion results show that the average inversion accuracy of marine environmental parameters with multiple characteristics of wind and waves can reach 90%, providing accurate monitoring of a variety of marine environmental parameters, supporting marine economic development and security guarantees.
Smart Images

Figure CN120448724A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of ship and ocean engineering technology, and relates to a method for time history inversion of an ocean environment, and in particular to a multi-feature ocean environment perception method and system based on a cross-scale attention mechanism. Background Art
[0002] Marine environmental monitoring plays a vital role in my country's marine economic development and national defense security. This monitoring relies on accurate knowledge of basic data and changing patterns of marine environmental parameters, including wind and waves. Offshore engineering activities such as offshore oil extraction, drilling, aquaculture, and offshore power generation require monitoring of wind and wave parameters. Furthermore, providing real-time and accurate data on wind and wave parameters is essential to strengthen defense against hostile threats emanating from the deep sea, enhance marine environmental security capabilities, and promote the development of strategic support points for marine environmental security.
[0003] The main method of obtaining wind and wave ocean environment parameters is to conduct on-site measurements using tools such as buoys, anemometers, optical cameras, and radars. Among them, in-situ sensors such as buoys and anemometers can accurately reflect the changing characteristics of the ocean environment at the observation point over time, but they also have inherent limitations: (1) In-situ sensors can only provide data at fixed locations and, due to the limitations of sea conditions and meteorological conditions, can only obtain data within a specific time and on a local point line, lacking the overall view required for offshore operations. (2) The measurement equipment is susceptible to natural or man-made damage and requires high maintenance costs. Among them, optical camera measurements are significantly affected by weather and lighting conditions, making it difficult to work normally in environments such as clouds, rain, and at night. Their monitoring range is limited and they are easily affected by sea surface reflection and refraction, affecting data accuracy and stability. In contrast, radar observations are not affected by external forces and have all-weather operation capabilities. They provide a stable, continuous, and comprehensive view of the ocean surface across time and space, and can quickly and widely obtain ocean environment information. At present, radar observation technology has gradually become the main means of obtaining wind and wave ocean environment information. It can conduct long-term, large-scale, and synchronous ocean environment monitoring and provide a variety of ocean environment elements. It is widely used in military and industrial fields.
[0004] When electromagnetic waves emitted by a radar resonate with sea surface waves, the echoes received by the radar antenna form a sea clutter image. This image contains information about wind and waves, allowing the extraction of wind and wave environmental parameters from the radar image. Common radars include synthetic aperture radar (SAR), high-frequency surface wave radar (HFSR), and X-band marine radar. SAR requires mobile platforms such as satellites or aircraft, resulting in high acquisition costs and poor real-time performance. HFSR also suffers from high cost, large near-shore blind spots, and low resolution. In contrast, X-band marine radar offers advantages such as low cost, ease of operation, high real-time performance, and high spatial and temporal resolution. Furthermore, X-band coherent marine radar uses phase difference technology to effectively suppress the Doppler effect, reduce interference from clutter and noise, and thus improve the signal-to-noise ratio (SNR) and produce high-precision radar images. Therefore, using X-band coherent marine radar to perceive the ocean environment has become a research hotspot.
[0005] At present, many scholars have conducted extensive research on the inversion of wind and wave parameters, but the following problems still exist: (1) The information collected by coherent radar includes echo intensity and phase information, but current research mainly uses the phase information of coherent radar images to obtain the speed information of the waves, and does not fully utilize the amplitude and phase information. In addition, the current ocean environment parameter inversion output is single, and most of the inversion is based on one of the wind and wave parameters. At the same time, intelligent inversion methods that consider wind and wave environment parameters have not yet been explored. (2) Traditional inversion methods based on radar remote sensing use empirical modulation functions, and the correlation between the signal-to-noise ratio and the significant wave height is small, resulting in inaccurate spectrum conversion, which affects the overall inversion accuracy, and the calculation is complex and time-consuming. (3) In recent years, more and more inversions tend to use convolutional neural networks, but convolutional neural networks usually require a large data set to train, while radar data is limited; convolutional neural networks are good at capturing edge information, while radar images have less edge texture information, so convolutional neural networks are limited in learning radar image feature extraction. Therefore, this type of technical solution has limitations in terms of practical engineering application significance.
[0006] In summary, existing X-band coherent marine radar environmental perception does not fully utilize amplitude and phase information, has limited image feature recognition, outputs single features, and has low measured inversion accuracy. Therefore, achieving more accurate measurement of multi-feature ocean environmental parameters has become a key issue that needs to be addressed. In terms of inversion of marine environmental statistical characteristics, the problems and defects of existing technologies are as follows:
[0007] (1) Chinese patent CN116958435A, invention title: A wave information inversion system based on X-band radar images, publication date: 2023.10.27, this invention mainly uses the phase information of coherent radar images to obtain the velocity information of the waves, and does not make full use of the amplitude and phase information. In addition, the current ocean environment parameter inversion output is single, and most of the inversion is based on one of the wind and wave parameters. Intelligent inversion methods that are suitable for considering wind and wave environment parameters at the same time have not yet been explored.
[0008] (2) Chinese patent CN103969643B, invention title: A method for inverting ocean wave parameters using X-band navigation radar based on a new type of ocean wave dispersion relation bandpass filter, publication date: August 6, 2014. This invention is based on the traditional inversion method of radar remote sensing using an empirical modulation function, which has a small correlation between the signal-to-noise ratio and the significant wave height, resulting in inaccurate spectrum conversion, thus affecting the overall inversion accuracy, and the calculation is complex and time-consuming.
[0009] (3) Chinese patent CN117647808B, invention name: Non-coherent radar phase-resolved wave time history inversion method based on deep learning, publication date: 2024.03.05, the invention uses convolutional neural network inversion, which requires a large data set training, while radar data is limited; and radar images have less edge texture information, and convolutional neural networks have limited learning for radar image feature extraction, and are mainly based on simulated wave field inversion. Summary of the Invention
[0010] To overcome the problems existing in the related art, the embodiments disclosed in the present invention provide a multi-feature ocean environment perception method and system, and in particular, a multi-feature ocean environment perception method and system based on a cross-scale attention mechanism. The technical solution is as follows:
[0011] The present invention is implemented as follows: a multi-feature ocean environment perception method based on a cross-scale attention mechanism, comprising the following steps:
[0012] S1, select the effective area of multiple working conditions for the measured complex value data of X-band coherent marine radar;
[0013] S2, preprocessing the input data, building the corresponding label data set and performing data segmentation, dividing the data into training set and test set;
[0014] S3, establish a multi-feature ocean environment perception model based on a cross-scale attention mechanism;
[0015] S4, initialize the model training parameters, input the obtained training set into the multi-feature ocean environment perception model for training, and adjust the model parameters according to the validation set results;
[0016] S5, based on the coherent radar complex value data read in real time, the pre-trained multi-feature ocean environment perception model obtained through step S4 is calculated, and the wind and wave parameter inversion results in ocean environment perception are output in real time.
[0017] In step S1, multi-condition effective area selection is performed on the measured complex value data of the X-band coherent marine radar, including:
[0018] Select the radar data analysis area, the initial data is the real echo data I=[[I x1 ,I x2 …I xn ],[I y1 ,I y2 …I yn ]] and imaginary echo data Q=[[Q x1 ,Q x2 …Qx n ],[Q y1 ,Q y2 …Q yn ]] is a two-dimensional image matrix, where I xn ,I yn ,Q xn ,Q yn are the real and imaginary echo data respectively, and x and y are the distance and the number of azimuth points respectively;
[0019] The expression for selecting the radar data analysis area is:
[0020] C(r,a)=(r / R) 0.5 ρ(r,a,t)
[0021] Where C(r,α) is the selected area, r is the position, R is the radar detection distance, and ρ(r,α,t) is the initial radar image.
[0022] In step S2, the input data is preprocessed to construct the corresponding label dataset and perform data segmentation, including:
[0023] The initial input data is normalized using the Z-score standardization method, and the data is converted into a standard normal distribution with a mean of 0 and a standard deviation of 1. Multidimensional tensor splicing and bilinear interpolation are performed. The final input data is the measured coherent radar complex-valued two-dimensional echo data matrix D = [[I x1 ,I x2 …I xn ],[I y1 ,I y2 …I yn ]],[[Q x1 ,Q x2 …Q xn ],[Q y1 ,Qy2 …Q yn ]], the dimension is 2×384×384, the corresponding labels are: [wave height, wave period, wave direction, wind speed, wind direction], Y=[y1,y2,y3,y4,y5], the dimension is 1×5, build the data set, and split it into training set The corresponding label dataset is Test set The corresponding label dataset is in, and are the real and imaginary echo data for training and testing respectively. The normal normalization formula of the initial input data is:
[0024] Z=(X-μ) / σ
[0025] Where Z is the transformed data, X is the original data, μ is the mean of the original data, and σ is the standard deviation of the original data.
[0026] In step S3, a multi-feature ocean environment perception model based on a cross-scale attention mechanism is established, including:
[0027] The model adopts a pyramid structure, gradually reducing the resolution of feature maps and increasing the number of channels through four stages, generating a cross-scale hierarchical feature representation from local details to global context. A dynamic self-attention layer is introduced to reduce the resolution of feature maps and generate multi-scale feature representations. A cross-scale embedding layer is introduced to increase the number of channels in feature maps to compensate for information loss caused by reduced resolution and extract higher-level features. Activation value cooling layers are added in stages 2, 3, and 4. Each stage consists of a cross-scale embedding layer and several dynamic self-attention modules.
[0028] The cross-scale embedding layer appears at the beginning of each stage and is used to receive the output or input image of the previous stage as input. It samples the patch using multiple kernels of different scales and constructs each token by embedding and connecting these convolution kernels. In this process, the number of embeddings is reduced to a quarter and the size of the pyramid structure is doubled. Multiple dynamic self-attention modules are set after the cross-scale embedding layer. Each dynamic self-attention module consists of a long-distance attention module, a multi-layer perceptron, a dynamic position bias module, and a residual connection and normalization module. The long-distance attention module includes a short-distance attention module or a long-distance attention module. The short-distance attention module and the long-distance attention module appear alternately in different blocks. The dynamic position bias module works in both the short-distance attention module and the long-distance attention module to obtain the embedded position representation, and a residual connection is used in each block.
[0029] The specific contents of each stage are:
[0030] The first stage: The original feature map F has a dimension of H×W×C. The number of channels is doubled by the cross-scale embedding layer, and the dimension is H×W×C1. The feature map is divided into 4×4 patches by the dynamic self-attention module, and each patch is embedded in the space of C1 dimension. Feature extraction and global context modeling are performed on the output feature map F1, and the dimension is Where C1 = 2C;
[0031] The second stage: the dimension of the feature map F1 is The number of channels is doubled by the cross-scale embedding layer, and the dimension is The feature map is divided into 2×2 patches through the dynamic self-attention module, and each patch is embedded in the C2 dimension space. Feature extraction and global context modeling are performed on the output feature map F2, and an activation value cooling layer is added. The dimension is Where C2 = 2C1;
[0032] The third stage: the dimension of the feature map F2 is The number of channels is doubled by the cross-scale embedding layer, and the dimension is The feature map is divided into 2×2 patches through the dynamic self-attention module, and each patch is embedded in the C3 dimension space. Feature extraction and global context modeling are performed on the output feature map F3, and an activation value cooling layer is added. The dimension is Where, C3=2C2;
[0033] The fourth stage: the dimension of feature map F3 is The number of channels is doubled by the cross-scale embedding layer, and the dimension is The feature map is divided into 2×2 patches through the dynamic self-attention module, and each patch is embedded in the C4 dimension space. Feature extraction and global context modeling are performed on the output feature map F4, and an activation value cooling layer is added. The dimension is Among them, C4=2C3.
[0034] In step S4, the obtained training set is input into the multi-feature ocean environment perception model for training, including:
[0035] The model is trained with the preset number of training cycles, initial learning rate, and mean square error as the loss function. The AdamW optimizer is used, including regularization to suppress overfitting, and the CosineAnnealing scheduler is used for smooth annealing. After the parameters are set, the resulting training set is input into the multi-feature ocean environment perception model. After each round of training, gradient zeroing, error backpropagation, parameter update, and learning rate adjustment are performed. The model's generalization ability is verified using a validation set, and the model parameters and hyperparameters are adjusted based on the verification results.
[0036] During model training, the cross-scale embedding layer receives the measured radar amplitude and phase two-dimensional matrix or the output of the previous level as input, and uses four different kernel sizes of 4×4, 8×8, 16×16 and 32×32 to sample the patch with the same stride of 4×4 to complete multi-scale feature extraction and fusion of the data. In the 2 / 3 / 4 stages, two different kernel sizes of 2×2 and 4×4 are used to gradually downsample with a stride of 2×2; the multi-scale feature data enters the dynamic range self-attention module, where short-range attention and long-range attention mechanisms appear alternately to capture the dependency between local interactions and global features, and dynamically generate relative position biases; an activation value cooling layer is inserted in the 2 / 3 / 4 stages, and five classification heads are set to perform wave height, wave period, wave direction, wind speed and wind direction inversion respectively.
[0037] Furthermore, the specific method of the cross-scale embedding layer is:
[0038] (1) Multi-scale feature sampling: Use multiple convolution kernels of different scales to downsample the input. The step size S of the convolution kernel is set to 2. The convolution operation for the i-th scale is:
[0039]
[0040] Where, P i is the feature block obtained by sampling, For a kernel size K i And the convolution operation of step size S, E is the input feature matrix;
[0041] (2) Multi-scale feature embedding and splicing: Linearly embed the feature blocks of each scale and project them to the specified dimension before splicing. The expression is:
[0042]
[0043] Where, T i P i Perform linear transformation to embed the feature blocks of the linear layer, Linear i For linear transformation, the feature is transformed from C i Embedded in D i , T is all T i The splicing result.
[0044] Furthermore, the specific method of the long and short distance attention module is:
[0045] Near attention and long-distance attention are used alternately. For the near-distance attention module, it is used for the dependency between adjacent embeddings to capture local information. The feature map with an input size of S×S is adjacently embedded in a local group of G×G. The near-distance attention module performs self-attention calculation in each G×G area. The long-distance attention module is used to process long-distance dependencies and realize the interaction between farther embeddings in the model. The feature map with an input size of S×S is sampled at an interval of I. All embeddings with an interval of I are divided into the same group, and the embeddings in each group will be self-attention calculated. For the long-distance attention module with an input size of S×S, the embeddings are sampled at a fixed interval of I.
[0046] The specific method of dynamic position bias is as follows: relative position bias represents the relative position of embedding by adding a bias to the embedded attention; nonlinear transformation consists of normalization and activation function RELU and fully connected layer; the input dimension of dynamic position bias is 2, and the dimension of the middle layer is set to D / 4, where D is the embedding dimension; output B i,j is a scalar, encoding i th and j th Embed the relative position features between them and add the dynamic position bias value to the attention score calculation;
[0047] B i,j =MLP(Δx i,j ,Δy i,j )
[0048] Where, (Δx i,j ,Δy i,j ) is the relative coordinate difference between the i-th and j-th feature block units; MLP( ) is a lightweight multi-layer perceptron, which consists of normalization and activation function RELU and fully connected layers;
[0049]
[0050] Where Attn is the attention size, Q,K,V∈R N×d are query, key, and value matrices respectively, is the square root of the dimension of k, B∈R N×N is a dynamically generated bias matrix.
[0051] Furthermore, the activation value cooling layer consists of a 3×3 deep convolution layer and a normalization layer. For deep convolution, local operations are used to smooth the features. Each channel is convolved separately without introducing information interaction between channels. The convolution kernel size is 3×3, and the features are adjusted in the spatial dimension. For normalization, each channel is normalized independently so that the mean of the feature is 0 and the variance is 1, reducing the activation value amplitude of the feature to a stable range.
[0052] In step S5, based on the real-time read coherent radar complex value data, the pre-trained multi-feature ocean environment perception model obtained in step S4 is calculated and the wind and wave parameter inversion results in ocean environment perception are output in real time, including:
[0053] For the test data set obtained in step S2 Input the model trained in step S4, and solve the model to obtain the inversion results of the ocean environment perception time history parameters in, Real and imaginary echo data for testing, D te The measured coherent radar complex-valued echo data matrix used for testing, Y te The model output is the inverted wind and wave parameters [wave height, wave period, wave direction, wind speed, wind direction].
[0054] Another object of the present invention is to provide a multi-feature ocean environment perception system based on a cross-scale attention mechanism, which is used to regulate the multi-feature ocean environment perception method based on a cross-scale attention mechanism. The system includes:
[0055] The region selection module is used to select effective regions under multiple working conditions for the measured complex value data of the X-band coherent marine radar;
[0056] The data partitioning module is used to preprocess the input data, build the corresponding label data set and perform data segmentation, dividing the data into training set and test set;
[0057] A multi-feature ocean environment perception model building module, which is used to build a multi-feature ocean environment perception model based on a cross-scale attention mechanism;
[0058] The model training module is used to initialize the model training parameters, input the obtained training set into the multi-feature ocean environment perception model for training, and adjust the model parameters according to the validation set results;
[0059] The parameter inversion module is used to calculate based on the coherent radar complex value data read in real time, obtain the pre-trained multi-feature ocean environment perception model, and output the wind and wave parameter inversion results in ocean environment perception in real time.
[0060] In combination with all the above technical solutions, the beneficial effects of the present invention are as follows:
[0061] First, in view of the technical problems existing in the above-mentioned prior art and the difficulty of solving the problems, the present invention closely combines the technical solutions to be protected and the results and data in the research and development process, and analyzes in detail and deeply how the technical solutions of the present invention solve the technical problems, and some creative technical effects brought about after solving the problems, which are specifically described as follows: In view of the problems that the existing X-band coherent marine radar environment perception does not fully utilize the amplitude and phase information, has limited image feature recognition, and has a single output feature, the present invention proposes a multi-feature marine environment perception method based on a cross-scale attention mechanism, firstly, the X-band coherent marine radar time history data is subjected to multi-condition effective area selection and data segmentation, and is divided into a training set and a test set; then the training set data is input into the test set. The model is introduced. Inside the model, the resolution of the feature map is reduced by introducing a dynamic self-attention layer to generate a multi-scale feature representation. The cross-scale embedding layer is introduced to increase the number of channels of the feature map to compensate for the information loss caused by the resolution reduction and extract higher-level features. In this way, the multi-scale features of radar echo intensity and Doppler velocity (amplitude and phase information) are explicitly extracted, and the dependency between their local interactions and global features is captured. Then, an activation value cooling layer is used to improve the amplitude explosion problem in the model performance. Finally, the pre-trained optimal multi-feature ocean environment perception model is used to establish the mapping relationship between radar amplitude and phase information and ocean environment parameters such as wind (wind speed and wind direction) and wave (significant wave height, characteristic period, and wave direction), thus realizing high-precision real-time perception of the offshore wind and wave environment.
[0062] Second, considering the technical solution as a whole or from the perspective of the product, the technical effects and advantages of the technical solution to be protected by this invention are described in detail as follows: By constructing a multi-feature marine environment perception model, this invention provides an accurate multi-feature marine environmental information monitoring method for ship navigation, offshore operations, etc., providing fundamental technical support for my country's marine economic development and security, and having important engineering significance. Inversion results show that the average inversion accuracy of the wind and wave multi-feature marine environmental parameters of this invention reaches 90%, achieving high precision in the inversion of measured data.
[0063] Third, the positive effects of the present invention are also reflected in the following important aspects:
[0064] (1) Full Utilization of Amplitude and Phase Information: Previous techniques typically use only either the phase or amplitude information of coherent radar images to invert ocean wave parameters, underutilizing both information. The present invention, by simultaneously considering both amplitude and phase information, allows for a more comprehensive extraction of ocean environmental parameters, thereby improving the accuracy and reliability of the inversion.
[0065] (2) II is applicable to intelligent inversion of multi-feature ocean environment parameters: the inversion output of previous technologies is single, and most of them are inverted based on one of the wind and wave parameters, without considering the inversion of multiple ocean environment parameters at the same time. In recent years, more and more inversions tend to use convolutional neural networks, but convolutional neural networks usually require a large data set to train, while radar data is limited; convolutional neural networks are good at capturing edge information, while radar images have less edge texture information, so convolutional neural networks are limited in learning radar image feature extraction. However, this patent introduces a cross-scale embedding layer and a dynamic range self-attention layer to explicitly extract the multi-scale features of radar echo intensity and Doppler velocity (amplitude and phase information), captures the dependency between their local interactions and global features, and then uses an activation value cooling layer to improve the amplitude explosion problem in the model performance. It can simultaneously invert multiple ocean environment parameters with high precision and provide more comprehensive ocean environment information. This full utilization of amplitude and phase information and the intelligent inversion applicable to multi-feature ocean environment parameters have not been fully explored in existing technologies at home and abroad, so the present invention fills the technical gap in this regard.
[0066] (3) Researchers are committed to developing a method that can simultaneously invert multiple features of ocean environmental parameters with high precision. Existing traditional inversion methods based on radar remote sensing often use empirical modulation functions, which have problems such as low correlation between signal-to-noise ratio and significant wave height, inaccurate spectrum conversion, etc., resulting in low measured inversion accuracy and complex and time-consuming calculations. Therefore, more and more inversions tend to use convolutional neural networks, but convolutional neural networks usually require a large data set to train, while radar data is limited; convolutional neural networks are good at capturing edge information, while radar images have less edge texture information, so convolutional neural networks are limited in learning radar image feature extraction. This patent introduces a new intelligent method and pre-trained models to be able to invert multiple ocean environmental parameters in real time, providing multi-feature ocean environmental information with high measured inversion accuracy, thus solving this technical problem.
[0067] (4) Coherent radar usually only uses its phase information without considering amplitude and phase information at the same time. Traditional inversion methods and intelligent methods have excellent effects on simulated data, but have low accuracy in measured data. The present invention overcomes the prejudice that algorithms are generally difficult to adapt to measured data. By improving the inversion method, the average inversion accuracy of its multi-feature ocean environment parameters can reach 90%, and the average inversion accuracy of 4-5 level sea conditions can reach 93%, realizing real-time inversion of multiple ocean environment parameters and providing multi-feature ocean environment information with high measured inversion accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure;
[0069] Figure 1 This is a flow chart of a multi-feature ocean environment perception method provided by an embodiment of the present invention;
[0070] Figure 2 This is a schematic diagram of a multi-feature ocean environment perception method provided by an embodiment of the present invention;
[0071] Figure 3 Schematic diagram of radar data area selection provided by an embodiment of the present invention;
[0072] Figure 4 This is a diagram of the architecture of a multi-feature ocean environment perception model provided by an embodiment of the present invention;
[0073] Figure 5 is an internal structure diagram of two consecutive dynamic range self-attention layers provided by an embodiment of the present invention;
[0074] Figure 6 is a schematic diagram of wave height inversion results provided by an embodiment of the present invention;
[0075] Figure 7 is a schematic diagram of a wave period inversion result provided by an embodiment of the present invention;
[0076] Figure 8 is a schematic diagram of wave direction inversion results provided by an embodiment of the present invention;
[0077] Figure 9 1 is a schematic diagram of wind speed inversion results provided by an embodiment of the present invention;
[0078] Figure 10 Schematic diagram of wind direction inversion results provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0079] To make the above-mentioned objects, features, and advantages of the present invention more readily apparent, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. The following description sets forth numerous specific details to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art may make similar modifications without departing from the scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0080] The innovation of the present invention lies in: In response to the problems of the existing X-band coherent marine radar environment perception that the amplitude and phase information is not fully utilized, the image feature recognition is limited, the output feature is single, and the actual measurement accuracy is not high, the present invention proposes a multi-feature ocean environment perception method based on the cross-scale attention mechanism, and constructs a multi-feature ocean environment perception model based on the cross-scale attention mechanism. The model gradually reduces the resolution of the feature map and increases the number of channels through four stages to generate a hierarchical feature representation from local details to global context. At each stage, a cross-scale embedding layer is first introduced to increase the number of channels of the feature map to enhance the expressive power of the features. Subsequently, a dynamic range self-attention layer is introduced to reduce the resolution of the feature map to generate a multi-scale feature representation. The model can not only capture visual information of different scales, but also improve the richness of features by increasing the number of channels, so as to better process the feature maps after the resolution is reduced in the subsequent stage, and then explicitly extract the multi-scale features of radar echo intensity and Doppler velocity (amplitude and phase information), capture the dependency of their local interactions and global features, make the model focus on areas with obvious wave characteristics, and use the activation value cooling layer to improve the gradient explosion problem of model performance. Then, the measured complex-valued echo data obtained by radar measurement can be subjected to feature extraction and analysis to directly obtain high-precision wind (wind speed, wind direction) and wave (significant wave height, characteristic period, wave direction) ocean environmental parameters. Finally, a mapping relationship between the original measured radar echo data and the wind (wind speed, wind direction) and wave (significant wave height, characteristic period, wave direction) ocean environmental parameters is established, realizing real-time and high-precision perception of the marine wind and wave environment, providing an accurate multi-feature marine environmental information monitoring method for ship navigation, marine operations, etc., and providing basic technical means to support my country's marine economic development and security, which has important engineering significance.
[0081] Example 1, as Figure 1 and Figure 2 As shown, based on the real-time read coherent radar complex value data, the pre-trained optimal multi-feature ocean environment perception model is obtained through step S4 for calculation, and the wind and wave parameter inversion result in ocean environment perception can be output in real time. The multi-feature ocean environment perception method provided by the embodiment of the present invention specifically includes the following steps:
[0082] S1, select the effective area of multiple working conditions for the measured complex value data of X-band coherent marine radar;
[0083] In order to achieve universality and robustness, a variety of ocean environment inputs including wave height, wave period, wave direction, wind speed and wind direction parameters are considered. The measured coherent radar complex echo data of 35 working conditions are used. According to formula (1), the appropriate radar data analysis area is selected. The initial data is the real part echo data I = [[I x1 ,I x2 …I xn ],[I y1 ,I y2…I yn ]] and imaginary echo data Q=[[Q x1 ,Q x2 …Q xn ],[Q y1 ,Q y2 …Q yn ]]Two-dimensional image matrix, where I xn ,I yn ,Q xn ,Q ym are the real and imaginary echo data respectively, x and y are the distance and the number of azimuth points respectively; the formula for selecting the radar data analysis area is:
[0084] C(r,a)=(r / R) 0.5 ρ(r,a,t)(1)
[0085] Where C(r,α) is the selected area, r is the location, R is the radar detection range, which is set to 0.8-1.9 km; ρ(r,α,t) is the initial radar image; α is the antenna azimuth angle, which is set to 80°;
[0086] like Figure 3 As shown in the figure, the raw radar data is selected with moderate echo intensity, avoiding invalid areas (land) and data areas with weak echo intensity, providing high-quality data for the deep learning model.
[0087] S2, preprocessing the input data, building the corresponding label data set and performing data segmentation, dividing the data into training set and test set;
[0088] The initial input data is normalized using the Z-score normalization method, as shown in Equation (2). The data is converted into a standard normal distribution with a mean of 0 and a standard deviation of 1, which stabilizes model training, accelerates convergence, alleviates the gradient vanishing and explosion problems, improves model generalization, and performs multi-dimensional tensor splicing and bilinear interpolation to meet the deep learning input data framework (NCHW) N–Batch C–Channel H–Height W–Width. The final input data is the measured coherent radar complex-valued two-dimensional echo data matrix, D = [[I x1 ,I x2 …I xn ],[I y1 ,I y2 …I yn ]],[[Q x1 ,Q x2 …Q xn ],[Q y1 ,Q y2 …Q yn]], the dimension is 2×384×384, the corresponding labels are: [wave height, wave period, wave direction, wind speed, wind direction], Y=[y1,y2,y3,y4,y5], the dimension is 1×5, build the data set, and split it into training set The corresponding label dataset is Test set The corresponding label dataset is in, and The real and imaginary echo data are used for training and testing respectively. The data set has a total of 1050 sets of data, 90% of which are used as training sets and 10% as test sets. The normal normalization formula for the initial input data is:
[0089] Z=(X-μ) / σ(2)
[0090] Where Z is the transformed data, X is the original data, μ is the mean of the original data, and σ is the standard deviation of the original data.
[0091] S3, establish a multi-feature ocean environment perception model based on a cross-scale attention mechanism;
[0092] like Figure 4 As shown in , the model adopts a pyramid structure and is divided into four stages. Each stage consists of a cross-scale embedding layer and several dynamic self-attention modules. The cross-scale embedding layer appears at the beginning of each stage. It receives the output (or input image) of the previous stage as input and samples the patch using multiple kernels of different scales (e.g., 4×4 or 8×8), and constructs each token by embedding and connecting these convolution kernels. In this way, some small dimensions (e.g., 4×4 convolution kernels) are forced to focus only on small-scale features, while other large dimensions (e.g., 8×8 convolution kernels) only focus on learning large-scale features, thereby producing tokens with clear cross-scale features. Multiple dynamic self-attention modules are set after the cross-scale embedding layer. Each dynamic self-attention module consists of a long-short distance attention module (including a short-distance attention module or a long-distance attention module), a multi-layer perceptron, a dynamic position bias module, and a residual connection and normalization module. As shown Figure 5 As shown, the short-distance attention module and the long-distance attention module appear alternately in different blocks. The dynamic position bias module operates in both the short-distance attention module and the long-distance attention module to obtain the embedded position representation. Residual connections are used in each block. In particular, extreme values in the late stage can make the training process unstable, hinder model convergence, and lead to amplitude explosion. To this end, activation value cooling layers are added in stages 2, 3, and 4 to improve the gradient explosion problem in model performance.
[0093] Through four stages, the feature map resolution is gradually reduced and the number of channels is increased to generate a cross-scale hierarchical feature representation from local details to global context. A dynamic self-attention layer is introduced to reduce the feature map resolution and generate multi-scale feature representations. A cross-scale embedding layer is introduced to increase the number of feature map channels to compensate for the information loss caused by the resolution reduction and extract higher-level features. Activation value cooling layers are added in stages 2, 3, and 4 to improve the gradient explosion problem in model performance. Each stage consists of a cross-scale embedding layer and several dynamic self-attention modules.
[0094] The cross-scale embedding layer appears at the beginning of each stage and is used to receive the output or input image of the previous stage as input. It samples the patch using multiple kernels of different scales and constructs each token by embedding and connecting these convolution kernels. In this process, the number of embeddings is reduced to a quarter, while the size of the pyramid structure is doubled. Multiple dynamic self-attention modules are set after the cross-scale embedding layer. Each dynamic self-attention module consists of a long-distance attention module, a multi-layer perceptron, a dynamic position bias module, and a residual connection and normalization module. The long-distance attention module includes a short-distance attention module or a long-distance attention module. The short-distance attention module and the long-distance attention module appear alternately in different blocks. The dynamic position bias module works in both the short-distance attention module and the long-distance attention module to obtain the embedded position representation, and a residual connection is used in each block.
[0095] The specific content of each stage,
[0096] The first stage: The original feature map F has a dimension of H×W×C. The number of channels is doubled by the cross-scale embedding layer, and the dimension is H×W×C1. The feature map is divided into 4×4 patches by the dynamic self-attention module, and each patch is embedded in the space of C1 dimension. Feature extraction and global context modeling are performed on the output feature map F1, and the dimension is Where C1 = 2C;
[0097] The second stage: the dimension of the feature map F1 is The number of channels is doubled by the cross-scale embedding layer, and the dimension is The feature map is divided into 2×2 patches through the dynamic self-attention module, and each patch is embedded in the C2 dimension space. Feature extraction and global context modeling are performed on the output feature map F2, and an activation value cooling layer is added. The dimension is Where C2 = 2C1;
[0098] The third stage: the dimension of the feature map F2 is The number of channels is doubled by the cross-scale embedding layer, and the dimension is The feature map is divided into 2×2 patches through the dynamic self-attention module, and each patch is embedded in the C3 dimension space. Feature extraction and global context modeling are performed on the output feature map F3, and an activation value cooling layer is added. The dimension is Where, C3=2C2;
[0099] The fourth stage: the dimension of feature map F3 is The number of channels is doubled by the cross-scale embedding layer, and the dimension is The feature map is divided into 2×2 patches through the dynamic self-attention module, and each patch is embedded in the C4 dimension space. Feature extraction and global context modeling are performed on the output feature map F4, and an activation value cooling layer is added. The dimension is Among them, C4=2C3.
[0100] S4, initialize the model training parameters, input the obtained training set into the multi-feature ocean environment perception model for training, and adjust the model parameters according to the validation set results;
[0101] The specific model training method is as follows: the model is trained for 1000 rounds, with an initial learning rate of 0.001 and a loss function of mean squared error (MSE). The AdamW optimizer is used for regularization to suppress overfitting, and the CosineAnnealing scheduler (cosine curve) is used for smooth annealing to prevent premature model convergence and improve training stability. After the parameters are set, the training set obtained in step S2 is input into the model. After each round of training, gradients are cleared, errors are backpropagated, parameters are updated, and the learning rate is adjusted. In addition, a validation set is used to verify the model's generalization ability, and the model parameters and hyperparameters are adjusted based on the validation results. The data flow during model training is as follows: First, the cross-scale embedding layer receives the measured radar amplitude and phase two-dimensional matrix (or the output of the previous level) as input. It uses four kernels of different sizes (4×4, 8×8, 16×16, and 32×32) to sample the patch with the same stride of 4×4 to complete multi-scale feature extraction and fusion. In the 2nd, 3rd, and 4th stages, it uses two kernels of different sizes (2×2 and 4×4) with a stride of 2×2 to gradually downsample the features layer by layer to provide higher-level semantic information. Next, the multi-scale feature data enters the dynamic range self-attention module, where short-range and long-range attention mechanisms alternate to capture the dependencies between local interactions and global features. Dynamically generated relative position biases are used to enhance the spatial perception ability of the attention mechanism. In particular, activation value cooling layers are inserted at each stage to periodically suppress feature amplitude explosion. Finally, five classification heads are set up to perform temporal inversion of wave height, wave period, wave direction, wind speed, and wind direction, respectively.
[0102] The specific method of the cross-scale embedding layer of the present invention is:
[0103] (1) Multi-scale feature sampling: Use multiple convolution kernels of different scales to downsample the input. The step size S of the convolution kernel remains consistent (set to 2) to ensure that the same number of feature blocks are generated at each scale. For the convolution operation of the i-th scale, as shown in formula (3):
[0104]
[0105] Where, P i is the feature block obtained by sampling, with a size of H'×W'×C'; For a kernel size K i and convolution operation with stride S, E is the input feature matrix with size H×W×C;
[0106] (2) Multi-scale feature embedding and splicing: Linearly embed the feature blocks of each scale, project them into the specified dimension (small-scale feature blocks are assigned more dimensions, while large-scale feature blocks are assigned fewer dimensions), and then splice them, as shown in Equation (4):
[0107]
[0108] Where, T i P i Perform linear transformation to embed the feature block of the linear layer, with a size of H'×W'×D'; Linear i For linear transformation, the feature is transformed from C i Embedded in D i , T is all T i The stitching result has a size of H'×W'×D'.
[0109] The specific method of the long and short distance attention module of the present invention is:
[0110] Near-range and long-range attention are used alternately to balance local and global features. The near-range attention module focuses on dependencies between adjacent embeddings, capturing local information. Feature maps of input size S×S are embedded adjacently into local groups of G×G. The near-range attention module performs self-attention within each G×G region. The long-range attention module handles long-range (global) dependencies, namely, interactions between embeddings farther away in the model. Embeddings are sampled from feature maps of input size S×S at an interval of I. All embeddings with an interval of I are grouped into the same group, and self-attention is performed within each group to avoid calculating global dependencies across the entire image. For the long-range attention module, embeddings are sampled at a fixed interval of I. Since the embeddings here are constructed only from single-scale features, it is difficult to establish dependencies. The upper cross-scale embedding layer, with adjacent large-scale features, provides sufficient context for connection, compensating for the shortcomings of long-range cross-scale attention.
[0111] The specific method of dynamic position biasing of the present invention is:
[0112] The relative position bias represents the relative position of the embedding by adding a bias to the embedded attention. Its nonlinear transformation consists of normalization and activation function RELU and fully connected layer. The input dimension of dynamic position bias is 2, that is, (Δx i,j ,Δy i,j ), let the dimension of the middle layer be D / 4, where D is the embedding dimension. Output B i,j is a scalar that encodes i th and j th The relative position features between embeddings are shown in formula (5), and the dynamic position bias value is added to the attention score calculation, as shown in formula (6).
[0113] B i,j =MLP(Δx i,j ,Δy i,,j )5)
[0114] Where, (Δx i,j ,Δy i,j ) is the relative coordinate difference between the i-th and j-th feature block units; MLP() is a lightweight multi-layer perceptron, which consists of normalization and activation function RELU and fully connected layers;
[0115]
[0116] Where Attn is the attention size, Q,K,V∈R N×d are query, key, and value matrices respectively, is the square root of the dimension of k, B∈R N×N is a dynamically generated bias matrix.
[0117] The activation value cooling layer of the present invention is composed of a 3×3 deep convolution layer and a normalization layer. For deep convolution, local operations are used to smooth the features and remove certain high-frequency components. Each channel is convolved separately, and no information interaction between channels is introduced, thereby maintaining lightweight. The convolution kernel size is 3×3, and the features are adjusted in the spatial dimension. For normalization, each channel is normalized independently so that the mean of the feature is 0 and the variance is 1, and the amplitude activation value amplitude of the feature is reduced to a stable range to prevent gradient explosion without losing information. These two steps are equivalent to "resetting" the feature amplitude, avoiding the explosive growth of the gradient accumulated by the amplitude.
[0118] S5, based on the coherent radar complex value data read in real time, the pre-trained multi-feature ocean environment perception model obtained through step S4 is calculated, and the wind and wave parameter inversion results in ocean environment perception are output in real time.
[0119] For the test data set obtained in step S2 Input the model trained in step S4, and solve the model to obtain the inversion results of the ocean environment perception time history parameters in, Real and imaginary echo data for testing, D te The measured coherent radar complex-valued echo data matrix used for testing, Y te The model output is the inverted wind and wave parameters [wave height, wave period, wave direction, wind speed, wind direction].
[0120] In Example 2, a software interface is added to realize a visualization system from radar data collection to wind and wave parameter inversion. The multi-feature ocean environment perception system provided by the embodiment of the present invention includes:
[0121] The region selection module is used to select effective regions under multiple working conditions for the measured complex value data of the X-band coherent marine radar;
[0122] The data partitioning module is used to preprocess the input data, build the corresponding label data set and perform data segmentation, dividing the data into training set and test set;
[0123] A multi-feature ocean environment perception model building module, which is used to build a multi-feature ocean environment perception model based on a cross-scale attention mechanism;
[0124] The model training module is used to initialize the model training parameters, input the obtained training set into the multi-feature ocean environment perception model for training, and adjust the model parameters according to the validation set results;
[0125] The parameter inversion module is used to calculate based on the coherent radar complex value data read in real time, obtain the pre-trained multi-feature ocean environment perception model, and output the wind and wave parameter inversion results in ocean environment perception in real time.
[0126] In order to further demonstrate the positive effects of the above embodiment, the present invention conducts the following experiments based on the above technical solution.
[0127] In existing classical wave inversion methods, the image spectrum obtained through Fourier transform requires the use of a modulation transfer function to convert it into the actual wave spectrum. This results in inaccurate spectrum conversion, which affects the overall inversion accuracy. Furthermore, methods based on signal-to-noise ratio (SNR) have low accuracy in inverting significant wave heights.
[0128] In the existing classical wave inversion method, the image spectrum obtained by Fourier transform is first subjected to dispersion relation filtering and converted into the actual wave spectrum with the help of a modulation function. Since the pixel intensity of the radar image does not directly reflect the wave height information, the current engineering mainly inverts the significant wave height based on the energy ratio between the wave spectrum and the background noise in the image spectrum, that is, the empirical relationship between the signal-to-noise ratio and the significant wave height. However, due to the inaccurate dispersion bandpass filtering effect and the small correlation between the signal-to-noise ratio and the significant wave height, the wave inversion accuracy of the existing spectrum analysis method is poor. The inversion effect of the present invention and the classical wave inversion effect are shown in Table 1, wherein the wave height inversion, wave period inversion, wave direction inversion, wind speed inversion and wind direction inversion results of the present invention are shown in Table 1 respectively. Figure 6 、 7 As shown in Figures , 8, 9, and 10, the average inversion accuracy of the wind and wave multi-characteristic ocean environment parameters can reach 90%, and the average inversion accuracy of sea conditions of level 4-5 can reach 93%, which achieves high accuracy in the measured inversion.
[0129] Table 1 Evaluation indicators
[0130]
[0131] The above description is only a preferred specific implementation method of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A multi-feature ocean environment perception method based on a cross-scale attention mechanism, characterized by: The method comprises the following steps: S1, select the effective area of multiple working conditions for the measured complex value data of X-band coherent marine radar; S2, preprocessing the input data, building the corresponding label data set and performing data segmentation, dividing the data into training set and test set; S3, establish a multi-feature ocean environment perception model based on a cross-scale attention mechanism; S4, initialize the model training parameters, input the obtained training set into the multi-feature ocean environment perception model for training, and adjust the model parameters according to the validation set results; S5, based on the coherent radar complex value data read in real time, the pre-trained multi-feature ocean environment perception model obtained through step S4 is calculated, and the wind and wave parameter inversion results in ocean environment perception are output in real time.
2. The multi-feature ocean environment perception method based on the cross-scale attention mechanism according to claim 1 is characterized in that: In step S1, multi-condition effective area selection is performed on the measured complex value data of the X-band coherent marine radar, including: Select the radar data analysis area, the initial data is the real echo data I=[[I x1 ,I x2 …I xn ],[I y1 ,I y2 …I yn ]] and imaginary echo data Q=[[Q x1 ,Q x2 …Q xn ],[Q y1 ,Q y2 …Q yn ]] is a two-dimensional image matrix, where I xn ,I yn ,Q xn ,Q yn are the real and imaginary echo data respectively, and x and y are the distance and the number of azimuth points respectively; The expression for selecting the radar data analysis area is: C(r,a)=(r / R) 0.5 ρ(r,a,t) Where C(r,α) is the selected area, r is the position, R is the radar detection distance, and ρ(r,α,t) is the initial radar image.
3. The multi-feature ocean environment perception method based on the cross-scale attention mechanism according to claim 2 is characterized in that: In step S2, the input data is preprocessed to construct the corresponding label dataset and perform data segmentation, including: The initial input data is normalized using the Z-score standardization method, and the data is converted into a standard normal distribution with a mean of 0 and a standard deviation of 1. Multidimensional tensor splicing and bilinear interpolation are performed. The final input data is the measured coherent radar complex-valued two-dimensional echo data matrix D = [[I x1 ,I x2 …I xn ],[I y1 ,I y2 …I yn ]],[[Q x1 ,Q x2 …Q xn ],[Q y1 ,Q y2 …Q yn ]], the dimension is 2×384×384, the corresponding labels are: [wave height, wave period, wave direction, wind speed, wind direction], Y=[y1,y2,y3,y4,y5], the dimension is 1×5, build the data set, and split it into training set The corresponding label dataset is Test set The corresponding label dataset is in, and are the real and imaginary echo data for training and testing respectively. The normal normalization formula of the initial input data is: Z=(X-μ) / σ Where Z is the transformed data, X is the original data, μ is the mean of the original data, and σ is the standard deviation of the original data.
4. The multi-feature ocean environment perception method based on the cross-scale attention mechanism according to claim 1 is characterized in that: In step S3, a multi-feature ocean environment perception model based on a cross-scale attention mechanism is established, including: The model adopts a pyramid structure, gradually reducing the resolution of feature maps and increasing the number of channels through four stages, generating a cross-scale hierarchical feature representation from local details to global context. A dynamic self-attention layer is introduced to reduce the resolution of feature maps and generate multi-scale feature representations. A cross-scale embedding layer is introduced to increase the number of channels in feature maps to compensate for information loss caused by reduced resolution and extract higher-level features. Activation value cooling layers are added in stages 2, 3, and 4. Each stage consists of a cross-scale embedding layer and several dynamic self-attention modules. The cross-scale embedding layer appears at the beginning of each stage and is used to receive the output or input image of the previous stage as input. It samples the patch using multiple kernels of different scales and constructs each token by embedding and connecting these convolution kernels. In this process, the number of embeddings is reduced to a quarter and the size of the pyramid structure is doubled. Multiple dynamic self-attention modules are set after the cross-scale embedding layer. Each dynamic self-attention module consists of a long-distance attention module, a multi-layer perceptron, a dynamic position bias module, and a residual connection and normalization module. The long-distance attention module includes a short-distance attention module or a long-distance attention module. The short-distance attention module and the long-distance attention module appear alternately in different blocks. The dynamic position bias module works in both the short-distance attention module and the long-distance attention module to obtain the embedded position representation, and a residual connection is used in each block. The specific contents of each stage are: The first stage: The original feature map F has a dimension of H×W×C. The number of channels is doubled by the cross-scale embedding layer, and the dimension is H×W×C1. The feature map is divided into 4×4 patches by the dynamic self-attention module, and each patch is embedded in the space of C1 dimension. Feature extraction and global context modeling are performed on the output feature map F1, and the dimension is Where C1 = 2C; The second stage: the dimension of feature map F1 is The number of channels is doubled by the cross-scale embedding layer, and the dimension is The feature map is divided into 2×2 patches through the dynamic self-attention module, and each patch is embedded in the C2 dimension space. Feature extraction and global context modeling are performed on the output feature map F2, and an activation value cooling layer is added. The dimension is Where C2 = 2C1; The third stage: the dimension of the feature map F2 is The number of channels is doubled by the cross-scale embedding layer, and the dimension is The feature map is divided into 2×2 patches through the dynamic self-attention module, and each patch is embedded in the C3 dimension space. Feature extraction and global context modeling are performed on the output feature map F3, and an activation value cooling layer is added. The dimension is Where, C3=2C2; The fourth stage: the dimension of feature map F3 is The number of channels is doubled by the cross-scale embedding layer, and the dimension is The feature map is divided into 2×2 patches through the dynamic self-attention module, and each patch is embedded in the C4 dimension space. Feature extraction and global context modeling are performed on the output feature map F4, and an activation value cooling layer is added. The dimension is Among them, C4=2C3.
5. The multi-feature ocean environment perception method based on the cross-scale attention mechanism according to claim 4 is characterized in that: In step S4, the obtained training set is input into the multi-feature ocean environment perception model for training, including: The model is trained with the preset number of training cycles, initial learning rate, and mean square error as the loss function. The AdamW optimizer is used, including regularization to suppress overfitting, and the CosineAnnealing scheduler is used for smooth annealing. After the parameters are set, the resulting training set is input into the multi-feature ocean environment perception model. After each round of training, gradient zeroing, error backpropagation, parameter update, and learning rate adjustment are performed. The model's generalization ability is verified using a validation set, and the model parameters and hyperparameters are adjusted based on the verification results. During model training, the cross-scale embedding layer receives the measured radar amplitude and phase two-dimensional matrix or the output of the previous level as input, and uses four different kernel sizes of 4×4, 8×8, 16×16 and 32×32 to sample the patch with the same stride of 4×4 to complete multi-scale feature extraction and fusion of the data. In the 2 / 3 / 4 stages, two different kernel sizes of 2×2 and 4×4 are used to gradually downsample with a stride of 2×2; the multi-scale feature data enters the dynamic range self-attention module, where short-range attention and long-range attention mechanisms appear alternately to capture the dependency between local interactions and global features, and dynamically generate relative position biases; an activation value cooling layer is inserted in the 2 / 3 / 4 stages, and five classification heads are set to perform wave height, wave period, wave direction, wind speed and wind direction inversion respectively.
6. The multi-feature ocean environment perception method based on the cross-scale attention mechanism according to claim 5 is characterized in that: The specific method of the cross-scale embedding layer is: (1) Multi-scale feature sampling: Use multiple convolution kernels of different scales to downsample the input. The step size S of the convolution kernel is set to 2. The convolution operation for the i-th scale is: Where, P i is the feature block obtained by sampling, For a kernel size K i And the convolution operation of step size S, E is the input feature matrix; (2) Multi-scale feature embedding and splicing: Linearly embed the feature blocks of each scale and project them to the specified dimension before splicing. The expression is: Where, T i P i Perform linear transformation to embed the feature blocks of the linear layer, Linear i For linear transformation, the feature is transformed from C i Embedded in D i , T is all T i The splicing result.
7. The multi-feature ocean environment perception method based on the cross-scale attention mechanism according to claim 5 is characterized in that: The specific method of the long and short distance attention module is: Near attention and long-distance attention are used alternately. For the near-distance attention module, it is used for the dependency between adjacent embeddings to capture local information. The feature map with an input size of S×S is adjacently embedded in a local group of G×G. The near-distance attention module performs self-attention calculation in each G×G area. The long-distance attention module is used to process long-distance dependencies and realize the interaction between farther embeddings in the model. The feature map with an input size of S×S is sampled at an interval of I. All embeddings with an interval of I are divided into the same group, and the embeddings in each group will be self-attention calculated. For the long-distance attention module with an input size of S×S, the embeddings are sampled at a fixed interval of I. The specific method of dynamic position bias is as follows: relative position bias represents the relative position of embedding by adding a bias to the embedded attention; nonlinear transformation consists of normalization and activation function RELU and fully connected layer; the input dimension of dynamic position bias is 2, and the dimension of the middle layer is set to D / 4, where D is the embedding dimension; output B i,j is a scalar, encoding i th and j th Embed the relative position features between them and add the dynamic position bias value to the attention score calculation; B i,j =MLP(Δx i,j ,Δy i,j ) Where, (Δx i,j ,Δy i,j ) is the relative coordinate difference between the i-th and j-th feature block units; MLP() is a lightweight multi-layer perceptron, which consists of normalization and activation function RELU and fully connected layers; Where Attn is the attention size, Q,K,V∈R N×d are query, key, and value matrices respectively, is the square root of the dimension of k, B∈R N×N is a dynamically generated bias matrix.
8. The multi-feature ocean environment perception method based on the cross-scale attention mechanism according to claim 5 is characterized in that: The activation value cooling layer consists of a 3×3 depth convolution layer and a normalization layer. For depth convolution, local operations are used to smooth the features. Each channel is convolved separately without introducing information interaction between channels. The convolution kernel size is 3×3, and the features are adjusted in the spatial dimension. For normalization, each channel is independently normalized so that the mean of the feature is 0 and the variance is 1, reducing the activation value amplitude of the feature to a stable range.
9. The multi-feature ocean environment perception method based on the cross-scale attention mechanism according to claim 1 is characterized in that: In step S5, based on the real-time read coherent radar complex value data, the pre-trained multi-feature ocean environment perception model obtained in step S4 is calculated and the wind and wave parameter inversion results in ocean environment perception are output in real time, including: For the test data set obtained in step S2 Input the model trained in step S4, and solve the model to obtain the inversion results of the ocean environment perception time history parameters in, Real and imaginary echo data for testing, D te The measured coherent radar complex-valued echo data matrix used for testing, Y te The model output is the inverted wind and wave parameters [wave height, wave period, wave direction, wind speed, wind direction].
10. A multi-feature ocean environment perception system based on a cross-scale attention mechanism, characterized by: The system is used to control the multi-feature ocean environment perception method based on the cross-scale attention mechanism according to any one of claims 1 to 9, and the system includes: The region selection module is used to select effective regions under multiple working conditions for the measured complex value data of the X-band coherent marine radar; The data partitioning module is used to preprocess the input data, build the corresponding label data set and perform data segmentation, dividing the data into training set and test set; A multi-feature ocean environment perception model building module, which is used to build a multi-feature ocean environment perception model based on a cross-scale attention mechanism; The model training module is used to initialize the model training parameters, input the obtained training set into the multi-feature ocean environment perception model for training, and adjust the model parameters according to the validation set results; The parameter inversion module is used to calculate based on the coherent radar complex value data read in real time, obtain the pre-trained multi-feature ocean environment perception model, and output the wind and wave parameter inversion results in ocean environment perception in real time.
Citation Information
Patent Citations
A method of inverting sea wave parameters for x-band navigation radar based on a new wave dispersion relation band-pass filter
CN103969643B
Ocean wave information inversion system based on X-band radar image
CN116958435A
Phase-resolved wave time history inversion method for non-coherent radar based on deep learning
CN117647808B