Mesoscale eddy detection method and system based on physical prior enhanced deep learning

By employing a deep learning method based on physical prior enhancement, and combining multi-scale visual and physical features, the problems of weak signal omission and complex background interference in mesoscale eddy detection are solved, achieving higher detection accuracy and reliability.

CN121962878AActive Publication Date: 2026-05-01OCEAN UNIV OF CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
OCEAN UNIV OF CHINA
Filing Date
2026-04-03
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies for mesoscale eddy detection suffer from problems such as missed detection in weak signal areas, background noise interference, difficulty in detection under complex sea conditions, and false detection or redundant prediction in dense target scenarios, which affect the accuracy and reliability of detection.

Method used

We employ a deep learning approach based on physical prior enhancement to improve the accuracy and reliability of mesoscale eddy detection through multi-scale visual feature extraction, physical feature enhancement, feature encoding, and physical query guidance.

Benefits of technology

It improves the performance of mesoscale eddy detection, especially under weak signal and complex background conditions, and can more accurately identify and locate mesoscale eddies, reducing the rate of missed detection and false detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962878A_ABST
    Figure CN121962878A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of marine remote sensing data intelligent analysis and deep learning target detection, and particularly relates to a mesoscale eddy detection method and system based on physical prior enhanced deep learning. The method comprises the following steps: inputting a sea level abnormal image to carry out multi-scale visual feature extraction, inputting the multi-scale visual features into a physical feature enhancement module, constructing mesoscale vortex physical priori, and generating comprehensive physical features, physical enhancement features and updated multi-scale features. And inputting the updated multi-scale features into an encoder to obtain memory features. Meanwhile, the physical features and the physical enhancement features are integrated and input into a physical query guide module, and a reference frame and a physical guide query vector are generated; and inputting the memory features, the physical guidance query vector and the reference frame into a decoder and a detection and prediction unit to carry out target decoding and detection and prediction, and outputting a category result and a bounding box of the mesoscale vortex target. And the detection precision and reliability of the mesoscale vortex under the complex ocean background are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent analysis of marine remote sensing data and deep learning target detection technology, specifically involving a method and system for detecting mesoscale eddies based on physical priors and enhanced deep learning. Background Technology

[0002] Mesoscale eddies, as a ubiquitous dynamic phenomenon in the ocean, have a significant impact on key ocean processes such as meridional heat transport, spatial distribution of salinity, material exchange between water bodies, and biogeochemical cycles. Researchers can acquire sea-level anomaly data using remote sensing methods such as satellite altimetry, and use this data to identify and analyze oceanic mesoscale eddies. Therefore, how to achieve automatic and accurate detection of mesoscale eddies from sea-level anomaly remote sensing data has become a key scientific issue in the field of in-depth utilization of ocean remote sensing data and intelligent extraction of marine environmental information.

[0003] Early researchers primarily employed detection methods based on physical indicators. These methods typically construct relevant physical discrimination methods based on sea level anomaly data to identify potentially rotating structural regions in the ocean. They possess certain physical significance and interpretability, and can achieve good identification results when the data quality is high and the target structure is clear. However, these methods are usually quite sensitive to threshold settings and data quality. The detection results are prone to fluctuations under different sea areas, different observation conditions, and different background disturbances, and their stability and adaptability have certain limitations.

[0004] In recent years, with the continuous accumulation of remote sensing observation data and the development of deep learning technology, researchers have begun to introduce deep learning methods into mesoscale eddy detection tasks. By constructing deep neural networks, the mapping relationship between sea level anomaly data and mesoscale eddy targets is learned, which shows good flexibility in complex scenarios and provides a new technical approach for automatic detection of mesoscale eddies.

[0005] However, the above methods still have shortcomings, which will be discussed in detail below: On the one hand, some mesoscale eddy targets in sea level anomaly data have weak signals and are easily affected by background disturbances, leading to missed detections in weak signal areas by the model. Specifically, the intensity of sea surface height anomaly signals of mesoscale eddies varies significantly at different life stages, especially for eddies in the nascent and decaying stages, whose closed contour structures are weak or even incomplete and easily submerged by ocean background noise. Traditional physical index methods are limited by threshold settings and have difficulty effectively identifying such weak signal eddies. Deep learning methods are also prone to misclassifying them as background during feature extraction. In addition, factors such as cloud cover, orbital gaps, and interpolation errors in remote sensing data further exacerbate the difficulty of detecting weak signal areas, limiting the applicability and reliability of existing methods in complex sea conditions.

[0006] On the other hand, non-vortex structures with local morphology highly similar to mesoscale vortices exist in the ocean background, and these structures are prone to false detection or redundant predictions in scenarios with densely distributed targets. Phenomena such as ocean front curvature, serpentine flow patterns in strong current areas, filamentary vortex structures, and local upwelling may present vortex-like isoline closures or rotating flow field characteristics in sea level anomaly data, interfering with the model's ability to distinguish real vortices. Especially in areas with dense vortices, the flow field interaction between adjacent targets can easily lead to blurred boundaries and morphological distortions, making it difficult for the model to accurately distinguish independent targets, resulting in false identification or overlapping detections. These problems not only reduce detection accuracy but also introduce uncertainty into subsequent vortex parameter extraction and physical process analysis, limiting the reliability of existing methods in practical ocean monitoring and scientific applications. Summary of the Invention

[0007] To address the aforementioned technical challenges, this invention proposes a mesoscale eddy detection method and system based on physical prior enhancement deep learning. Using sea level anomaly remote sensing data as input, the method enhances the structural representation capability of mesoscale eddy targets through physical prior information, thereby improving the detectability of the targets. Furthermore, by guiding the detection network to focus on key regions of the mesoscale eddy, the method suppresses interference from similar backgrounds and redundant predictions in dense scenes, thus improving the accuracy and reliability of the mesoscale eddy detection results.

[0008] Mesoscale eddy detection methods based on physical priors and deep learning enhancement include: S1. Multi-scale visual feature extraction: Input sea level anomaly image The image is processed through the backbone network. Hierarchical feature extraction is performed to obtain a multi-scale visual feature set; S2, Enhanced Physical Features: Based on the multi-scale visual feature set, a physical prior is constructed and multi-scale fusion is performed to form a comprehensive physical feature. The comprehensive physical characteristics Enhanced visual features selected from the multi-scale visual feature set Modulation and gating fusion are performed to generate physical enhancement features. And update it to a multi-scale feature set; S3, Feature Coding: The multi-scale feature set is input into the encoder for global interactive modeling to obtain the memory feature set. ; S4, Physical Query Guidance: Based on the aforementioned comprehensive physical characteristics With the physical enhancement features Generate saliency score plot With polarity proxy graph And perform local aggregation and refinement, and the significance score plot after refinement. With polarity proxy graph Candidate center coordinates are obtained through screening and refinement. These candidate center coordinates are then normalized to reference points, and an initial reference frame is generated based on these reference points. Based on the reference point, local features are extracted in the corresponding local neighborhood to generate a physically guided initial query vector. The initial reference frame; As the initial query vector The corresponding reference box; S5. Target Decoding and Detection Prediction: Based on the memory feature set and the initial query vector and corresponding reference boxes Construct the initial query set for the decoder and corresponding reference boxes And perform target decoding and detection prediction to output the detection output set. .

[0009] Furthermore, a physical prior is constructed based on the multi-scale visual feature set, and multi-scale fusion is performed to form a comprehensive physical feature. ,include: S2.1, Physical Prior Construction: S2.1.1 Based on each scale feature in the multi-scale visual feature set, a scalar surrogate field of the corresponding scale is generated through single-channel mapping. The scalar proxy field A fixed discrete operator is used to calculate the five-channel physical prior. The five-channel physical prior is then normalized and mapped to the physical feature space through a learnable mapping layer to obtain the physical features at the corresponding scale. ; S2.2 Multi-scale physical feature fusion: Physical prior features at each scale By resampling, the physical features are aligned to the spatial scale corresponding to the highest semantic level visual features to obtain aligned physical features. The aligned physical features at each scale are then concatenated along the channel dimension to obtain multi-scale aggregated physical features. And obtain comprehensive physical characteristics through fusion mapping. .

[0010] Furthermore, the comprehensive physical characteristics Enhanced visual features selected from the multi-scale visual feature set Modulation and gating fusion are performed to generate physical enhancement features. And updated to a multi-scale feature set, including: S2.3, Physical Enhancement Feature Generation: S2.3.1 Select the highest semantic level feature in the multi-scale visual feature set as the visual feature to be enhanced. The comprehensive physical characteristics For the visual features Modulation is performed to obtain modulated visual features. ; S2.3.2, the aforementioned visual features With the aforementioned comprehensive physical characteristics The features are concatenated in the channel dimension, and the gating coefficients g are generated by gating mapping based on the concatenated features. S2.3.3, Based on the gating coefficient g, the original visual features and modulated visual features Weighted fusion is performed to obtain physical enhancement features , is represented as: ; S2.3.4, the physical enhancement feature Replace the visual features This yields the updated multi-scale feature set.

[0011] Furthermore, the comprehensive physical characteristics For the visual features Modulation is performed to obtain modulated visual features. Including: the comprehensive physical characteristics The mapping generated by modulation parameters produces an effect on visual features. Channel-level affine modulation parameters and parameters and visual features Modulation is performed to obtain modulated visual features. , is represented as: ; in, This indicates element-wise multiplication.

[0012] Furthermore, the multi-scale feature set is input into the encoder for global interactive modeling to obtain a memory feature set. ,include: Each scale feature in the multi-scale feature set is encoded with its corresponding location and then input into the encoder. The encoder uses a multi-layer self-attention and feedforward network to achieve cross-spatial location and cross-scale information interaction, and outputs a memory feature set. , is represented as: ; in, This represents the set of selected feature layer numbers. Indicates the number is The feature map corresponding to the feature layer. This represents the updated multi-scale feature set. Represents the set of encoded memory features. This indicates the encoder.

[0013] Furthermore, based on the aforementioned comprehensive physical characteristics With the physical enhancement features Generate saliency score plot With polarity proxy graph And perform local aggregation and refinement, and the significance score plot after refinement. With polarity proxy graph Candidate center coordinates are obtained through screening and refinement. These candidate center coordinates are then normalized to reference points, and an initial reference frame is generated based on these reference points. Based on the reference point, local features are extracted in the corresponding local neighborhood to generate a physically guided initial query vector. The initial reference frame; As the initial query vector The corresponding reference boxes include: S4.1, Significance and Polarity Response Modeling: Based on comprehensive physical characteristics and physical enhancement features Spatial saliency score map is generated by fusing score mapping and weights. Meanwhile, the comprehensive physical characteristics A polar proxy graph is generated through polarity mapping. ; S4.2, Localized Aggregation and Refining: The saliency scoring map is evaluated using local aggregation operators with different local neighborhood sizes. and the polarity proxy graph Smoothing was performed, and the results of each local aggregation were merged and refined to obtain the refined significance response map. and refined polar response diagram ; S4.3 Candidate Position Filtering: The refined polar response diagram The data is divided into two categories: positive and negative, and score maps are calculated for each category. Candidate seed points are selected based on these score maps to generate a candidate seed point set. Within a preset local neighborhood of each candidate seed point, the neighborhood coordinates are weighted and averaged using the weights obtained after normalizing the neighborhood saliency response values ​​to obtain refined candidate center coordinates. These candidate center coordinates are then normalized to reference points, and initial reference boxes corresponding to these reference points are generated according to a preset initial scale. ; S4.4 Query Initialization: The comprehensive physical characteristics With the physical enhancement features The data is concatenated along the channel dimension, and the query features are obtained through a mapping layer. Within the local neighborhood corresponding to the reference point, features are constructed from the query. Extract local features to generate the physical guidance initial query vector. The initial query vector With the initial reference frame This is the corresponding reference box.

[0014] Furthermore, the refined polar response diagram The data is divided into two categories: positive and negative, and scores are calculated for each category. The resulting graph is shown below: , ; in, This means that only the positive values ​​in the corresponding response graph are retained. This represents element-wise multiplication. This represents the candidate score map for positive polarity. This represents the negative polarity candidate score map.

[0015] Furthermore, based on the memory feature set and the initial query vector and corresponding reference boxes Construct the initial query set for the decoder and corresponding reference boxes And perform target decoding and detection prediction to output the detection output set. ,include: Physically guided initial query vector and the corresponding reference box The target matching query portion is injected into the decoder, while the remaining queries retain their original learnable query representations and reference boxes unchanged; only the target matching query portion is initialized with the aforementioned query vector. and corresponding reference boxes Replace to form the initial query set and corresponding reference boxes The initial query set The corresponding reference box and the memory feature set Decoding is performed by detecting the output set of the prediction head. The detection output set This includes the category prediction results and bounding box prediction results for mesoscale eddy candidate targets.

[0016] A mesoscale eddy detection system based on physical prior enhancement deep learning, the system including a physical feature enhancement module and a physical query guidance module; First, input the sea level anomaly image. The image is processed through the backbone network. Hierarchical feature extraction is performed to obtain a multi-scale visual feature set; Then, the physical feature enhancement module constructs physical priors based on the multi-scale visual feature set and performs multi-scale fusion to form comprehensive physical features. The comprehensive physical characteristics Enhanced visual features selected from the multi-scale visual feature set Modulation and gating fusion are performed to generate physical enhancement features. And update it to a multi-scale feature set; The physical query guidance module is based on the comprehensive physical characteristics. With the physical enhancement features Generate saliency score plot With polarity proxy graph And perform local aggregation and refinement, and the significance score plot after refinement. With polarity proxy graph Candidate center coordinates are obtained through screening and refinement. These candidate center coordinates are then normalized to reference points, and an initial reference frame is generated based on these reference points. Based on the reference point, local features are extracted in the corresponding local neighborhood to generate a physically guided initial query vector. The initial reference frame; As the initial query vector The corresponding reference box; Finally, based on the aforementioned memory feature set and the initial query vector and corresponding reference boxes Construct the initial query set for the decoder and corresponding reference boxes And perform target decoding and detection prediction to output the detection output set. .

[0017] A computer-readable storage medium stores a program for a mesoscale eddy detection method based on physical priors-enhanced deep learning, which, when run on a computer processor, improves the performance of ocean mesoscale eddy detection.

[0018] Compared with the prior art, the advantages of this invention are: To address the issue that visual features are easily interfered with and affect detection results when mesoscale vortex signals are weak and background disturbances are strong, a physical prior feature enhancement module was designed. This module transforms physical information that reflects the characteristics of mesoscale vortex structures into prior expressions that can participate in deep feature learning. It then integrates these prior expressions with visual information in multi-scale features, thereby enhancing the model's ability to perceive real vortex structures and improving its feature expression capabilities under weak target and complex background conditions.

[0019] On the other hand, existing mesoscale eddy detection methods often struggle to accurately focus on the location of mesoscale eddies when faced with densely distributed eddies and similar backgrounds, leading to issues such as detection box offset and missed or false detections. To address this, a physical query guidance module is designed. By introducing physical guidance information consistent with the characteristics of mesoscale eddies during the target localization stage, the model can more effectively focus on reasonable candidate target areas. Through these two innovative designs, this invention significantly improves the performance of ocean mesoscale eddy detection. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the mesoscale eddy detection method and system based on physical priors and deep learning in this embodiment. Figure 2 This is a schematic diagram of the physical feature enhancement module in the mesoscale eddy detection method and system based on physical prior enhancement deep learning in this embodiment. Figure 3 This is a schematic diagram of the physical query guidance module in the mesoscale eddy detection method and system based on physical priors and deep learning in this embodiment. Detailed Implementation

[0022] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0023] Example 1: like Figure 1 As shown, this embodiment provides a mesoscale eddy detection method and system based on physical prior reinforcement deep learning, specifically including the following steps: S1. Multi-scale visual feature extraction: See Figure 1 Input sea level anomaly image The image is processed through the backbone network. Hierarchical feature extraction is performed to obtain a multi-scale visual feature set, represented as follows: ; Among them, sea level anomaly images , and These represent the height and width of the sea level anomaly image, respectively. Represents a set of visual features at multiple scales. This represents the set of selected feature layer numbers. Indicates the number is Feature maps corresponding to feature layers, multi-scale visual features , , and These represent the number of channels, height, and width of the feature map at the corresponding level, respectively. This represents the feature extraction operation of the backbone network.

[0024] In step S1 of this embodiment, the backbone network can be implemented using a visual feature extraction network capable of outputting multi-level feature maps, such as the ResNet series of convolutional neural networks, the Swing Transformer, and other backbone networks.

[0025] S2, Enhanced Physical Features: See Figure 1 and Figure 2 Physical priors are constructed based on multi-scale visual feature sets, and multi-scale fusion is performed to form comprehensive physical features. Comprehensive physical characteristics Enhanced visual features selected from a multi-scale visual feature set Modulation and gating fusion are performed to generate physical enhancement features. And update it to a multi-scale feature set.

[0026] In step S2 of this embodiment, a physical prior is constructed based on a multi-scale visual feature set, and multi-scale fusion is performed to form a comprehensive physical feature. Specifically: S2.1, Physical Prior Construction: S2.1.1 Based on each scale feature in the multi-scale visual feature set, a scalar surrogate field of the corresponding scale is generated through single-channel mapping. Scalar Agent Field A fixed discrete operator is used to compute the five-channel physical prior. The five-channel physical prior is then normalized and mapped to the physical feature space through a learnable mapping layer to obtain the physical features at the corresponding scale. .

[0027] In step S2.1.1 of this embodiment, the fixed discrete operators include: a first-order gradient operator, a second-order Laplace operator, and a local operator for smoothing or constructing a surrogate flow field. The five-channel physical prior includes: a scalar surrogate channel, a gradient magnitude channel, a Laplace response channel, a Laplace sign response channel, and an Okubo–Weiss surrogate channel.

[0028] Scalar Agent Field The five-channel physical prior is calculated using a fixed discrete operator, specifically: first, a smoothing local operator is used to process the scalar surrogate field. Local smoothing is performed to obtain the smoothed scalar surrogate response. Then, the smoothed scalar surrogate response is calculated using the first-order gradient operator. direction and The gradient in the direction is calculated, and the gradient magnitude is calculated. Then, the Laplace response is calculated using the second-order Laplace operator, and the sign of the Laplace response is taken to obtain the Laplace sign response. Among them, the first-order gradient operator can be the Sobel operator, the second-order Laplace operator can be the 3×3 Laplace operator, and the smoothing local operator can be the 3×3 Box smoothing kernel. For the Okubo–Weiss surrogate quantity channel, the surrogate flow field is first constructed based on the spatial gradient of the smoothed scalar surrogate response. Then, the gradient of the surrogate flow field is calculated to obtain the relative vorticity and strain components. Finally, the Okubo–Weiss surrogate quantity is calculated according to the relationship between the strain components and the relative vorticity, and is used as one of the five-channel physical priors.

[0029] The five-channel physical prior is represented as follows: ; in, This represents the physical prior computation function, i.e., the fixed discrete operator mapping function. , representing the five-channel physical prior.

[0030] The five-channel physical prior is normalized and then mapped to the physical feature space through a learnable mapping layer to obtain the physical prior features at the corresponding scale. Specifically, to reduce the differences in numerical ranges between different physical priors and suppress the influence of anomalous responses, the five-channel physical priors are normalized. Normalization can be achieved using median-based, quantile-based, or other normalization methods. Subsequently, a learnable mapping layer maps the normalized five-channel physical priors to the physical feature space, obtaining physical features at the corresponding scale. Forming a multi-scale set of physical features The learnable mapping layer can be implemented using convolutional mapping or linear mapping. , The number of channels representing physical characteristics.

[0031] S2.2 Multi-scale physical feature fusion: To enhance the consistency of physical information at different scales and form a unified physical prior representation, the multi-scale physical feature set obtained in step S2.1 is... Multi-scale fusion is performed. Specifically, scale-based physical prior features... The physical features are aligned to the spatial scale corresponding to the highest semantic level visual features through resampling. The aligned physical features at each scale are then concatenated along the channel dimension to obtain multi-scale aggregated physical features. And obtain comprehensive physical characteristics through fusion mapping. This is used for subsequent visual feature enhancement processing and serves as input for saliency and polarity response modeling and query initialization construction in physical query guidance processing; the resampling method can be implemented using interpolation resampling or other equivalent spatial resolution transformation methods; expressed as: ; in, This indicates a channel splicing operation. This represents a fusion mapping operation, which can be implemented using convolutional mapping or linear mapping. This indicates the physical characteristics after alignment.

[0032] In step S2 of this embodiment, the comprehensive physical characteristics are considered. Enhanced visual features selected from a multi-scale visual feature set Modulation and gating fusion are performed to generate physical enhancement features. And it is updated to a multi-scale feature set, specifically: S2.3, Physical Enhancement Feature Generation: S2.3.1 Select the highest semantic level feature in the multi-scale visual feature set as the visual feature to be enhanced. Comprehensive physical characteristics visual features Modulation is performed to obtain modulated visual features. .

[0033] S2.3.2, Visual Features With comprehensive physical characteristics The features are concatenated in the channel dimension, and the gating coefficients g are generated by gating mapping based on the concatenated features.

[0034] S2.3.3, Based on the gating coefficient g, the original visual features and modulated visual features Weighted fusion is performed to obtain physical enhancement features , is represented as: ; Wherein, the gating coefficient g is related to visual features Dimensionally appropriate weighting coefficients.

[0035] S2.3.4, Physical Enhancement Features Replace visual features This yields the updated multi-scale feature set.

[0036] In step S2.3.1 of this embodiment, the comprehensive physical characteristics are considered. visual features Modulation is performed to obtain modulated visual features. Specifically: Comprehensive physical characteristics The mapping generated by modulation parameters produces an effect on visual features. Channel-level affine modulation parameters and parameters and visual features Modulation is performed to obtain modulated visual features. , is represented as: ; in, Represents element-wise multiplication, visual features , , and These represent the number of channels, width, and height of the visual features, respectively.

[0037] S3, Feature Coding: See Figure 1 and Figure 2 The multi-scale feature set is input into the encoder for global interactive modeling to obtain the memory feature set. .

[0038] Step S3 in this embodiment is specifically as follows: After each scale feature in the multi-scale feature set is encoded with its corresponding location, it is input into the encoder. The encoder achieves cross-spatial and cross-scale information interaction through a multi-layer self-attention and feedforward network, and outputs a memory feature set. , is represented as: ; in, This represents the updated multi-scale feature set. Represents the set of encoded memory features. This indicates the encoder.

[0039] S4, Physical Query Guidance: See Figure 1 and Figure 3 Based on comprehensive physical characteristics With physical enhancement features Generate saliency score plot With polarity proxy graph And perform local aggregation and refinement, and the significance score plot after refinement. With polarity proxy graph Candidate center coordinates are obtained through screening and refinement. These coordinates are then normalized to reference points, and an initial reference box is generated based on these reference points. Based on the reference point, local features are extracted within the corresponding local neighborhood to generate a physically guided initial query vector. ; where the initial reference box As the initial query vector The corresponding reference box.

[0040] Step S4 in this embodiment is specifically as follows: S4.1, Significance and Polarity Response Modeling: Based on comprehensive physical characteristics and physical enhancement features Spatial saliency score map is generated by fusing score mapping and weights. Simultaneously, considering physical characteristics A polar proxy graph is generated through polarity mapping. .

[0041] Specifically, comprehensive physical characteristics Physical saliency maps are generated through physical saliency score mapping, thereby enhancing physical features. A visual saliency map is generated through a visual saliency rating map. This rating map can be implemented using a combination of convolutional mapping and nonlinear activation, such as a 1×1 convolutional mapping with a sigmoid activation function. The physical saliency map and the visual saliency map are then fused according to preset weights or learnable weights to obtain the final spatial saliency rating map. , , and These represent the width and height of the spatial saliency scoring map, respectively, while also considering the comprehensive physical characteristics. A polar proxy graph is generated through polarity mapping. Polarity mapping can be achieved by combining convolutional mapping with nonlinear activation, such as using 1×1 convolutional mapping with hyperbolic tangent activation function to characterize cyclones or anticyclones corresponding to different spatial locations; polarity surrogate maps are used for polarity balance selection in subsequent candidate location screening.

[0042] S4.2, Localized Aggregation and Refining: The saliency scoring map is evaluated using local aggregation operators with different local neighborhood sizes. and the polarity proxy graph Smoothing was performed, and the results of each local aggregation were merged and refined to obtain the refined significance response map. and refined polar response diagram Among them, fusion can be achieved by averaging, weighted summation or mapping after concatenation; local aggregation operators can be achieved by local average pooling.

[0043] S4.3 Candidate Position Filtering: Refined polar response diagram The samples were divided into positive and negative polarity categories, and score graphs were calculated for each category. The sample with the highest score was selected from the score graphs. 1. Positions are selected as candidate seed points, where... , This represents the total number of queries by the detector. After merging candidate seed points, a candidate seed point set is generated. To improve the localization accuracy of candidate locations, a weighted average of the neighborhood coordinates is applied within a preset local neighborhood of each candidate seed point, using the weights obtained after normalizing the neighborhood saliency response values. This yields refined center coordinates, which are then normalized to a reference point. An initial reference box corresponding to the reference point is generated based on a preset initial scale. Initial reference box use Formal representation; in which, and Represents the normalized center coordinates. and This represents the initial width and height of the reference box. The preset initial scale can be a fixed default width and height or a preset proportional width and height. The preset local neighborhood can be a local window of a fixed size. Normalization can be achieved using softmax normalization or other equivalent normalization methods.

[0044] S4.4 Query Initialization: Comprehensive physical characteristics With physical enhancement features The data is concatenated along the channel dimension, and the query features are obtained through a mapping layer. Features are constructed from the query within the local neighborhood corresponding to the reference point. Extract local features to generate the physical guidance initial query vector. Initialize query vector With the initial reference box This is the corresponding reference box.

[0045] In step S4.2 of this embodiment, the refined polarity response diagram The data is divided into two categories: positive and negative, and scores are calculated for each category. The resulting graph is shown below: , ; in, This means that only the positive values ​​in the corresponding response graph are retained. This represents element-wise multiplication. This represents the candidate score map for positive polarity. This represents the negative polarity candidate score map.

[0046] S5. Target Decoding and Detection Prediction: See Figure 1 Based on the memory feature set and initializing query vector and corresponding reference boxes Construct the initial query set for the decoder and corresponding reference boxes And perform target decoding and detection prediction to output the detection output set. .

[0047] Step S5 in this embodiment is specifically as follows: Physically guided initial query vector and corresponding reference boxes The part used for target matching queries is injected into the decoder, while the other queries retain their original learnable query representations and reference boxes. Specifically, denoised queries are introduced during the training phase. These denoised queries maintain their original generation method and are only initialized for the part used for target matching queries. and corresponding reference boxes Replace to form the initial query set and corresponding reference boxes Initial query set Corresponding reference box and memory feature set Decode the data and detect the output set by detecting the prediction head. The detection output set includes the category prediction results and bounding box prediction results for mesoscale eddy candidate targets; represented as: , in, Represents the decoder function. This represents the detection prediction head function, with the bounding box using normalized parameter form. This indicates that the detection results of the mesoscale eddy target are obtained.

[0048] Model training: During the training phase, a one-to-one matching method is used to establish the optimal matching relationship between the prediction set and the real labeled set. The matching method is implemented using Hungarian matching. The matching cost is composed of the classification cost and the bounding box regression cost. The detection loss is calculated only for the prediction results that are successfully matched, and the prediction results that are not matched are regarded as background.

[0049] Given a training set The optimization objective of the detection network is to minimize the joint detection loss between the predicted results and the ground truth labels. The loss function is expressed as: , in, This represents the set of trainable parameters for the detection network. Represents classification loss. This represents the regression loss for the bounding box parameters. This represents the bounding box overlap constraint loss. , and This represents the weight coefficient of the corresponding loss term; the classification loss can be focal loss or other classification loss forms suitable for imbalanced class scenarios; based on the above joint detection loss, the AdamW gradient optimization algorithm is used to iteratively update the detection network parameters. A learning rate scheduling strategy is then employed to complete model training.

[0050] During training, the detection network's detection results on the validation data are evaluated, and the network model with the best detection performance is saved as the final mesoscale eddy detection model.

[0051] Example 2: This embodiment also provides a mesoscale eddy detection system based on physical prior enhancement deep learning, the system including a physical feature enhancement module and a physical query guidance module; First, input the sea level anomaly image. The image is processed through the backbone network. Hierarchical feature extraction is performed to obtain a multi-scale visual feature set; Then, the physical feature enhancement module constructs physical priors based on multi-scale visual feature sets and performs multi-scale fusion to form comprehensive physical features. Comprehensive physical characteristics Enhanced visual features selected from a multi-scale visual feature set Modulation and gating fusion are performed to generate physical enhancement features. And update it to a multi-scale feature set; Based on comprehensive physical characteristics With physical enhancement features Generate saliency score plot With polarity proxy graph And perform local aggregation and refinement, and the significance score plot after refinement. With polarity proxy graph Candidate center coordinates are obtained through screening and refinement. These coordinates are then normalized to reference points, and an initial reference box is generated based on these reference points. Based on the reference point, local features are extracted within the corresponding local neighborhood to generate a physically guided initial query vector. ; where the initial reference box As the initial query vector The corresponding reference box; Finally, based on the memory feature set and initializing query vector and corresponding reference boxes Construct the initial query set for the decoder and corresponding reference boxes And perform target decoding and detection prediction to output the detection output set. .

[0052] Example 3: Experimental verification: To verify the technical effectiveness of this invention, this embodiment uses a self-built mesoscale eddy annotation dataset as the training and testing dataset. The dataset is constructed based on multi-source fusion satellite sea level anomaly data products from the Copernicus Ocean Environment Monitoring Service, and mesoscale eddy targets in the data are manually annotated. The annotation is based on closed streamline features and mesoscale eddy size range constraints, and bounding boxes are used to annotate the eddy targets. The mesoscale eddy annotation dataset constructed in the above manner is used to verify the accuracy and effectiveness of the method of this invention in mesoscale eddy detection tasks.

[0053] Specifically, this embodiment selects data from a preset sea area, with a latitude range of 5°N to 37°N, a longitude range of 105°E to 125°E, a spatial resolution of 0.25° × 0.25°, and a temporal resolution of daily scale. The data spans from August 20, 2016 to February 9, 2022, and includes 2000 sea level anomaly images. To avoid information leakage between training and test data, the dataset is divided according to time. Data from August 20, 2016 to February 9, 2021 is used as the training set, containing 1635 images; data from February 10, 2021 to February 9, 2022 is used as the test set, containing 365 images. A total of 182,678 mesoscale eddies are labeled in the above data, including 93,250 anticyclones and 89,428 cyclones; the training set contains 156,937 mesoscale eddies, and the test set contains 25,741 mesoscale eddies.

[0054] The experiment was conducted on two NVIDIA RTX 2080Ti GPUs, using the AdamW optimizer to train the network for 36 epochs with a batch size of 2. During training, the learning rate was initialized to [value missing]. The weight decay coefficient is set to 0.05.

[0055] To verify the effectiveness of this invention, this embodiment selected Deformable-DETR, DINO, Lite-DETR, RT-DETR, and DEIM as comparison methods. Deformable-DETR, DINO, Lite-DETR, and RT-DETR are representative Transformer detection methods in the field of object detection, while DEIM is a recently proposed method and the best-performing existing comparison method selected on this dataset. Experiments were conducted on each comparison method under the same dataset partitioning and uniform training settings. The experimental results are detailed in Table 1 below.

[0056] Table 1 Comparison of experimental evaluation indicators for each method .

[0057] Among all evaluation metrics, the method of this invention exhibits the best performance. On the AP50 metric, the method of this invention achieves 84.6%, higher than the selected comparative methods; on the mAP metric, the method of this invention achieves 50.2%, also the best result. In particular, compared to DEIM, the best-performing existing comparative method selected for this dataset, the method of this invention still improves the AP50 and mAP metrics by 1.9 percentage points and 0.5 percentage points, respectively. These results demonstrate that in the mesoscale eddy target detection task, the method of this invention has a stronger detection effect than existing methods.

[0058] To further verify the effectiveness of the physical feature enhancement module and the physical query guidance module in this invention, module ablation experiments were also conducted. The basic detection model was a Transformer-based visual detection model without the physical feature enhancement module and the physical query guidance module. Since the physical query guidance module needs to generate physical guidance queries based on the output of the physical feature enhancement module, it cannot be used independently of the physical feature enhancement module. Therefore, this embodiment sets up three schemes for comparison and verification: the basic detection model, the detection model with the physical feature enhancement module, and the detection model with both the physical feature enhancement module and the physical query guidance module. The experimental results are shown in Table 2 below.

[0059] Table 2 Comparison of experimental evaluation indicators for various ablation protocols .

[0060] As shown in Table 2, after introducing the physical feature enhancement module into the basic detection model, the AP50 increased from 83.6% to 84.1%, and the mAP increased from 47.7% to 49.5%, indicating that the physical feature enhancement module can improve the detection effect of mesoscale vortex targets. Furthermore, after introducing the physical query guidance module, the AP50 further increased to 84.6%, and the mAP further increased to 50.2%, indicating that the physical query guidance module can further improve the detection effect based on the physical features provided by the physical feature enhancement module. The above results show that both the physical feature enhancement module and the physical query guidance module in this invention can improve the detection effect of mesoscale vortex targets, and the physical query guidance module can further improve the detection effect based on the physical feature enhancement module.

[0061] Example 4:

[0062] In another embodiment of the present invention, a computer-readable storage medium is provided, which stores a computer program; when the program is executed by a processor, it can realize the mesoscale eddy detection method based on physical prior enhancement deep learning as described in the foregoing embodiments; the computer-readable storage medium can be a disk, optical disk, read-only memory, flash memory, solid-state drive or other forms of non-transitory storage device, all of which can support the operation and reproduction of the method in different computing platforms and hardware environments.

[0063] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should be protected by the present invention.

Claims

1. A mesoscale eddy detection method based on physical prior reinforcement deep learning, characterized in that, Includes the following steps: S1. Multi-scale visual feature extraction: Input sea level anomaly image The image is processed through the backbone network. Hierarchical feature extraction is performed to obtain a multi-scale visual feature set; S2, Physical Feature Enhancement: Based on the multi-scale visual feature set, a physical prior is constructed and multi-scale fusion is performed to form a comprehensive physical feature. The comprehensive physical characteristics Enhanced visual features selected from the multi-scale visual feature set Modulation and gating fusion are performed to generate physical enhancement features. And update it to a multi-scale feature set; S3, Feature Coding: The multi-scale feature set is input into the encoder for global interactive modeling to obtain the memory feature set. ; S4, Physical Query Guidance: Based on the aforementioned comprehensive physical characteristics With the physical enhancement features Generate saliency score plot With polarity proxy graph And perform local aggregation and refinement, and the significance score plot after refinement. With polarity proxy graph Candidate center coordinates are obtained through screening and refinement. These candidate center coordinates are then normalized to reference points, and an initial reference frame is generated based on these reference points. Based on the reference point, local features are extracted in the corresponding local neighborhood to generate a physically guided initial query vector. ; wherein, the initial reference frame As the initial query vector The corresponding reference box; S5. Target Decoding and Detection Prediction: Based on the memory feature set and the initial query vector and corresponding reference boxes Construct the initial query set for the decoder and corresponding reference boxes And perform target decoding and detection prediction to output the detection output set. .

2. The method according to claim 1, characterized in that, Based on the multi-scale visual feature set, a physical prior is constructed and multi-scale fusion is performed to form a comprehensive physical feature. ,include: S2.1, Physical Prior Construction: S2.1.1 Based on each scale feature in the multi-scale visual feature set, a scalar surrogate field of the corresponding scale is generated through single-channel mapping. The scalar proxy field A fixed discrete operator is used to calculate the five-channel physical prior. The five-channel physical prior is then normalized and mapped to the physical feature space through a learnable mapping layer to obtain the physical features at the corresponding scale. ; S2.2 Multi-scale physical feature fusion: Physical prior features at each scale By resampling, the physical features are aligned to the spatial scale corresponding to the highest semantic level visual features to obtain aligned physical features. The aligned physical features at each scale are then concatenated along the channel dimension to obtain multi-scale aggregated physical features. And obtain comprehensive physical characteristics through fusion mapping. .

3. The method according to claim 1, characterized in that, The comprehensive physical characteristics Enhanced visual features selected from the multi-scale visual feature set Modulation and gating fusion are performed to generate physical enhancement features. And updated to a multi-scale feature set, including: S2.3, Physical Enhancement Feature Generation: S2.3.1 Select the highest semantic level feature in the multi-scale visual feature set as the visual feature to be enhanced. The comprehensive physical characteristics Regarding the visual features Modulation is performed to obtain modulated visual features. ; S2.3.2, the aforementioned visual features With the aforementioned comprehensive physical characteristics The features are concatenated in the channel dimension, and the gating coefficients g are generated by gating mapping based on the concatenated features. S2.3.3, Based on the gating coefficient g, the original visual features and modulated visual features Weighted fusion is performed to obtain physical enhancement features , is represented as: ; S2.3.4, the physical enhancement feature Replace the visual features This yields the updated multi-scale feature set.

4. The method according to claim 3, characterized in that, The comprehensive physical characteristics Regarding the visual features Modulation is performed to obtain modulated visual features. Including: the comprehensive physical characteristics The mapping generated by modulation parameters produces an effect on visual features. Channel-level affine modulation parameters and parameters and visual features Modulation is performed to obtain modulated visual features. , is represented as: ; in, This indicates element-wise multiplication.

5. The method according to claim 1, characterized in that, The multi-scale feature set is input into the encoder for global interactive modeling to obtain the memory feature set. ,include: Each scale feature in the multi-scale feature set is encoded with its corresponding location and then input into the encoder. The encoder uses a multi-layer self-attention and feedforward network to achieve cross-spatial location and cross-scale information interaction, and outputs a memory feature set. , is represented as: ; in, This represents the set of selected feature layer numbers. Indicates the number is The feature map corresponding to the feature layer. This represents the updated multi-scale feature set. Represents the set of encoded memory features. This indicates the encoder.

6. The method according to claim 1, characterized in that, Based on the aforementioned comprehensive physical characteristics With the physical enhancement features Generate saliency score plot With polarity proxy graph And perform local aggregation and refinement, and the significance score plot after refinement. With polarity proxy graph Candidate center coordinates are obtained through screening and refinement. These candidate center coordinates are then normalized to reference points, and an initial reference frame is generated based on these reference points. Based on the reference point, local features are extracted in the corresponding local neighborhood to generate a physically guided initial query vector. ; wherein, the initial reference frame As the initial query vector The corresponding reference boxes include: S4.1, Significance and Polarity Response Modeling: Based on comprehensive physical characteristics and physical enhancement features Spatial saliency score map is generated by fusing score mapping with weights. Meanwhile, the comprehensive physical characteristics A polar proxy graph is generated through polarity mapping. ; S4.2, Localized Aggregation and Refining: The saliency scoring map is evaluated using local aggregation operators with different local neighborhood sizes. and the polarity proxy graph Smoothing was performed, and the results of each local aggregation were merged and refined to obtain the refined significance response map. and refined polar response diagram ; S4.3 Candidate Position Filtering: The refined polar response diagram The data is divided into two categories: positive and negative, and score maps are calculated for each category. Candidate seed points are selected based on these score maps to generate a candidate seed point set. Within a preset local neighborhood of each candidate seed point, the neighborhood coordinates are weighted and averaged using the weights obtained after normalizing the neighborhood saliency response values ​​to obtain refined candidate center coordinates. These candidate center coordinates are then normalized to reference points, and initial reference boxes corresponding to these reference points are generated according to a preset initial scale. ; S4.4 Query Initialization: The comprehensive physical characteristics With the physical enhancement features The data is concatenated along the channel dimension, and the query features are constructed through a mapping layer. Within the local neighborhood corresponding to the reference point, features are constructed from the query. Extract local features to generate the physical guidance initial query vector. The initial query vector With the initial reference frame This is the corresponding reference box.

7. The method according to claim 6, characterized in that, The refined polar response diagram The data is divided into two categories: positive and negative, and scores are calculated for each category. The resulting graph is shown below: , ; in, This means that only the positive values ​​in the corresponding response graph are retained. This represents element-wise multiplication. This represents the candidate score map for positive polarity. This represents the negative polarity candidate score map.

8. The method according to claim 1, characterized in that, Based on the memory feature set and the initial query vector and corresponding reference boxes Construct the initial query set for the decoder and corresponding reference boxes And perform target decoding and detection prediction to output the detection output set. ,include: Physically guided initial query vector The corresponding reference boxes are injected into the decoder for the target matching query part, while the remaining queries retain the original learnable query representation and reference boxes unchanged; only the target matching query part is initialized with the aforementioned query vector. and corresponding reference boxes Replace to form the initial query set and corresponding reference boxes The initial query set The corresponding reference box and the memory feature set Decoding is performed by detecting the output set of the prediction head. The detection output set This includes the category prediction results and bounding box prediction results for mesoscale eddy candidate targets.

9. A mesoscale eddy detection system based on physical prior reinforcement deep learning, characterized in that, The system includes a physical feature enhancement module and a physical query guidance module; First, input the sea level anomaly image. The image is processed through the backbone network. Hierarchical feature extraction is performed to obtain a multi-scale visual feature set; Then, the physical feature enhancement module constructs physical priors based on the multi-scale visual feature set and performs multi-scale fusion to form comprehensive physical features. The comprehensive physical characteristics Enhanced visual features selected from the multi-scale visual feature set Modulation and gating fusion are performed to generate physical enhancement features. And update it to a multi-scale feature set; The physical query guidance module is based on the comprehensive physical characteristics. With the physical enhancement features Generate saliency score plot With polarity proxy graph And perform local aggregation and refinement, and the significance score plot after refinement. With polarity proxy graph Candidate center coordinates are obtained through screening and refinement. These candidate center coordinates are then normalized to reference points, and an initial reference frame is generated based on these reference points. Based on the reference point, local features are extracted in the corresponding local neighborhood to generate a physically guided initial query vector. ; wherein, the initial reference frame As the initial query vector The corresponding reference box; Finally, based on the aforementioned memory feature set and the initial query vector and corresponding reference boxes Construct the initial query set for the decoder and corresponding reference boxes And perform target decoding and detection prediction to output the detection output set. .

Citation Information

Patent Citations

  • Marine mesoscale eddy detection method and system based on target segmentation

    CN115953394A

  • Graph neural network sea surface temperature prediction method based on ocean observation physical enhancement data, electronic equipment and storage medium

    CN118709517A

  • Double-spectrum target detection method based on physical prior constraint

    CN121214067A

  • Monocular 3D target detection method and system based on comparative learning

    CN121686053A

  • Multi-view perception method for 3D point position coding and memory queue time sequence modeling based on depth prior

    CN121746814A