Multimodal ocean data evaluation methods, systems, and storage media for large models

By constructing a multimodal ocean dataset and performing spatiotemporal alignment and physical constraint, the problems of spatiotemporal inconsistency in multimodal ocean data and lack of physical basis for evaluation were solved, and high-precision evaluation of large models in the ocean field was achieved.

CN122046266BActive Publication Date: 2026-07-17GUANGDONG LABORATORY OF SOUTHERN OCEAN SCIENCE AND ENGINEERING (GUANGZHOU)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG LABORATORY OF SOUTHERN OCEAN SCIENCE AND ENGINEERING (GUANGZHOU)
Filing Date
2026-04-17
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing technologies, multimodal ocean data processing methods have failed to achieve deep integration, resulting in fragmented spatiotemporal correlations, insufficient accuracy and objectivity in evaluation results, and a lack of physical basis, leading to misjudgments in ocean event identification.

Method used

A multimodal ocean dataset is constructed, and high-precision spatiotemporal alignment is performed. Combined with physical constraints, feature information is extracted through a large model and consistency judgment is performed to generate comprehensive evaluation results.

Benefits of technology

It improves the accuracy and objectivity of multimodal ocean data evaluation, ensures the scientific validity and reliability of evaluation results, and achieves deep fusion and accurate identification of different modal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122046266B_ABST
    Figure CN122046266B_ABST
Patent Text Reader

Abstract

This application relates to the field of marine monitoring technology and discloses a method, system, and storage medium for evaluating multimodal marine data for large-scale models. The method includes: constructing a multimodal marine dataset containing text reports, remote sensing images, monitoring videos, and marine event information; performing spatiotemporal alignment processing on the dataset to obtain a registered multimodal marine dataset; extracting evaluation feature information from the registered data using a pre-set large-scale model; calculating a physical conformity score based on the evaluation feature information; determining preliminary event categories of marine phenomena based on the physical conformity score; calculating event matching probability values ​​based on the preliminary event categories; fusing the physical conformity score and event matching probability values ​​to determine the final event category and confidence level, and integrating to generate a comprehensive evaluation result. This application solves the problems of spatiotemporal inconsistency in multimodal data and lack of physical basis for evaluation by multimodal fusion, spatiotemporal alignment, and physical constraint, thereby improving the accuracy and reliability of the evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of marine monitoring technology, and in particular to a method, system and storage medium for evaluating multimodal marine data for large models. Background Technology

[0002] Marine monitoring is a core support for marine resource development, environmental governance, and disaster early warning, and the accurate analysis of multimodal marine data has become a key technology. The development of large-scale models has provided new pathways for multi-source data processing. The marine field urgently needs adapted professional evaluation methods to achieve effective evaluation and accurate identification of marine events for multimodal marine data such as text reports, remote sensing images, and surveillance videos, supporting the practical application of large-scale models in the marine field.

[0003] Current technologies for processing marine data often employ single-modal independent analysis methods, performing simple semantic parsing and keyword extraction on text reports, or separate visual feature extraction on remote sensing images and surveillance videos. The processing results for each modality are independent, failing to achieve deep fusion processing of multimodal marine data and hindering the utilization of the complementary value between environmental parameter information from text, spatial feature information from remote sensing images, and dynamic evolution information from surveillance videos. While some multimodal processing methods attempt to integrate different types of marine data, they only perform simple format unification and data stitching, without establishing a standardized spatiotemporal reference coordinate system for high-precision registration. They rely on manually labeled time points and spatial locations for rough correspondence, commonly resulting in timestamp resolution bias and inconsistent spatial projection transformations. This leads to fragmented spatiotemporal correlations among multimodal data, making it impossible to achieve accurate correspondence between different modalities under the same marine event. Meanwhile, existing methods lack constraints based on physical laws such as ocean hydrodynamics, relying solely on surface-level features for simple similarity matching without verifying the physical validity of feature information using environmental parameters like wind and currents. This easily leads to misclassification of natural ocean phenomena such as wind-generated ripples, natural oil slicks, and Langmuir circulation as target ocean events like oil spills or ship anomalies, resulting in low accuracy. Furthermore, traditional evaluation methods lack standardized feature extraction and comprehensive verification processes, and there is no unified standard for consistent judgment of ocean data feature information output by large models, nor a multi-dimensional event matching system. Evaluation results are determined solely by single feature similarity, making the evaluation process highly subjective and lacking quantitative scoring and verification mechanisms. Consequently, the results lack objectivity and reliability, failing to provide a scientific and effective evaluation basis for the application of large models in the ocean field.

[0004] To address the above deficiencies, this application combines the construction of multimodal ocean datasets, high-precision spatiotemporal alignment, physical constraints, and multimodal information matching to solve the problems of spatiotemporal inconsistency in multimodal ocean data, lack of physical basis for evaluation, and poor feature fusion, thereby improving the evaluation accuracy, objectivity, and reliability of large-scale model processing of ocean data. Summary of the Invention

[0005] This application provides a method, system, and storage medium for evaluating multimodal ocean data for large models, which solves the problems of spatiotemporal inconsistency of multimodal ocean data, lack of physical basis for evaluation, and poor feature fusion, and improves the accuracy, objectivity, and reliability of evaluation results for ocean data processed by large models.

[0006] Firstly, this application provides a method for evaluating multimodal ocean data for large models, the method comprising: S1. Construct a multimodal ocean dataset, which includes text reports, remote sensing images, surveillance videos, and corresponding ocean event information; S2. Perform spatiotemporal alignment processing on the multimodal ocean dataset to obtain the registered multimodal ocean dataset; S3. Extract feature information from the registered multimodal ocean dataset using a preset large model to obtain evaluation feature information, which includes environmental parameters, image region features, and dynamic texture sequences. S4. The physical law constraints built into the preset large model are used to make a consistency judgment on the evaluation feature information to obtain a physical conformity score. S5. Determine the preliminary event category based on the physical compliance score and the evaluation feature information; S6. Using the preliminary event category as the matching basis, perform multimodal information matching and comparison between the evaluation feature information and the marine event information to obtain the event matching probability value; S7. A comprehensive evaluation result of the multimodal ocean dataset is generated by comprehensively analyzing and verifying the physical conformity score and the event matching probability value.

[0007] Secondly, this application provides a multimodal ocean data evaluation system for large models, used to implement the aforementioned multimodal ocean data evaluation method for large models, the system comprising: The data acquisition module is used to construct a multimodal ocean dataset, which includes text reports, remote sensing images, surveillance videos, and corresponding ocean event information. The spatiotemporal alignment module is used to perform spatiotemporal alignment processing on the multimodal ocean dataset to obtain a registered multimodal ocean dataset. The feature extraction module is used to extract feature information from the registered multimodal ocean dataset using a preset large model to obtain evaluation feature information, which includes environmental parameters, image region features, and dynamic texture sequences. The physical judgment module is used to make a consistency judgment on the evaluation feature information based on the physical law constraints built into the preset large model, and obtain a physical conformity score. The category determination module is used to determine the preliminary event category based on the physical conformity score and the evaluation feature information; The information matching module is used to perform multimodal information matching and comparison between the evaluation feature information and the marine event information based on the preliminary event category to obtain the event matching probability value; The comprehensive analysis module is used to perform comprehensive analysis and verification using the physical conformity score and the event matching probability value, and generate a comprehensive evaluation result for the multimodal ocean dataset.

[0008] Thirdly, this application provides a computer-readable storage medium storing instructions that, when executed by a processor, implement the aforementioned multimodal ocean data evaluation method for large models.

[0009] This application proposes a method, system, and storage medium for evaluating multimodal ocean data for large-scale models. It addresses the problems of spatiotemporal inconsistency in multimodal ocean data, lack of physical basis for evaluation, and poor feature fusion, thereby improving the accuracy, objectivity, and reliability of evaluation results for large-scale models processing ocean data. Compared with existing technologies, the beneficial effects of this application's technical solution are at least as follows: First, by constructing a multimodal ocean dataset that includes text reports, remote sensing images, monitoring videos, and corresponding ocean event information, the cleaning and standardization of multi-source raw data is completed, realizing the systematic integration of multimodal ocean data and solving the problem that traditional single-modal analysis methods cannot leverage the complementary value of multi-source data information.

[0010] Second, by establishing a spatiotemporal reference coordinate system and combining it with a cross-modal spatiotemporal alignment model, high-precision spatiotemporal registration and correction of multimodal ocean datasets are performed, eliminating temporal and spatial matching deviations between data of different modalities. This solves the problem that traditional multimodal integration methods only perform simple splicing and break the spatiotemporal correlation of data.

[0011] Third, by pre-setting a large model to extract environmental parameters, image region features, and dynamic texture sequences from multimodal ocean data and integrating cross-modal features, a standardized evaluation feature information extraction process is formed, which solves the problems of traditional methods having no unified feature extraction standards and strong evaluation subjectivity.

[0012] Fourth, based on the physical constraints built into the pre-set large model, the physical consistency of the evaluation feature information is verified by combining the Stokes drift model and a physical conformity score is generated, providing a physical basis for the determination of marine events and solving the problem that traditional methods rely solely on surface feature matching and are prone to misjudging natural phenomena.

[0013] Fifth, preliminary event categories are defined through physical conformity scoring, and the temporal morphological evolution characteristics of suspected target areas are extracted and multimodal matched with marine event information. Combined with weighted fusion calculation, accurate event determination is achieved, and a standardized multidimensional event matching system is established, which solves the problems of traditional methods lacking quantitative verification mechanisms and having insufficient objectivity of evaluation results. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a flowchart illustrating the multimodal ocean data evaluation method for large models in this application; Figure 2 This is a schematic diagram of the structure of the large-scale model for multimodal ocean data evaluation in this application; Figure 3 This is a graph showing the comparison results of ROC curves in this application; Figure 4 This is a graph showing the performance stability comparison results of 20 independent experiments in this application; Figure 5 This is a schematic diagram of the system structure of the multimodal ocean data evaluation method for large models in this application. Detailed Implementation

[0016] This application provides a method, system, and storage medium for evaluating multimodal ocean data for large models. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use of terms can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0017] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of a multimodal ocean data evaluation method for large models in this application includes: S1. Construct a multimodal ocean dataset, which includes text reports, remote sensing images, surveillance videos, and corresponding ocean event information.

[0018] In one specific embodiment, performing step S1 includes the following steps: Collect multi-source raw data in the marine field, including text reports, remote sensing images, and surveillance videos; Data cleaning is performed on the multi-source raw data to remove invalid and noisy data, resulting in standardized multi-source ocean data. For each type of data in the standardized multi-source ocean data, the corresponding ocean event information is labeled. The ocean event information includes the ocean event type, the spatiotemporal information of the event occurrence, and the event characteristic description. Standardized multi-source ocean data is correlated and matched with corresponding ocean event information, then integrated and packaged to generate a multimodal ocean dataset.

[0019] Specifically, multi-source raw data in the marine field were collected, including text reports, remote sensing images, and surveillance videos. The collected text reports included marine station observation records, offshore operation logs, and marine environmental assessment documents. Remote sensing images were obtained from ocean satellite multispectral imagers and aerial remote sensing equipment, while surveillance videos were collected from offshore buoy monitoring cameras, ship-mounted recording equipment, and nearshore monitoring station camera systems. All three types of data were acquired within the same marine monitoring area and time frame, ensuring the data's relevance to the actual marine environment from the source. After collection, the multi-source raw data underwent data cleaning. For text reports, regular expression matching algorithms were used to remove invalid text lacking valid environmental parameters and spatiotemporal identifiers. Noise filtering algorithms were used to remove noise interference data such as garbled characters and duplicate fields caused by input errors and transmission interference. The matching degree adopts a weighted calculation method, and the calculation formula is: matching degree = feature item matching rate × 0.6 + feature value effectiveness rate × 0.4, with a value range of 0~1. The feature item matching rate is the ratio of the number of valid data feature items matched in the text to the total number of feature items. The feature value effectiveness rate is the ratio of the number of feature items that are successfully matched and whose values ​​are within a reasonable business range to the total number of matched feature items. Among them, the valid data features are the core effective information set consisting of essential spatiotemporal identification features (time stamp, latitude and longitude, monitoring station code, etc.) and marine environmental parameter features (wind speed, wave height, water temperature, salinity, etc.) required for marine text. For example, the matching threshold is 0.85. When the matching degree between the text field and the valid data features is lower than this threshold, it is judged as invalid data and is removed. For remote sensing images, a median filtering algorithm is used for noise removal. For example, the filtering window is set to 3×3. At the same time, invalid images with blurred images or a pixel missing ratio of more than 15% are removed by pixel value detection, and the image pixels are normalized to the standard grayscale range of 0-255. For surveillance videos, a frame difference method is used to detect and remove invalid video segments with black screens or frame rates lower than 10fps. A Gaussian filtering algorithm is used to filter salt-and-pepper noise in the video. For example, the filtering kernel function is set to 5×5, and the resolution of the video frames is uniformly adjusted to 1920×1080. After the above targeted processing by different algorithms, various types of raw data are transformed into standardized multi-source marine data with unified format and valid information. The text report forms a structured field-based document, the remote sensing image becomes a standardized digital image file, and the surveillance video is transformed into video stream data with unified parameters, realizing a one-to-one correspondence transformation of multi-source data from the raw state to the standardized state.

[0020] After data standardization, marine event information was labeled for each type of data in the standardized multi-source marine data. The labeling process relied on the marine domain labeling system, labeling each standardized text report, each standardized remote sensing image, and each standardized monitoring video with corresponding marine event information. Marine event types included oil spills, ship anomalies, ocean circulation, red tides, and other marine phenomena. The spatiotemporal information of the events was labeled with specific UTC timestamps and spatial coordinates based on the WGS84 coordinate system. Event feature descriptions recorded the specific characteristics of the marine phenomena presented in the corresponding data. For example, text reports labeled parameters such as wind speed, current direction, and sea surface temperature; remote sensing images labeled the shape, area, and pixel features of suspected target areas; and monitoring videos labeled the trajectory and speed of dynamic sea surface changes. Multimodal labeling association rules were used during the labeling process. The rules included: firstly, unique identifier association, assigning a globally unique marine event identifier code to each modality of data corresponding to the same marine event, embedding it in the header of all labeled information as the core index, and realizing the association of a single event with a unique identifier. The system employs the following criteria: 1. **Identification and Multimodal Co-identification:** 2. **Consistent Spatiotemporal Information:** Event occurrence times in all modal data annotations are converted to UTC standard timestamps, and spatial information is annotated using the WGS84 geographic coordinate system. The spatiotemporal range of annotations in remote sensing images and surveillance videos must completely cover the text report, with deviations not exceeding preset thresholds of ≤30 seconds in time and ≤0.01 degrees in space. 3. **Unified Event Categories:** Marine event types in all modal data annotations strictly adhere to the unified marine classification standards. Inconsistent or ambiguous annotations across modalities are prohibited. Annotations that cannot be directly determined are to be verified and associated with a clearly defined category in other modal data related to the same event. 4. **Complementary Feature Descriptions:** Text reports focus on environmental parameter features, remote sensing images on spatial morphological features, and surveillance videos on dynamic evolution features. Annotations are developed around the core features of the same event and mutually corroborate each other, with cross-modal association fields noted in the annotations. 5. **Association of Missing Information:** When information in a particular modal data is missing, the missing type and reason are clearly indicated, and corresponding valid information in other modal data related to the same event is simultaneously associated to ensure that core feature information is not entirely missing. For example, when an oil spill occurs in a certain monitoring area, the corresponding standardized text report, remote sensing image, and monitoring video are all labeled as oil spill. The spatiotemporal information of the event is the same timestamp and spatial coordinates. The event feature description is developed around the environmental parameters, image features, and dynamic features of the oil spill event.

[0021] After annotation, standardized multi-source marine data is associated and matched with corresponding marine event information. The matching process adopts a hash matching algorithm based on spatiotemporal features. This algorithm uses UTC standard timestamps and WGS84 latitude and longitude coordinates as core inputs, constructs a "timestamp-longitude-latitude" formatted string as the input value of the hash function, and uses the MD5 hash function to convert the input string into a unique spatiotemporal hash value for each group of standardized text reports, remote sensing images, and monitoring videos with the same marine event information. The algorithm performs hash operations on the timestamps and spatial coordinates of events, establishing a one-to-one correspondence between standardized multi-source ocean data and ocean event information through the hash values. Simultaneously, it establishes association mappings between multimodal data, enabling text, image, and video data corresponding to the same ocean event to form associated data groups. After completing the association matching, the data is integrated and encapsulated using a structured data encapsulation format. Each associated standardized text report, remote sensing image, monitoring video, and corresponding ocean event information is encapsulated into a unified dataset file. This file contains a data index module, a data ontology module, and a labeling information module. The data index module records the storage paths and associations of each modality of data, the data ontology module stores standardized multi-source ocean data, and the labeling information module stores the corresponding ocean event information. This encapsulation method forms a complete multimodal ocean dataset. Each dataset unit contains multimodal data and annotation information corresponding to the same ocean event. It solves the technical problems of simple splicing and fragmented spatiotemporal correlation in traditional multimodal integration methods. At the same time, it realizes the systematic integration of multimodal ocean data, making the environmental parameter information of text, the spatial feature information of remote sensing images, and the dynamic evolution information of monitoring videos form an organic whole, which can give full play to the complementary value of multi-source data.

[0022] S2. Perform spatiotemporal alignment processing on the multimodal ocean dataset to obtain the registered multimodal ocean dataset.

[0023] In one specific embodiment, performing step S2 includes the following steps: By using a pre-established spatiotemporal reference coordinate system, timestamp parsing and spatial projection transformation are performed on text reports, remote sensing images, and surveillance videos in the multimodal ocean dataset to obtain preliminarily aligned multimodal data. A cross-modal spatiotemporal alignment model is adopted to extract the time markers and spatial key points corresponding to each modality in the initially aligned multimodal data; Calculate the time deviation between each modal data based on time markers, and calculate the spatial overlap between each modal data based on spatial key points; Based on temporal deviation and spatial overlap, high-precision registration correction is performed on the initially aligned multimodal data to eliminate spatiotemporal matching deviations between different modal data. The registered and corrected multimodal data is associated and bound with the corresponding marine event information in the multimodal ocean dataset to generate a registered multimodal ocean dataset.

[0024] Specifically, spatiotemporal alignment processing of multimodal ocean datasets requires basic processing based on a pre-established spatiotemporal reference coordinate system. This system is built upon the UTC time standard and the WGS84 geographic coordinate system, covering the entire spatiotemporal range of ocean monitoring. For the three types of data in the multimodal ocean datasets—text reports, remote sensing images, and surveillance videos—timestamp parsing and spatial projection transformation operations are performed respectively. For text reports, the event occurrence time information is extracted and converted to a UTC standard timestamp. Spatial information such as latitude, longitude, and sea area location is extracted and converted to spatial coordinates in the WGS84 coordinate system. For remote sensing images, the imaging equipment... The original imaging coordinates of the remote sensing image are converted to planar coordinates in the WGS84 coordinate system using a projection transformation algorithm with a projection error threshold of 0.001 degrees. The timestamps of the monitoring video frame sequence are analyzed and uniformly calibrated to UTC time. The installation latitude and longitude and shooting angle parameters of the camera are extracted. The video shooting area is mapped to the WGS84 coordinate system through spatial coordinate transformation. After the above processing, the three types of data form a preliminary aligned multimodal data with unified time reference and spatial coordinates, realizing a one-to-one correspondence between the original multimodal ocean dataset and the preliminary aligned multimodal data.

[0025] After initial alignment, a cross-modal spatiotemporal alignment model was used to extract time markers and spatial keypoints corresponding to each modality's data. This model is a Siamese network built on the Transformer architecture, containing an encoder layer, a cross-attention matching layer, and a decoder layer. After dedicated training, the model possesses cross-modal spatiotemporal feature matching capabilities. Training was based on the multimodal ocean dataset constructed by S1, with 80% of the data extracted as the training set and the remaining 20% ​​as the validation set. Manually labeled multimodal "spatiotemporal anchor pairs" were used as training labels. Each anchor pair contained trimodal correspondence information, including text latitude and longitude coordinates, corresponding pixel coordinates of the image, and corresponding pixel coordinates of the video keyframe. The training loss function adopted was the contrastive loss function, with the formula: , where d is the Euclidean distance between the anchor point and the feature vector, y=1 is a positive sample pair and y=0 is a negative sample pair, and the margin is set to 1.0; the training optimizer is Adam, the initial learning rate is 0.0001, and it decays to 0.1 times the original value every 20 epochs. The batch size is 16 sets of multimodal data. The training rounds are 100 epochs and the training is stopped early when the validation set loss does not decrease for 10 consecutive epochs. The model encoder layer consists of three identical 6-layer Transformer encoders, processing preliminary alignment data for text, image, and video modalities respectively. Text data is converted into 512-dimensional feature vectors via word embedding layers. Remote sensing images are processed by a convolutional neural network (3 convolutional layers, 2 pooling layers, 3×3 kernel size) to extract 2048-dimensional image feature maps. Surveillance video is processed using optical flow to extract inter-frame motion features and converted into 1024-dimensional video feature vectors. The cross-attention matching layer receives the encoder output and calculates the spatiotemporal similarity between features of different modalities using an 8-head multi-head cross-attention mechanism. The similarity between text and image is calculated using the cosine similarity of the feature vectors. The similarity between images and videos is calculated by weighting spatial feature matching degree and temporal series overlap degree. The decoder layer iteratively optimizes the feature matching relationship based on the similarity calculation results, and finally extracts the time stamps and spatial key points of each modality data. The time stamps of the text report are the UTC timestamps of the event, and the spatial key points are the core geographic coordinates of the WGS84 coordinate system corresponding to the event. The time stamps of the remote sensing images are the imaging UTC timestamps, and the spatial key points are the core geographic coordinates and four corner geographic coordinates of the image imaging area. The time stamps of the surveillance video are the UTC timestamps of the key frames, and the spatial key points are the core geographic coordinates and shooting boundary geographic coordinates of the WGS84 coordinate system corresponding to the video shooting scene.

[0026] Based on the extracted time markers and spatial key points, the temporal deviation and spatial overlap between each modality of data are calculated. The temporal deviation is obtained by calculating the absolute value of the UTC time difference between the time markers of any two types of modality data, in seconds. The corresponding temporal deviations are calculated between text reports and remote sensing images, text reports and surveillance videos, and remote sensing images and surveillance videos. The spatial overlap is calculated using the intersection-union algorithm. For the regions enclosed by the spatial key points of any two types of modality data, the ratio of the intersection area to the union area is calculated, with the ratio ranging from 0 to 1. Similarly, the spatial overlap between each pair of the three types of modality data is calculated. This calculation method enables a quantitative characterization of the spatiotemporal deviation between each modality of data, providing calculable and comparable quantitative indicators for the spatiotemporal differences between different modalities.

[0027] Based on the calculated time deviation and spatial overlap, high-precision registration and correction are performed on the initially aligned multimodal data. The preset time deviation threshold is 30 seconds and the spatial overlap threshold is 0.7. For rapidly evolving marine events (such as ship anomalies and sudden oil spills), the time deviation threshold is tightened to 10 seconds and the spatial overlap threshold is increased to 0.85. For slowly evolving marine events (such as red tide spread and ocean circulation), the time deviation threshold is relaxed to 60 seconds and the spatial overlap threshold is lowered to 0.6. When the time deviation between any two types of modal data exceeds a threshold, time-lagging modal data is corrected by time translation. For surveillance video, time calibration is achieved by adjusting the timestamps of the frame sequence. For remote sensing images, the imaging timestamps are directly corrected. After correction, the time deviation is recalculated until it falls below the threshold. When the spatial overlap is below the threshold, spatial geometric correction is performed on modal data with spatial coordinate deviation. An affine transformation algorithm is used to sequentially perform rotation, translation, and scaling on the spatial coordinates of remote sensing images and surveillance videos. Gradient descent is used to optimize the affine transformation parameters, with a learning rate of 0.00001 degrees / step and a maximum iteration count of 100. Iteration is terminated under two conditions: parameter adjustment is immediately stopped when the spatial overlap exceeds a preset threshold or the maximum iteration count is reached. For the spatial coordinates of the text report, fine-tuning is performed using linear interpolation based on the spatial key points of the corrected remote sensing images and surveillance videos to ensure a high degree of spatial matching among the three types of modal data: text, remote sensing images, and surveillance videos. After all correction operations are completed, the temporal deviation and spatial overlap between each modal data are recalculated to verify whether they all meet the preset threshold requirements, thus forming corrected multimodal data without spatiotemporal matching deviation, achieving accurate conversion from initially aligned multimodal data to corrected multimodal data.

[0028] The registered and corrected multimodal data is associated with corresponding marine event information in the multimodal ocean dataset. A feature association coding algorithm is used to match the time stamps and spatial keypoints of the corrected multimodal data with the spatiotemporal information of the events in the marine event information, generating a unique association code. This code consists of a time feature code, a spatial feature code, and an event identifier code. The time feature code is obtained by converting the corrected UTC timestamp, the spatial feature code is obtained by converting the core coordinates of the WGS84 coordinate system, and the event identifier code is a unique number for the marine event information. This association coding establishes a one-to-one correspondence between the corrected multimodal data and the marine event information, ensuring that each set of corrected data... Textual and remote sensing images and surveillance video data are precisely bound to corresponding marine event types and event feature descriptions. All data that have been linked and bound are then integrated to generate a registered multimodal marine dataset. All multimodal data in this dataset achieves high-precision temporal and spatial alignment and maintains a precise correlation with marine event information. This solves the technical problems of fragmented spatiotemporal correlation of data and inability to achieve precise correspondence between different modalities under the same marine event in traditional multimodal integration methods. It enables deep fusion of environmental parameter information from text, spatial feature information from remote sensing images, and dynamic evolution information from surveillance videos around the same marine event, fully leveraging the complementary information value of multimodal data.

[0029] S3. Extract feature information from the registered multimodal ocean dataset using a pre-set large model to obtain evaluation feature information, which includes environmental parameters, image region features, and dynamic texture sequences.

[0030] In one specific embodiment, performing step S3 includes the following steps: By using a pre-set large model, semantic parsing and keyword extraction are performed on the text reports in the registered multimodal ocean dataset to obtain the corresponding environmental parameters, including wind speed values, flow direction angles, and event occurrence times. By using a pre-set large model, suspected target regions are segmented in remote sensing images of the registered multimodal ocean dataset, and feature vectors of the suspected target regions are extracted to form image region features. By using a pre-set large model, inter-frame difference and motion feature analysis were performed on the monitoring videos in the registered multimodal ocean dataset to extract the dynamic texture sequence of sea surface flow. By using a pre-set large model, environmental parameters, image region features, and dynamic texture sequences are aligned and integrated across modal feature dimensions to generate evaluation feature information.

[0031] Specifically, please refer to Figure 2Feature extraction from the registered multimodal ocean dataset requires a pre-built large model. This model is designed for multimodal data processing in the ocean domain, integrating natural language processing and computer vision. The model is trained using labeled multimodal data from the ocean domain, including text reports, remote sensing images, surveillance videos, and corresponding ocean event information. The training batch size is set to 32, the initial learning rate is set to 0.001, and the model converges after 100 epochs of training. The input to the model is the registered multimodal ocean dataset, and the output is the feature information of each modality and the integrated evaluation feature information. First, the registered text report is input into the natural language processing module of the pre-set large model for semantic parsing and keyword extraction. This module includes a word segmentation layer, a semantic encoding layer, and a keyword extraction layer. The word segmentation layer uses the bidirectional maximum matching method to segment the text report at the word level. The semantic encoding layer uses a pre-trained marine domain BERT model for semantic encoding, with the hidden layer dimension set to 768 and the number of attention heads set to 12. Marine domain keyword matching is performed on the encoded feature vector. The matching dictionary contains words related to environmental parameters such as wind speed, current direction, and time. The similarity between the feature vector and the words in the dictionary is calculated using cosine similarity. For example, the similarity threshold is set to 0.85. Words higher than the threshold are extracted as core keywords. At the same time, the numerical information corresponding to the keywords is parsed and extracted to form environmental parameters including wind speed values, current direction angles, and the time of event occurrence.

[0032] The registered remote sensing image is input into the computer vision module of a pre-defined large model to complete the segmentation of suspected target regions and feature vector extraction. The suspected target region segmentation submodule of this module adopts the U-Net network architecture, which includes 4 downsampling layers and 4 upsampling layers. The convolutional kernel size is set to 3×3, the activation function is ReLU, and the output layer generates a segmentation mask through the Sigmoid function. For example, the mask threshold is set to 0.5, and pixel regions higher than the threshold are identified as suspected target regions. After segmentation, the feature extraction submodule extracts feature vectors from the suspected target regions. This submodule uses a convolutional neural network with a 5-layer structure, specifically including 3 convolutional layers, 1 pooling layer, and 1 fully connected layer. The 3 convolutional layers all use 3×3 kernels with a stride of 1, same padding, and ReLU activation function to progressively extract low-level visual features and high-level semantic features from the suspected target regions. The pooling layer uses 2×2 max-pooling kernels with a stride of 2 to reduce the dimensionality of the feature maps output by the convolutional layers, reducing the number of parameters while retaining key features. The fully connected layer contains 1 hidden layer and 1 output layer. The hidden layer has 1024 neurons with ReLU activation function, and the output layer has 256 neurons to map the extracted multi-dimensional features into 256-dimensional feature vectors, maintaining consistency with subsequent feature dimension alignment requirements. This convolutional neural network extracts low-level visual features and high-level semantic features of the region. Low-level features include gray-level mean, texture variance, and edge density, while high-level features include region shape, area ratio, and pixel distribution. The extracted features are normalized to 256 dimensions to form corresponding feature vectors. Multiple feature vectors are integrated to form image region features.

[0033] The registered surveillance video is input into the computer vision module of the preset large model to perform inter-frame difference and motion feature analysis. First, the surveillance video is preprocessed with frame sequence, and key frames are extracted at a fixed frame rate with a frame interval of 2 frames. The inter-frame difference algorithm is executed on consecutive key frames to calculate the change in pixel value between frames by pixel-by-pixel subtraction. The change threshold is set to 30, and pixels with values ​​higher than the threshold are identified as moving pixels. Connectivity analysis is performed on the moving pixel region to extract the contour and position information of the moving region. Then, the motion feature analysis of the moving region is performed using the optical flow method. The Lucas-Kanade optical flow algorithm is used with a window size of 5×5 to calculate the pixel motion vector of the moving region and obtain motion features such as motion direction and motion speed. Based on the time series, the motion features of consecutive frames are integrated to generate a dynamic texture sequence representing the state of sea surface flow in the order of the frame sequence. The sequence dimension is consistent with the number of key frames, and each sequence node contains feature information such as motion direction, speed, and texture distribution.

[0034] The evaluation feature information is generated by aligning and integrating the environmental parameters, image region features, and dynamic texture sequences across modal feature dimensions using a pre-defined large model. Specifically, the environmental parameters are first vectorized into 256-dimensional vectors consistent with the image region features. The dynamic texture sequences are compressed to 256 dimensions using principal component analysis with a principal component contribution rate of 0.95, thus unifying the three feature dimensions. Then, corresponding fusion weights are assigned according to the correlation. First, the dot product similarity between the environmental parameters, image region features, and dynamic texture sequences and the marine event features is calculated, denoted as the correlation. (Environmental parameters) (Image region features) (Dynamic texture sequence), and then the correlation is evaluated using the softmax function. , , Normalization is performed to obtain the dynamic fusion weights corresponding to the three types of features. The dynamic weight is the core basis for feature fusion. At the same time, the preset correlation threshold is 0.2. If the correlation of any single-modal feature is lower than the threshold, in order to avoid the failure of single-modal features affecting the fusion effect, the fallback mechanism is immediately activated, and the fusion weight is switched to the preset initial weight (environmental parameter 0.3, image region feature 0.4, dynamic texture sequence 0.3). Finally, each type of feature vector is multiplied by the corresponding fusion weight (dynamic weight / initial fallback weight) and then added element by element to obtain the fused feature information. This fused feature information is the evaluation feature information.

[0035] S4. Consistency judgment is made on the evaluation feature information by using the physical law constraints built into the preset large model to obtain the physical compliance score.

[0036] In one specific embodiment, performing step S4 includes the following steps: The image region features, environmental parameters, and dynamic texture sequences in the evaluation feature information are input into the physical law constraint module built into the preset large model; Based on environmental parameters, the Stokes drift model is used to simulate the expected driving direction of the sea surface under wind-driven conditions, and the corresponding flow vector and velocity threshold range are obtained. Based on the dynamic texture sequence, the displacement vector sequence of the target area on the sea surface is extracted, and the actual movement direction and actual movement speed of the target area are obtained by fitting. Calculate the degree of matching between the actual movement direction of the target area and the expected driving flow direction, as well as the degree of conformity between the actual movement speed and the speed threshold range; Based on the calculation results of matching degree and conformity degree, a corresponding physical conformity score is generated. The physical conformity score is used to characterize the degree of conformity between feature information and physical laws.

[0037] Specifically, image region features, environmental parameters, and dynamic texture sequences extracted from registered multimodal ocean data are used as input data and fed into a self-contained, callable physical constraint calculation unit built into the pre-defined large model. This unit is a lightweight embedded functional module built into the pre-defined large model to verify the laws of ocean hydrodynamics. It is not an independent neural network. It achieves end-to-end data interaction with the feature extraction module of the large model through an internal tensor channel. It directly receives the three types of evaluation feature information output by the feature extraction module and performs consistency verification calculations without external data conversion. This data input method directly connects to the output results of the preceding feature extraction stage, realizing the continuous connection of the data flow in the evaluation process. Moreover, all the calculation logic and parameters of this unit are solidified in the inference framework of the pre-defined large model, which is a native built-in function of the large model and is loaded, called, and inferred synchronously with the large model. Based on the input environmental parameters, the physical constraint module calls the Stokes drift model to perform simulation calculations of the expected driving direction of the sea surface under wind drive. The calculation of the Stokes drift model is based on the wind stress formula. To solve for wind stress, air density, Here, V represents the surface drag coefficient, and V is the actual wind speed value obtained from the environmental parameters. The model incorporates the flow direction angle from the environmental parameters as a basic reference, simulating the expected driving flow vector of the sea surface under this wind field condition through fluid dynamics equations. Simultaneously, it calculates the corresponding sea surface flow velocity threshold range based on a pre-built wind stress-current response lookup table. This method is an engineering-oriented quantitative calculation approach, eliminating the need for real-time complex fluid dynamics solutions: First, for the target monitoring sea area, the high-precision regional ocean model ROMS (Regional Ocean Modeling System, a universally used high-precision numerical simulation tool in the marine field for simulating ocean currents, wind fields, and other marine environments) is pre-run. A full combination of working conditions is input, including wind speeds of 0-30 m / s (step size 1 m / s, covering common ocean wind speed ranges) and wind directions of 0-360° (step size 10°, covering all wind directions, with a 10° step size balancing accuracy and efficiency). The ocean current response at a 10-meter sea surface level is simulated, and the average sea surface velocity output under each working condition is recorded. and standard deviation Using (wind speed, wind direction) as the index key A two-dimensional lookup table is constructed for the values, and the lookup table is updated annually to adapt to seasonal changes. During real-time evaluation, bilinear interpolation is performed on the lookup table based on the input wind speed value and flow direction angle to obtain the interpolated mean value. and standard deviation Define the speed threshold range as The 95% confidence interval for the flow velocity serves as a quantitative reference standard for velocity compliance calculation. A one-to-one data relationship is formed between the wind speed values ​​and flow direction angles in the environmental parameters and the flow direction vector and velocity threshold range output by the model. Each set of environmental parameters can yield a unique expected driving flow direction and velocity threshold reference through the model. For example, inputting a wind speed of 8 m / s and a wind direction of 135°, the expected driving flow direction and velocity threshold reference can be obtained through lookup table interpolation. , The velocity threshold range is [0.194m / s, 0.506m / s]. This process provides a quantitative reference standard for the physical verification of ocean data, solving the technical problem that traditional methods rely solely on surface feature matching and lack quantitative basis based on physical laws.

[0038] Based on the input dynamic texture sequence, the physical law constraint calculation unit extracts the displacement vector sequence of the target area on the sea surface using the Lucas-Kanade optical flow algorithm. This algorithm sets a 5×5 calculation window and solves the motion vector of pixels in consecutive frames of the dynamic texture sequence. The motion displacement of each pixel is vector synthesized to obtain the displacement vector sequence of the target area on the sea surface in the time series. Then, the displacement vector sequence is linearly fitted by the least squares method. The fitting process uses time as the horizontal axis and displacement as the vertical axis to solve for the actual movement direction of the target area in space. At the same time, the actual movement speed is calculated based on the ratio of displacement to time. Each frame sequence node data in the dynamic texture sequence corresponds to a displacement vector in the displacement vector sequence. After fitting, a unique actual movement direction and actual movement speed data are formed. This data processing process realizes the quantitative extraction of dynamic features of the video, allowing the dynamic changes of the sea surface to be transformed into computable physical parameters, solving the technical problem that it is difficult to effectively utilize video dynamic information in traditional multimodal processing.

[0039] After obtaining the expected driving flow vector, velocity threshold range, and the actual movement direction and velocity of the target area, the physical law constraint calculation unit performs matching and conformity calculations on the two types of data. The cosine similarity algorithm is used to calculate the matching degree between the actual movement direction and the expected driving flow vector. The algorithm converts the two direction vectors into unit vectors and then performs a dot product operation. The result ranges from 0 to 1, with higher values ​​indicating higher direction matching. For the conformity of the actual movement velocity with the velocity threshold range, a normalization calculation method is used. If the actual movement velocity is within the velocity threshold range, the conformity is directly set to 1. If it exceeds the threshold range, a linear decay calculation is performed based on the magnitude of the exceedance; the greater the exceedance, the lower the conformity value. The actual movement direction and the expected driving flow vector have a one-to-one correspondence, while the actual movement velocity and the velocity threshold range have a matching relationship between a single data point and a numerical range. These two types of calculation results quantify the degree of conformity between the target area's motion in direction and velocity and the physical laws, solving the technical problem of the lack of quantitative calculation standards for physical conformity in traditional evaluations.

[0040] Based on the calculated matching and compliance results, the physical law constraint calculation unit generates a physical compliance score using a weighted fusion method. For example, the weight coefficient for direction matching is set to 0.6 and the weight coefficient for velocity compliance is set to 0.4. The physical compliance score is calculated using the formula S=0.6×M+0.4×F, where S is the physical compliance score, M is the direction matching, and F is the velocity compliance. The score ranges from 0 to 1. Each set of matching and compliance values ​​can yield a unique physical compliance score using this formula. This score directly characterizes the degree of compliance between the evaluation feature information and the physical laws of marine fluid dynamics. For example, in the evaluation of an oil spill, if the cosine similarity between the actual movement direction of the oil spill area and the expected wind-driven flow direction is 0.95, and the actual movement speed is within the speed threshold range with a compliance of 1, then the physical compliance score calculated using the formula is 0.97. This quantitative scoring result is fed back to the subsequent category determination module through the internal channel of the large model, providing a clear physical basis for the subsequent determination of marine event categories. This effectively avoids misjudging natural marine phenomena such as wind-blown ripples and natural oil slicks as target marine events, and solves the technical problems of easy misjudgment and low accuracy of evaluation results in traditional evaluation methods. At the same time, this score is incorporated into the comprehensive evaluation and reasoning process of the large model as a quantitative indicator, so that the entire evaluation process has an objective judgment standard, and solves the technical problems of strong subjectivity and lack of quantitative verification mechanism in traditional evaluation.

[0041] S5. Determine the preliminary event category based on the physical compliance score and the evaluation feature information.

[0042] In one specific embodiment, performing step S5 includes the following steps: Physical compliance scoring thresholds are set based on historical data from marine event statistics. Determine whether the physical compliance score is lower than the physical compliance score threshold; If so, then by combining the environmental parameters and dynamic texture sequence in the evaluation feature information, the corresponding marine phenomenon is determined to be a natural phenomenon representation and marked as a non-target event category; If not, then by combining the image region features in the evaluation feature information, the corresponding marine phenomenon is determined to be a potential target event and marked as a target event category to be verified; The marked non-target event categories are integrated with the target event categories to be verified to generate preliminary event categories.

[0043] Specifically, relying on historical data of marine events accumulated in the marine field, statistical analysis is conducted on the physical conformity scores corresponding to different types of marine phenomena. The threshold for physical conformity scores is determined using the quantile method. The specific process is carried out step by step based on the long-term accumulated historical data of marine events in the marine field: First, the historical data is preprocessed to remove abnormal score data caused by monitoring equipment failure or data transmission interference, and invalid data with missing marine phenomenon category labels are screened out. Valid and fully labeled score data sets are retained and divided into natural phenomenon score datasets and target event score datasets according to actual categories. Then, the natural phenomenon score datasets are sorted in ascending order of score values ​​to form an ordered numerical sequence. The quantile position is determined using the quantile position calculation formula K=P×(n+1) associated with linear interpolation in statistics, where K is the continuous value of the quantile in the ordered sequence. The location value, P, is the upper quantile ratio selected based on the requirements of marine monitoring misjudgment rate control, and n is the number of valid data in the natural phenomenon scoring dataset. Then, the specific score value corresponding to the K location is calculated using linear interpolation to obtain the preliminary threshold reference value. Subsequently, it is verified and adjusted in conjunction with the target event scoring dataset, and the lower quantile value of the target event scoring dataset is extracted. If the preliminary threshold reference value is lower than the lower quantile value, it is directly determined as the final threshold. If there is overlap, it is fine-tuned based on the score overlap interval to ensure that the majority of natural phenomenon scores are distributed below the threshold and the majority of target event scores are distributed above the threshold. At the same time, for the differences in sea conditions in different sea areas, historical datasets for corresponding sea areas are constructed, and personalized thresholds are calculated using the same process to form a threshold system adapted to different sea areas. Furthermore, the thresholds are dynamically updated regularly based on newly added monitoring data to ensure the timeliness and adaptability of the thresholds.

[0044] The physical compliance score obtained in the previous step is compared with the set physical compliance score threshold to complete the logical judgment of whether the score is lower than the threshold. The physical compliance score serves as a single numerical judgment basis and forms a one-to-one comparison relationship with the threshold. Each physical compliance score can obtain a clear judgment result through this comparison. If the physical conformity score is determined to be below the threshold, the environmental parameters and dynamic texture sequences extracted from the previous evaluation feature information are retrieved. These two types of feature information are used as the basis for judgment, and the corresponding marine phenomena are comprehensively analyzed in conjunction with the laws of marine hydrodynamics. The wind and current-related parameters in the environmental parameters and the dynamic changes in the sea surface in the dynamic texture sequence form complementary data, jointly representing the natural evolution law of the marine phenomenon. Based on the feature correlation between the two types of data, the marine phenomenon is determined to be a natural phenomenon and is marked as a non-target event category. Each physical conformity score below the threshold corresponds to a set of environmental parameters and dynamic texture sequences. After analysis, a unique non-target event category label is obtained. For example, the physical conformity score corresponding to wind-blown ripples on the sea surface is below the threshold. Combined with the wind field characteristics in the corresponding environmental parameters and the dynamic characteristics of no continuous diffusion in the dynamic texture sequence, it is determined to be a natural phenomenon and marked as a non-target event category. If the physical conformity score is determined to be no lower than the threshold, the image region features in the evaluation feature information are retrieved. This feature information is used as the core judgment basis to analyze the spatial morphology, pixel distribution, and other features of the suspected target region in the image. The suspected target region features in the image region features are correlated with the physical conformity scores that are no lower than the threshold, indicating that the marine phenomenon has the physical law characteristics and spatial morphological characteristics of the target event. Based on this feature information, the corresponding marine phenomenon is determined to be a potential target event and is marked as a target event category to be verified. Each group of physical conformity scores that are no lower than the threshold corresponds to a unique image region feature. After analysis, a unique target event category label to be verified is obtained. For example, if the physical conformity score of an oil spill on the sea surface is not lower than the threshold, and the irregular planar suspected target area features exist in the corresponding image area features, it is determined to be a potential target event and marked as a target event category to be verified. This determination process achieves accurate screening of potential target events, allowing subsequent multimodal information matching to be carried out only for target events to be verified, improving the efficiency of the overall evaluation process, and solving the problem of resource waste caused by the lack of screening in traditional evaluation and the complex matching of all marine phenomena.

[0045] The non-target event categories that have been marked and the target event categories to be verified are classified and integrated. All marine phenomena that have been judged and marked are grouped by category. Each judged marine phenomenon corresponds to a unique category label. After integration, a preliminary event category containing category information of all evaluation objects is formed. This preliminary event category forms a complete data correspondence with the preceding physical conformity score and evaluation feature information. Each set of scores and feature information can find the corresponding category in the preliminary event category. This integration process realizes the systematic classification of marine phenomena, defines a clear matching range for subsequent multimodal information matching, and solves the technical problems of traditional multimodal matching having no clear range, low matching efficiency, and easy invalid matching.

[0046] S6. Based on the preliminary event category, the evaluation feature information and marine event information are matched and compared using multimodal information to obtain the event matching probability value.

[0047] In one specific embodiment, performing step S6 includes the following steps: Identify the evaluation feature information and marine event information corresponding to the target event category to be verified in the preliminary event category; Extract the temporal morphological evolution features of suspected target regions from the locked evaluation feature information image region features. The temporal morphological evolution features include the area change rate of the target region, centroid displacement sequence, boundary fractal dimension change amount and aspect ratio fluctuation value. Retrieve the standard morphological evolution template features corresponding to various types of marine events pre-stored in the locked marine event information; The similarity between the temporal morphological evolution features and the standard morphological evolution template features is calculated one by one to obtain the feature matching degree corresponding to each type of marine event. By combining the physical conformity score with the matching degree of each feature, a weighted fusion calculation is performed to generate the event matching probability value corresponding to the target event category to be verified.

[0048] Specifically, based on the initially defined scope of event categories, the corresponding evaluation feature information and marine event information are accurately locked for the target event categories to be verified. During this process, each group of target event categories to be verified uniquely corresponds to a set of evaluation feature information and a set of compliant marine event information, forming a one-to-one mapping relationship. The locking operation is based on the spatiotemporal hash values ​​previously allocated to the multimodal marine data. By matching hash values, all relevant data corresponding to the target event to be verified are directly located, avoiding invalid processing of non-target event data. This solves the technical problems of traditional evaluation methods, such as the lack of a clear matching range, low data processing efficiency, and susceptibility to invalid matches.

[0049] After data locking is completed, temporal morphological evolution features are extracted from the image region features contained in the locked evaluation feature information for suspected target regions. The extraction process is based on the registered remote sensing image sequence and the key frame sequence of the monitoring video. The area change rate of suspected target regions in the continuous time dimension is calculated sequentially. This value is obtained by the ratio of the area difference of suspected target regions at adjacent time nodes to the area of ​​the previous node. At the same time, the spatial coordinate change of the centroid of the suspected target region is calculated through the WGS84 coordinate system to form the centroid displacement sequence. The box-counting method is used to calculate the change in the fractal dimension of the boundary of the suspected target region. When calculating the box-counting method, the difference in fractal dimension is obtained by covering the boundary region with grids of different scales according to the set grid step size. Then, the aspect ratio fluctuation value of the minimum bounding rectangle of the suspected target region at adjacent time nodes is statistically analyzed. These feature indicators together constitute the temporal morphological evolution features. The image region features of each suspected target region will correspond to a set of temporal morphological evolution features including the area change rate, centroid displacement sequence, boundary fractal dimension change value, and aspect ratio fluctuation value.

[0050] After extracting the temporal morphological evolution features, the standard morphological evolution template features corresponding to various types of marine events are retrieved from the locked marine event information. Each type of marine event in the pre-stored marine event information corresponds to a set of standard morphological evolution template features obtained through statistical analysis of historical data in the marine field. These template features also include the standard reference range and variation law of area change rate, centroid displacement sequence, boundary fractal dimension change, and aspect ratio fluctuation value. The retrieval process is completed through event type identification code. The marine event information corresponding to each target event to be verified will be matched with the corresponding event type identification code according to its spatiotemporal characteristics. The corresponding standard morphological evolution template features are directly retrieved through this identification code, realizing the dimensional unity and type matching between the extracted temporal morphological evolution features and the standard template features.

[0051] The extracted temporal morphological evolution features are compared one by one with the retrieved standard morphological evolution template features for various marine event types. The calculation process employs a cosine similarity algorithm, converting both temporal morphological evolution features and standard morphological evolution template features into fixed-dimensional feature vectors. During calculation, the feature vectors are processed according to a set normalization threshold, and the similarity value is obtained through vector dot product. Each set of temporal morphological evolution features is compared with the standard morphological evolution template features for each marine event type, resulting in a corresponding feature matching degree. Multiple feature matching degrees are obtained for multiple marine event types. This calculation process allows for the quantitative determination of feature similarity, solving the technical problems of lacking a unified quantitative standard and high subjectivity in traditional evaluation of feature matching.

[0052] After obtaining the feature matching degree corresponding to each type of marine event, the physical conformity score obtained in the previous steps is combined with these feature matching degrees for weighted fusion calculation to generate the event matching probability value corresponding to the target event category to be verified. The weighted fusion calculation integrates the physical conformity score and feature matching degree according to the set weight allocation ratio. During the calculation, the physical conformity score and each feature matching degree are substituted into the calculation formula respectively. Each feature matching degree will generate an event matching probability value, corresponding to the matching possibility of different marine event types. The weight coefficient is set according to the evaluation needs of the marine field, so that the physical conformity and feature morphological similarity form a reasonable weight allocation in event judgment. This fusion calculation process combines physical basis with feature matching results, which solves the technical problem of low accuracy of evaluation results in traditional evaluation that relies only on single feature similarity judgment. For example, in the process of verifying an oil spill, the temporal morphological evolution features of the suspected target area extracted show a dynamic evolution law that conforms to the oil spill diffusion. After retrieving the standard morphological evolution template features of the oil spill, the corresponding feature matching degree is calculated by the cosine similarity algorithm. Combined with the physical conformity score of the target event to be verified, and substituted into the weighted fusion formula, the event matching probability value corresponding to the oil spill event is obtained, which clearly characterizes the possibility that the target event to be verified is an oil spill event.

[0053] S7. Through comprehensive analysis and verification using physical conformity scores and event matching probability values, a comprehensive evaluation result of the multimodal ocean dataset is generated.

[0054] In one specific embodiment, performing step S7 includes the following steps: Obtain the pre-set weighted fusion coefficient and joint verification threshold. The weighted fusion coefficient includes a first coefficient corresponding to the physical compliance score and a second coefficient corresponding to the event matching probability value. The joint verification score is obtained by multiplying the physical compliance score by the first coefficient and the event matching probability value by the second coefficient. Determine whether the joint verification score reaches the joint verification threshold. If it does, confirm that the corresponding marine phenomenon is the target event and match the corresponding marine event category. Otherwise, determine that the corresponding marine phenomenon is a non-target event category and form the final event category determination result. The joint validation scores are normalized to obtain the confidence level corresponding to the final event category determination result; By integrating the final event category determination results, corresponding confidence scores, evaluation feature information, and marine event information, a comprehensive evaluation result of the multimodal marine dataset is generated.

[0055] Specifically, weighted fusion coefficients and joint verification thresholds are obtained from a pre-set parameter configuration library. This parameter configuration library is built based on historical evaluation data in the field of marine monitoring and sea state characteristics of different sea areas. The weighted fusion coefficients include a first coefficient corresponding to the physical conformity score and a second coefficient corresponding to the event matching probability value. The values ​​of the two types of coefficients are set differently according to the characteristics of marine event types, and the sum of the values ​​of all coefficients is a fixed value. The joint verification threshold is determined by quantile analysis of historical joint verification data of marine target events and non-target events. Different categories of target events to be verified will retrieve a set of exclusive weighted fusion coefficients and joint verification thresholds, forming a unique correspondence. The parameter acquisition process relies on event type identification codes to complete the matching and retrieval, which solves the technical problems of lack of standardized quantitative verification parameters and strong subjectivity in the traditional evaluation process.

[0056] After obtaining the relevant parameters, the physical compliance score obtained in the previous steps is multiplied by the first coefficient, and the corresponding event matching probability value is multiplied by the second coefficient. The results of the two multiplications are then summed to obtain the joint verification score. During the calculation, each physical compliance score corresponds to an event matching probability value. The two are weighted and summed to generate a unique joint verification score. This calculation process deeply integrates the physical compliance and multimodal feature matching results through weighting, so that the two types of judgment criteria reflect the weight ratio that adapts to the evaluation needs of the marine field in the comprehensive verification. This solves the technical problem of relying on a single feature for judgment in traditional evaluation and the lack of multi-dimensional basis for the evaluation results.

[0057] After the joint verification score is calculated, it is compared with the retrieved joint verification threshold. Based on the comparison result, the event category of the marine phenomenon is determined. If the value of the joint verification score reaches or exceeds the joint verification threshold, the corresponding marine phenomenon is confirmed as the target event, and the marine event category corresponding to the matching probability value of the event is matched. If the joint verification score does not reach the joint verification threshold, the corresponding marine phenomenon is determined to be a non-target event. Each joint verification score will obtain a clear event category determination result after being compared with the threshold. This determination process realizes the standardized determination of event categories through quantitative threshold comparison, which solves the technical problems of no clear quantitative standard for event determination and easy misjudgment and omission in traditional evaluation.

[0058] After the final event category determination result is formed, the joint validation score is normalized to obtain the confidence level corresponding to the determination result. The normalization process adopts the min-max normalization algorithm, which maps the value range of the joint validation score to a fixed interval. During the calculation process, the upper and lower limits of the numerical mapping of the algorithm are set. The algorithm performs uniform standardization processing on joint validation scores of different numerical ranges. Each joint validation score generates a unique confidence level value after normalization, which directly represents the reliability of the final event category determination result.

[0059] After the confidence level calculation is completed, the final event category determination result, the corresponding confidence level, the previously extracted evaluation feature information, and the original marine event information are comprehensively integrated. The integration process employs a structured data encapsulation method, dividing various data types into corresponding modules according to data relationships. Specifically, the final event category determination result and the confidence level form a related data group; environmental parameters, image region features, and dynamic texture sequences in the evaluation feature information are encapsulated by modality; and marine event information is classified by event type, spatiotemporal information, and feature description. Each group of relevant data for the target event to be verified, after integration, forms a unique multimodal marine dataset comprehensive evaluation result. This integration process achieves a comprehensive correlation between various data types during the evaluation process and the final determination result, solving the technical problem of traditional evaluation results having a single data dimension and being unable to achieve source tracing analysis.

[0060] Please see Figure 3 , Figure 3 The figure shows the ROC curve comparison results, comparing the ROC curves of our method and traditional methods in marine event recognition tasks. It clearly presents the difference in AUC between the two curves, with our method's ROC curve closer to the upper left corner and having a higher AUC value. This figure illustrates the advantages of our evaluation method in marine event recognition accuracy: First, the physical conformity scoring effectively filters natural phenomena, reducing the false positive rate and solving the problem of traditional methods easily misclassifying wind ripples and natural oil films as target events, allowing the ROC curve to maintain a high true positive rate in the low false positive rate range. Second, the construction of a multimodal information matching system improves the accuracy of target event recognition, solving the problem of insufficient true positive rate caused by traditional single feature matching and expanding the AUC area of ​​the ROC curve. Third, the standardized quantitative verification mechanism makes the evaluation results comparable and reliable, solving the problems of strong subjectivity and unstable performance of traditional evaluation methods, and providing intuitive and scientific performance evidence for the application effect of large models in the marine field.

[0061] Please see Figure 4 , Figure 4The figure shows the performance stability comparison results of 20 independent experiments, comparing the proposed method with the traditional method. Using event recognition accuracy as the core indicator, it presents the dispersion and mean differences between the two sets of experimental results. The proposed method exhibits less dispersion and a higher mean. This figure illustrates the performance stability advantages of the proposed method: First, the standardized multimodal dataset construction process ensures the consistency and high quality of input data, solving the performance fluctuation problem caused by inconsistent data quality in traditional methods. Second, the high-precision spatiotemporal alignment and feature extraction process reduces error accumulation during data processing, solving the performance instability problems caused by spatiotemporal bias and non-standard feature extraction in traditional methods. Third, the quantitative weighted fusion and joint verification mechanism ensures that event judgment results are not affected by random factors, solving the problem of large differences in repeated experimental results caused by strong subjectivity and lack of unified standards in traditional methods, providing stability assurance for the practical application of the method.

[0062] Please see Figure 5 The following describes a multimodal ocean data evaluation method system for large models, as described in the embodiments of this application. The multimodal ocean data evaluation method system for large models includes: The data acquisition module is used to construct a multimodal ocean dataset, which includes text reports, remote sensing images, surveillance videos, and corresponding ocean event information. The spatiotemporal alignment module is used to perform spatiotemporal alignment processing on multimodal ocean datasets to obtain registered multimodal ocean datasets; The feature extraction module is used to extract feature information from the registered multimodal ocean dataset using a pre-set large model to obtain evaluation feature information, which includes environmental parameters, image region features, and dynamic texture sequences. The physics judgment module is used to make consistency judgments on the evaluation feature information based on the physical law constraints built into the preset large model, and obtain a physical compliance score. The category determination module is used to determine the preliminary event category based on the physical compliance score and evaluation feature information; The information matching module is used to perform multimodal information matching and comparison between the evaluation feature information and the marine event information based on the preliminary event category, and to obtain the event matching probability value. The comprehensive analysis module is used to perform comprehensive analysis and verification through physical conformity scores and event matching probability values, and generate comprehensive evaluation results for multimodal ocean datasets.

[0063] Through the collaborative efforts of the aforementioned components, the technical problems in traditional marine data evaluation—such as the spatiotemporal fragmentation of multimodal data, the lack of unified standards for feature extraction, the lack of physical basis for event determination, and the high subjectivity and insufficient stability of evaluation results—have been resolved. The data acquisition module and the spatiotemporal alignment module work together to standardize the acquisition of multi-source data and unify spatiotemporal benchmarks. By synchronously acquiring and registering three types of data—text reports, remote sensing images, and surveillance videos—the temporal deviation and spatial misalignment between multimodal data are eliminated. This provides a high-quality, spatiotemporally consistent data foundation for subsequent feature extraction and event determination, solving the problems of scattered storage of multi-source data, fragmented spatiotemporal correlation, and inability to fully leverage the complementary value of information in traditional evaluation. The feature extraction module achieves accurate extraction and dimensional unification of multimodal features. Differentiated algorithms are used to extract environmental parameters, image region features, and dynamic texture sequences based on the characteristics of different modal data. Then, the Transformer cross-modal attention mechanism and PCA dimensional alignment technology are used to achieve weighted fusion of the three types of features, forming unified dimensional evaluation feature information. This solves the problems of chaotic feature extraction standards, heterogeneous multimodal feature dimensions, and low correlation between features in traditional evaluation, fully exploring the complementary value of multi-source data. The physical judgment module and the category determination module work together to complete the initial event classification under the constraints of physical laws. Based on the Stokes drift model and fluid dynamics consistency verification, a physical conformity score is generated. This score is used to initially screen natural phenomena and potential target events, defining the scope for subsequent precise matching. This solves the problems of traditional evaluation methods that rely solely on surface feature matching, lack physical basis, and are prone to misclassifying natural phenomena as target events, thus improving the accuracy and scientific rigor of event classification. The information matching module and the comprehensive analysis module work together to accurately compare and comprehensively verify multimodal information. Based on the initial event category, the target to be verified is identified. The event matching probability value is calculated by comparing the temporal morphological evolution features with the standard template. Then, combined with the physical conformity score, weighted fusion and threshold verification are performed to generate a comprehensive evaluation result with confidence. This solves the problems of traditional evaluation methods that lack quantitative standards for event judgment, are highly subjective, and cannot quantify the reliability of results, achieving standardization of the evaluation process and traceability of the results.

[0064] This application also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the multimodal ocean data evaluation method for large models.

[0065] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the methods and systems described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0066] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for evaluating multimodal ocean data for large models, characterized in that, Includes the following steps: S1. Construct a multimodal ocean dataset, which includes text reports, remote sensing images, surveillance videos, and corresponding ocean event information; S2. Perform spatiotemporal alignment processing on the multimodal ocean dataset to obtain the registered multimodal ocean dataset; S3. Extract feature information from the registered multimodal ocean dataset using a preset large model to obtain evaluation feature information, which includes environmental parameters, image region features, and dynamic texture sequences. S4. The physical law constraints built into the preset large model are used to make a consistency judgment on the evaluation feature information to obtain a physical conformity score. S5. Determine the preliminary event category based on the physical compliance score and the evaluation feature information; S6. Using the preliminary event category as the matching basis, perform multimodal information matching and comparison between the evaluation feature information and the marine event information to obtain the event matching probability value; S7. A comprehensive evaluation result of the multimodal ocean dataset is generated by comprehensively analyzing and verifying the physical conformity score and the event matching probability value.

2. The method according to claim 1, characterized in that, Step S1 includes: Collect multi-source raw data in the marine field, including text reports, remote sensing images, and surveillance videos; The multi-source raw data is cleaned to remove invalid and noisy data, resulting in standardized multi-source ocean data. For each type of data in the standardized multi-source ocean data, corresponding ocean event information is labeled, including ocean event type, spatiotemporal information of event occurrence, and event feature description; The standardized multi-source ocean data is associated and matched with the corresponding ocean event information, and then integrated and encapsulated to generate a multimodal ocean dataset.

3. The method according to claim 1, characterized in that, Step S2 includes: Using a pre-established spatiotemporal reference coordinate system, timestamp parsing and spatial projection transformation are performed on text reports, remote sensing images, and surveillance videos in the multimodal ocean dataset to obtain preliminarily aligned multimodal data. A cross-modal spatiotemporal alignment model is used to extract the time markers and spatial key points corresponding to each modality in the initially aligned multimodal data; The time deviation between each modal data is calculated based on the time marker points, and the spatial overlap between each modal data is calculated based on the spatial key points. Based on the time deviation and the spatial overlap, the initially aligned multimodal data is registered and corrected with high precision to eliminate the spatiotemporal matching deviation between the various modal data. The registered and corrected multimodal data is associated and bound with the corresponding marine event information in the multimodal ocean dataset to generate the registered multimodal ocean dataset.

4. The method according to claim 1, characterized in that, Step S3 includes: The registered multimodal ocean dataset is semantically parsed and keywords are extracted by a pre-set large model to obtain the corresponding environmental parameters, including wind speed, current angle and event time. The pre-set large model is used to segment the suspected target regions in the registered multimodal ocean dataset remote sensing images, and the feature vectors of the suspected target regions are extracted to form image region features; The pre-defined large model is used to perform inter-frame difference and motion feature analysis on the monitoring videos in the registered multimodal ocean dataset to extract the dynamic texture sequence of sea surface flow. The evaluation feature information is generated by aligning and integrating the environmental parameters, image region features, and dynamic texture sequences across modal feature dimensions using the preset large model.

5. The method according to claim 1, characterized in that, Step S4 includes: The image region features, environmental parameters, and dynamic texture sequences in the evaluation feature information are input into the physical law constraint module built into the preset large model; Based on the environmental parameters, the Stokes drift model is used to simulate the expected driving direction of the sea surface under wind-driven conditions, and the corresponding flow vector and velocity threshold range are obtained. Based on the dynamic texture sequence, the displacement vector sequence of the target area on the sea surface is extracted, and the actual movement direction and actual movement speed of the target area are obtained by fitting. Calculate the degree of matching between the actual movement direction of the target area and the expected driving flow direction, as well as the degree of conformity between the actual movement speed and the speed threshold range; Based on the calculation results of matching degree and conformity degree, a corresponding physical conformity score is generated, which is used to characterize the degree of conformity between the feature information and physical laws.

6. The method according to claim 1, characterized in that, Step S5 includes: Physical compliance scoring thresholds are set based on historical data from marine event statistics. Determine whether the physical compliance score is lower than the physical compliance score threshold; If so, then by combining the environmental parameters and dynamic texture sequence in the evaluation feature information, the corresponding marine phenomenon is determined to be a natural phenomenon representation and marked as a non-target event category; If not, then by combining the image region features in the evaluation feature information, the corresponding marine phenomenon is determined to be a potential target event and marked as a target event category to be verified; The marked non-target event categories are integrated with the target event categories to be verified to generate preliminary event categories.

7. The method according to claim 6, characterized in that, Step S6 includes: Lock the evaluation feature information and marine event information corresponding to the target event category to be verified in the preliminary event category; Extract the temporal morphological evolution features of the suspected target region from the image region features in the locked evaluation feature information. The temporal morphological evolution features include the area change rate, centroid displacement sequence, boundary fractal dimension change, and aspect ratio fluctuation value of the target region. Retrieve the standard morphological evolution template features corresponding to various types of marine events pre-stored in the locked marine event information; The similarity between the temporal morphological evolution features and the standard morphological evolution template features is calculated one by one to obtain the feature matching degree corresponding to each type of marine event. The physical conformity score and the matching degree of each feature are combined and weighted to generate the event matching probability value corresponding to the target event category to be verified.

8. The method according to claim 1, characterized in that, Step S7 includes: Obtain a pre-set weighted fusion coefficient and a joint verification threshold, wherein the weighted fusion coefficient includes a first coefficient corresponding to the physical compliance score and a second coefficient corresponding to the event matching probability value; The joint verification score is obtained by multiplying the physical conformity score by the first coefficient and the event matching probability value by the second coefficient. Determine whether the joint verification score reaches the joint verification threshold. If it does, confirm that the corresponding marine phenomenon is the target event and match the corresponding marine event category. Otherwise, determine that the corresponding marine phenomenon is a non-target event category and form the final event category determination result. The joint verification score is normalized to obtain the confidence level corresponding to the final event category determination result; By integrating the final event category determination results, corresponding confidence levels, evaluation feature information, and marine event information, a comprehensive evaluation result of the multimodal marine dataset is generated.

9. A multimodal ocean data evaluation system for large models, used to implement the multimodal ocean data evaluation method for large models as described in any one of claims 1 to 8, characterized in that, The aforementioned multimodal ocean data evaluation system for large models includes: The data acquisition module is used to construct a multimodal ocean dataset, which includes text reports, remote sensing images, surveillance videos, and corresponding ocean event information. The spatiotemporal alignment module is used to perform spatiotemporal alignment processing on the multimodal ocean dataset to obtain a registered multimodal ocean dataset. The feature extraction module is used to extract feature information from the registered multimodal ocean dataset using a preset large model to obtain evaluation feature information, which includes environmental parameters, image region features, and dynamic texture sequences. The physical judgment module is used to make a consistency judgment on the evaluation feature information based on the physical law constraints built into the preset large model, and obtain a physical conformity score. The category determination module is used to determine the preliminary event category based on the physical conformity score and the evaluation feature information; The information matching module is used to perform multimodal information matching and comparison between the evaluation feature information and the marine event information based on the preliminary event category to obtain the event matching probability value; The comprehensive analysis module is used to perform comprehensive analysis and verification using the physical conformity score and the event matching probability value, and generate a comprehensive evaluation result for the multimodal ocean dataset.

10. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement the multimodal ocean data evaluation method for large models as described in any one of claims 1-8.