A meteorological disaster identification method and device based on a multi-modal large model

CN122528017APending Publication Date: 2026-08-07HUAYUNSHENGDA(BEIJING)METEROLOGICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAYUNSHENGDA(BEIJING)METEROLOGICAL TECH CO LTD
Filing Date
2026-04-20
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

但实际气象监测中,常因监测设备故障、区域监测盲区等问题,仅能获取部分常规气象要素观测资料,如何基于不完整多源数据实现高精度气象灾害识别,成为行业亟待解决的问题

Benefits of technology

[0060]本申请实施例可以获取气象要素特征序列和气象灾害图像数据集;对气象要素特征序列和气象灾害图像数据集进行多模态特征提取,得到气象要素特征序列对应的气象要素时序深度特征和气象灾害图像数据集对应的图像视觉深度特征;将气象要素时序深度特征和图像视觉深度特征进行融合,得到多模态综合特征;利用数值适配型多模态模型和图像适配型多模态模型分别对多模态综合特征进行识别,得到数值适配型多模态模型对应的第一气象灾害识别结果和图像适配型多模态模型对应的第二气象灾害识别结果;将第一气象灾害识别结果和第二气象灾害识别结果进行结果融合判断,得到第三气象灾害识别结果;对第三气象灾害识别结果进行二次校验,得到目标气象识别结果,提高基于多模态大模型的气象灾害识别的准确度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122528017A_ABST
    Figure CN122528017A_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a meteorological disaster identification method and device based on a multi-modal large model; the embodiment of the application can acquire a meteorological element feature sequence and a meteorological disaster image data set; the meteorological element feature sequence and the meteorological disaster image data set are processed to obtain meteorological element time sequence deep features and image visual deep features; the features are fused to obtain multi-modal comprehensive features; the multi-modal comprehensive features are identified by using a numerical adaptive multi-modal model and an image adaptive multi-modal model to obtain a first meteorological disaster identification result and a second meteorological disaster identification result; the first meteorological disaster identification result and the second meteorological disaster identification result are fused to obtain a third meteorological disaster identification result; the third meteorological disaster identification result is verified again to obtain a target meteorological identification result, the accuracy of meteorological disaster identification based on the multi-modal large model is improved, and the accuracy of meteorological disaster identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a method and apparatus for meteorological disaster identification based on a multimodal large model. Background Technology

[0002] Real-time and accurate identification of meteorological disasters is a core aspect of meteorological monitoring and disaster prevention and mitigation. Meteorological disasters such as heavy precipitation, dense fog with extremely low visibility, extreme winds, hail, and heavy snowfall, as well as derivative geological disasters such as landslides, are characterized by their suddenness, regionality, and high destructiveness. Rapid identification and assessment of these disasters are prerequisites for emergency response. With the development of artificial intelligence and multimodal large-scale modeling technology, disaster identification through the fusion of multi-source data has become a research hotspot. Meanwhile, the meteorological monitoring field has accumulated a large amount of observational data on conventional meteorological elements (temperature, humidity, wind speed, air pressure, etc.) and meteorological disaster image / video monitoring data, providing a data foundation for multimodal fusion disaster identification. However, in actual meteorological monitoring, due to problems such as monitoring equipment failure and regional monitoring blind spots, only partial observational data on conventional meteorological elements can be obtained. How to achieve high-precision meteorological disaster identification based on incomplete multi-source data has become an urgent problem to be solved in the industry. Summary of the Invention

[0003] This application proposes a method and apparatus for meteorological disaster identification based on a multimodal large model, which can improve the accuracy of meteorological disaster identification based on a multimodal large model.

[0004] This application provides a meteorological disaster identification method based on a multimodal large model, including:

[0005] Obtain meteorological element feature sequences and meteorological disaster image datasets;

[0006] Multimodal feature extraction is performed on the meteorological element feature sequence and the meteorological disaster image dataset to obtain the meteorological element temporal depth features corresponding to the meteorological element feature sequence and the image visual depth features corresponding to the meteorological disaster image dataset;

[0007] The temporal depth features of the meteorological elements and the visual depth features of the image are fused to obtain multimodal comprehensive features;

[0008] The multimodal integrated features are identified using a numerically adapted multimodal model and an image-adaptive multimodal model, respectively, to obtain the first meteorological disaster identification result corresponding to the numerically adapted multimodal model and the second meteorological disaster identification result corresponding to the image-adaptive multimodal model.

[0009] The first meteorological disaster identification result and the second meteorological disaster identification result are fused and judged to obtain the third meteorological disaster identification result;

[0010] The third meteorological disaster identification result is verified a second time to obtain the target meteorological identification result.

[0011] Accordingly, embodiments of this application also provide a meteorological disaster identification device based on a multimodal large model, including:

[0012] The acquisition unit is used to acquire meteorological element feature sequences and meteorological disaster image datasets;

[0013] A multimodal feature extraction unit is used to extract multimodal features from the meteorological element feature sequence and the meteorological disaster image dataset to obtain the meteorological element temporal depth features corresponding to the meteorological element feature sequence and the image visual depth features corresponding to the meteorological disaster image dataset.

[0014] The fusion unit is used to fuse the temporal depth features of the meteorological elements and the visual depth features of the image to obtain multimodal integrated features;

[0015] The identification unit is used to identify the multimodal integrated features using a numerically adapted multimodal model and an image-adapted multimodal model respectively, to obtain a first meteorological disaster identification result corresponding to the numerically adapted multimodal model and a second meteorological disaster identification result corresponding to the image-adapted multimodal model;

[0016] The judgment unit is used to perform result fusion judgment on the first meteorological disaster identification result and the second meteorological disaster identification result to obtain the third meteorological disaster identification result;

[0017] The verification unit is used to perform a secondary verification on the third meteorological disaster identification result to obtain the target meteorological identification result.

[0018] In some embodiments, the determining unit may include:

[0019] The confidence conversion subunit is used to convert the confidence of the first meteorological disaster identification result and the second meteorological disaster identification result respectively, so as to obtain the first confidence distribution corresponding to the first meteorological disaster identification result and the second confidence distribution corresponding to the second meteorological disaster identification result;

[0020] The judgment subunit is used to perform judgment processing on the first confidence distribution and the second confidence distribution based on a preset fusion judgment strategy to obtain the third meteorological disaster identification result.

[0021] In some embodiments, the determining subunit may include:

[0022] The comparison module is used to compare the first confidence distribution and the second confidence distribution with a preset confidence threshold based on the preset fusion judgment strategy.

[0023] The weighted fusion module is used to perform weighted fusion on the first confidence distribution and the second confidence distribution when the first confidence distribution is less than the preset confidence threshold and the second confidence distribution is less than the preset confidence threshold, so as to obtain the third meteorological disaster identification result;

[0024] The update module is used to use the fourth confidence distribution as the third meteorological disaster identification result when there is a fourth confidence distribution in the first confidence distribution or the second confidence distribution that is greater than or equal to the preset confidence threshold.

[0025] In some embodiments, the verification unit may include:

[0026] The first acquisition subunit is used to acquire the disaster feature matching rule base;

[0027] The verification subunit is used to perform matching verification between the third meteorological disaster identification result and the temporal depth features of meteorological elements based on the disaster feature matching rule library, and to perform matching verification between the third meteorological disaster identification result and the visual depth features of the image based on the disaster feature matching rule library;

[0028] An update subunit is used to use the third meteorological disaster identification result as the target meteorological identification result when the matching verification passes.

[0029] In some embodiments, the verification unit may further include:

[0030] The second acquisition subunit is used to acquire the temporal depth features of the target meteorological element and the visual depth features of the target image when the matching verification fails. The temporal depth features of the target meteorological element include the temporal depth features of meteorological elements that do not match the disaster feature matching rule base, and the visual depth features of the target image include the visual depth features of the image that do not match the disaster feature matching rule base.

[0031] The first adjustment subunit is used to adjust the numerical adaptation multimodal model using the temporal depth features of the target meteorological elements and the third meteorological disaster identification results, so as to obtain the adjusted numerical adaptation multimodal model.

[0032] The second adjustment subunit is used to adjust the image-adaptive multimodal model using the visual depth features of the target image and the third meteorological disaster identification result, to obtain the adjusted image-adaptive multimodal model;

[0033] The identification subunit is used to identify the meteorological element feature sequence using the adjusted numerical adaptation multimodal model, and to identify the meteorological disaster image dataset using the adjusted image adaptation multimodal model, so as to obtain the target meteorological identification result.

[0034] In some embodiments, the apparatus proposed in this application may further include:

[0035] The second acquisition unit is used to acquire the numerical adaptive multimodal model to be trained, the image adaptive multimodal model to be trained, meteorological element training data, disaster image training set, disaster labels, and disaster text description labels, wherein the meteorological element training data and the disaster image training set have a temporal correlation relationship;

[0036] The second identification unit is used to identify the meteorological element training data and the disaster image training set using the numerical adaptive multimodal model to be trained, and to obtain the fourth meteorological identification result.

[0037] The third identification unit is used to identify the disaster image training set using the image-adaptive multimodal model to be trained, and obtain the fifth meteorological identification result;

[0038] The first calculation unit is used to calculate the time-series causal inference loss based on the fourth meteorological identification result and the disaster label;

[0039] The second calculation unit is used to calculate the visual semantic understanding loss based on the fifth meteorological identification result and the disaster text description label;

[0040] The loss fusion unit is used to fuse the temporal causal reasoning loss and the visual semantic understanding loss to obtain a fusion loss;

[0041] An adjustment unit is used to adjust the numerical adaptive multimodal model and the image adaptive multimodal model to be trained using the fusion loss, respectively, to obtain the numerical adaptive multimodal model and the image adaptive multimodal model.

[0042] In some embodiments, the first computing unit includes:

[0043] The first calculation subunit is used to calculate the causal order prediction loss based on the fourth meteorological identification result and the disaster label;

[0044] The weight acquisition subunit is used to acquire the attention weights and prior weights of the numerically adapted multimodal model to be trained.

[0045] The second calculation subunit is used to calculate the disaster classification loss based on the attention weight and prior weight;

[0046] The third calculation subunit is used to calculate the first disaster classification cross-entropy loss based on the meteorological element training data and the fourth meteorological identification result;

[0047] The first fusion subunit is used to fuse the causal order prediction loss, the disaster classification loss, and the multi-class cross-entropy loss to obtain the temporal causal inference loss.

[0048] In some embodiments, the second computing unit includes:

[0049] The fourth calculation subunit is used to calculate the image-text contrast loss based on the fifth meteorological identification result and the disaster text description label;

[0050] The fifth calculation subunit is used to extract disaster text description information from the fifth meteorological identification result, and calculate the image description generation loss based on the fifth meteorological identification result and the disaster text description information;

[0051] The sixth calculation subunit is used to calculate the second disaster classification cross-entropy loss based on the disaster image training set and the fifth meteorological recognition result;

[0052] The seventh calculation subunit is used to obtain the positive sample training set of the disaster image training set, and calculate the visual feature discrimination loss based on the positive sample training set and the fifth meteorological recognition result;

[0053] The second fusion subunit is used to fuse the image-text contrast loss, the image description generation loss, the second disaster classification cross-entropy loss, and the visual feature discrimination loss to obtain the visual semantic understanding loss.

[0054] In some embodiments, the loss fusion unit includes:

[0055] The eighth calculation subunit is used to calculate the cross-modal contrast loss based on the fourth meteorological identification result and the fifth meteorological identification result;

[0056] The ninth calculation subunit is used to calculate the mutual information maximization loss based on the fourth meteorological identification result and the fifth meteorological identification result;

[0057] The third fusion subunit is used to perform weighted fusion of the cross-modal contrast loss, the mutual information maximization loss, the temporal causal reasoning loss, and the visual semantic understanding loss to obtain the fusion loss.

[0058] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative embodiments described above.

[0059] Accordingly, this application embodiment also provides a storage medium storing instructions, which, when executed by a processor, implement any of the meteorological disaster identification methods based on a multimodal large model provided in this application embodiment.

[0060] This application embodiment can acquire meteorological element feature sequences and meteorological disaster image datasets; perform multimodal feature extraction on the meteorological element feature sequences and meteorological disaster image datasets to obtain the meteorological element temporal depth features corresponding to the meteorological element feature sequences and the image visual depth features corresponding to the meteorological disaster image datasets; fuse the meteorological element temporal depth features and the image visual depth features to obtain multimodal comprehensive features; use numerical adaptation multimodal models and image adaptation multimodal models to identify the multimodal comprehensive features respectively to obtain the first meteorological disaster identification result corresponding to the numerical adaptation multimodal model and the second meteorological disaster identification result corresponding to the image adaptation multimodal model; fuse the first meteorological disaster identification result and the second meteorological disaster identification result to obtain the third meteorological disaster identification result; perform secondary verification on the third meteorological disaster identification result to obtain the target meteorological identification result, thereby improving the accuracy of meteorological disaster identification based on a multimodal large model. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] Figure 1 This is a schematic diagram of a scenario for the meteorological disaster identification method based on a multimodal large model provided in an embodiment of this application;

[0063] Figure 2 This is a flowchart illustrating the meteorological disaster identification method based on a multimodal large model provided in this application embodiment;

[0064] Figure 3 This is a schematic diagram of the structure of the meteorological disaster identification device based on a multimodal large model provided in the embodiments of this application;

[0065] Figure 4This is a schematic diagram of the structure of the computer device provided in the embodiments of this application. Detailed Implementation

[0066] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. However, the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0067] This application proposes a meteorological disaster identification method based on a multimodal large model. This method can be executed by a multimodal large model-based meteorological disaster identification device, which can be integrated into a computer device. The computer device can include at least one of a terminal and a server. That is, the multimodal large model-based meteorological disaster identification method proposed in this application can be executed by a terminal, a server, or jointly by a terminal and a server capable of communicating with each other.

[0068] The terminal may include, but is not limited to, smartphones, tablets, laptops, personal computers (PCs), smart home appliances, wearable electronic devices, VR / AR devices, in-vehicle terminals, intelligent voice interaction devices, etc.

[0069] A server can be an interconnecting server between multiple heterogeneous systems or a backend server. It can also be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, and big data and artificial intelligence platforms, etc.

[0070] It should be noted that the embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.

[0071] In one embodiment, such as Figure 1The meteorological disaster identification device based on a multimodal large model can be integrated into computer equipment such as a terminal or server to implement the meteorological disaster identification method based on a multimodal large model proposed in this application. Specifically, the server 11 or terminal 10 can acquire meteorological element feature sequences and meteorological disaster image datasets; perform multimodal feature extraction on the meteorological element feature sequences and meteorological disaster image datasets to obtain the meteorological element temporal depth features corresponding to the meteorological element feature sequences and the image visual depth features corresponding to the meteorological disaster image datasets; fuse the meteorological element temporal depth features and the image visual depth features to obtain multimodal comprehensive features; use numerical adaptation multimodal models and image adaptation multimodal models to identify the multimodal comprehensive features respectively to obtain the first meteorological disaster identification result corresponding to the numerical adaptation multimodal model and the second meteorological disaster identification result corresponding to the image adaptation multimodal model; perform result fusion judgment on the first meteorological disaster identification result and the second meteorological disaster identification result to obtain the third meteorological disaster identification result; perform secondary verification on the third meteorological disaster identification result to obtain the target meteorological identification result.

[0072] The following will provide a detailed description of each example. It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.

[0073] This application will describe the embodiments from the perspective of a meteorological disaster identification device based on a multimodal large model. This meteorological disaster identification device based on a multimodal large model can be integrated into a computer device, which can be a server or a terminal or other similar device.

[0074] like Figure 2 The present invention provides a meteorological disaster identification method based on a multimodal large model, the specific process of which includes:

[0075] 101. Obtain meteorological element feature sequences and meteorological disaster image datasets.

[0076] In some embodiments, multi-source raw data required for meteorological disaster identification can be collected. This multi-source raw data is divided into two main data types: meteorological element observation data and meteorological disaster image / video data.

[0077] In some embodiments, meteorological element observation data are collected in real time through ground meteorological monitoring stations and mobile meteorological monitoring equipment (such as vehicle-mounted meteorological instruments, UAV meteorological detection systems, portable handheld weather stations, etc.) to obtain partial or complete conventional meteorological numerical data within the monitoring area. The collected meteorological elements include at least the following six core parameters: temperature (unit: °C), relative humidity (unit: %), wind speed (unit: m / s, selectable instantaneous wind speed or 2-minute average wind speed), wind direction (unit: °), air pressure (unit: hPa), and precipitation (unit: mm, including minute-by-minute cumulative precipitation or hourly cumulative precipitation). Depending on actual operational needs, auxiliary elements such as visibility, ground temperature, and solar radiation can also be collected.

[0078] In some embodiments, the data acquisition time interval can be flexibly configured according to the monitoring scenario, ranging from 1 minute to 10 minutes. For example, for severe convective weather (such as short-duration heavy rainfall, thunderstorms, and strong winds), a 1-minute high-resolution sampling is used to capture rapid changes in meteorological elements. For other examples, for relatively slow-evolving disasters such as dense fog and persistent precipitation, the interval can be appropriately reduced to 5 minutes or 10 minutes to save storage and transmission resources.

[0079] In some embodiments, the raw meteorological element data is stored in the form of a structured numerical sequence, with each record being a feature vector of a time step, arranged in ascending order of time.

[0080] In some embodiments, meteorological disaster image / video data is collected using various devices such as meteorological monitoring cameras, drones, and satellite remote sensing, covering the disaster occurrence process within the monitoring area. Video data is extracted as static images at a frame rate of 1–5 frames per second, forming an unstructured visual dataset. High frame rates (e.g., 5 frames per second) are suitable for capturing rapidly evolving disaster details such as heavy precipitation and extreme winds, while low frame rates (e.g., 1–2 frames per second) are used for slowly changing processes such as dense fog and continuous snowfall, to reduce data redundancy. The collected data can present typical visual features of disasters such as heavy precipitation (rain curtain, water accumulation), dense fog (reduced visibility, blurred outlines), and landslides (cracks, debris), providing raw visual material for subsequent image feature extraction and disaster identification.

[0081] In some embodiments, the original meteorological element data can be cleaned, normalized, and completed to obtain standardized meteorological element data and image data, solving the problems of incomplete meteorological element data and inconsistent data formats, and preparing for subsequent feature extraction.

[0082] Data cleaning refers to removing outliers and missing values ​​(random missing values ​​not related to data completion) from time-series meteorological data. Specifically, the 3σ criterion can be used to identify outliers and delete invalid data.

[0083] Normalization refers to mapping the original values ​​of various meteorological elements (such as temperature, humidity, wind speed, air pressure, precipitation, etc.) to a unified range of [0,1] through linear transformation, so as to eliminate the adverse effects of differences in dimensions and orders of magnitude between different elements on model training.

[0084] Incomplete data completion refers to the process in actual meteorological monitoring where, due to equipment failure, communication interruption, or monitoring blind spots, only observational data for some conventional meteorological elements (e.g., only temperature, humidity, and wind speed, lacking air pressure, precipitation, etc.) can be obtained. Data completion techniques are used to generate effective feature representations of the missing elements to meet the requirements of subsequent time-series models for complete input. For scenarios where only some conventional meteorological elements are available, a lightweight completion model based on meteorological element association rules (such as a Bayesian network) is adopted. This model completes the feature trends of the missing elements based on existing elements (such as temperature, humidity, and wind speed), rather than completing specific numerical values, thus preserving the temporal variation characteristics of the data and adapting to subsequent time-series feature extraction.

[0085] In some embodiments, for meteorological disaster image / video data, the acquired images can be denoised (Gaussian filtering), normalized in size (uniformly adjusted to 224×224 / 448×448 pixels), and converted in color space (RGB to YUV) to eliminate the effects of image noise and size differences.

[0086] For video frames, static images can be extracted from the video data at a set frame rate, preserving the temporal visual features of the video and forming an image temporal sequence.

[0087] Through the above processing, meteorological element feature sequences and meteorological disaster image datasets can be obtained.

[0088] 102. Multimodal feature extraction is performed on the meteorological element feature sequence and the meteorological disaster image dataset to obtain the meteorological element temporal depth features corresponding to the meteorological element feature sequence and the image visual depth features corresponding to the meteorological disaster image dataset.

[0089] In some embodiments, an LSTM / GRU temporal neural network (preferably LSTM) can be used to adapt to the temporal variation patterns of meteorological elements and mine the temporal correlation features of the data.

[0090] For example, a standardized meteorological element feature sequence is input into an LSTM network. Through temporal operations in the input layer, hidden layer, and output layer, a temporal deep feature vector of the meteorological element is extracted. The vector dimension can be adjusted according to the model training effect, with 128 / 256 dimensions being preferred.

[0091] In some embodiments, meteorological element observation data (such as temperature, humidity, wind speed, air pressure, precipitation, etc.) are essentially structured numerical sequences sampled over time, exhibiting clear temporal variation patterns: for example, humidity continuously rises and air pressure decreases before heavy precipitation occurs; wind speed shows an increasing trend of fluctuation before strong winds arrive. Therefore, LSTM networks can be used to extract multimodal features from meteorological element feature sequences. LSTM has a stronger memory capacity and a more refined gating mechanism, making it particularly suitable for mixed temporal patterns in meteorological data that are long-lasting (e.g., fog lasting several hours) or sudden (e.g., short-duration heavy precipitation lasting 10 minutes). Through supervised training, LSTM networks can automatically learn the cooperative relationship between different meteorological elements, such as the association between the pattern of "wind speed first increasing and then decreasing, accompanied by air pressure first decreasing and then increasing" and extreme wind disasters, thereby achieving high-precision temporal feature extraction of disaster types. The LSTM network can include an input layer, hidden layers, and an output layer. The input layer receives the meteorological element feature sequence, the hidden layer consists of several LSTM units, which perform convolution processing on the meteorological element feature sequence, and then the output layer performs activation function processing on the convolutional meteorological element feature sequence to obtain the time-series deep feature vector of meteorological elements.

[0092] In some embodiments, for meteorological disaster image datasets, a pre-trained multimodal visual large model (such as CLIP or BLIP) can be used to extract multimodal features to obtain the image visual depth features corresponding to the meteorological disaster image dataset. Specifically, standardized meteorological disaster images can be input into the pre-trained multimodal visual large model, and the image visual depth features (image depth feature vectors) of the images can be extracted through the model's convolutional layers and attention layers. The vector dimension is preferably 128 / 256 dimensions.

[0093] 103. The temporal depth features of meteorological elements and the visual depth features of images are fused to obtain multimodal integrated features.

[0094] In one embodiment, the fusion in this application refers to the deep fusion of the temporal depth features of meteorological elements and the visual depth features of images, rather than simple splicing, so as to explore the inherent relationship between the two and form a fused multimodal comprehensive feature, thereby improving the feature representation capability.

[0095] Attention-based fusion algorithms (such as Cross-Modal Attention) can be used to fuse temporal deep features of meteorological elements and visual deep features of images to obtain multimodal comprehensive features. This algorithm can automatically learn the association weights between meteorological element features and visual features of images, highlighting feature information that plays an important role in disaster identification.

[0096] More specifically, a cross-modal attention matrix can be constructed, and the attention weights corresponding to the temporal deep features of meteorological elements and the visual deep features of images in the cross-modal attention matrix can be calculated to quantify the degree of correlation between the two. For example, important features can be assigned high weights, and secondary features can be assigned low weights. Then, the attention matrix can be used to perform a weighted summation of the temporal deep features of meteorological elements and the visual deep features of images, and then feature mapping can be performed through a fully connected layer to obtain a multimodal comprehensive feature vector with unified dimensions and deep feature fusion (the preferred dimension is 256 / 512).

[0097] In some embodiments, even if meteorological element data is incomplete, the attention mechanism can still highlight the role of visual features by associating the temporal characteristics of existing elements with the visual features of the image, thus ensuring the effectiveness of the fused features.

[0098] 104. Using a numerically adapted multimodal model and an image-adapted multimodal model respectively, the multimodal integrated features are identified to obtain the first meteorological disaster identification result corresponding to the numerically adapted multimodal model and the second meteorological disaster identification result corresponding to the image-adapted multimodal model.

[0099] In some embodiments, a multimodal large model collaborative reasoning mechanism can be constructed, using two multimodal large models adapted to different modalities (e.g., numerically adapted multimodal model M1 and image adapted multimodal model M2) to perform collaborative reasoning on the fused multimodal comprehensive feature vector and output a preliminary judgment result of the disaster type.

[0100] In some embodiments, a large multimodal model (such as GPT-4V or Qwen-VL) with strong adaptability to numerical time-series data can be selected. This model has multimodal processing capabilities and can provide causal logical reasoning relationships between preceding and following time series by providing time-series text data and image data. Based on this large multimodal model, it can be fine-tuned using a meteorological disaster dataset (including meteorological element data, disaster images, and disaster labels) to enhance the model's ability to reason about the correlation between meteorological elements and disaster types.

[0101] In some embodiments, a multimodal large model with strong adaptability to visual image features, such as BLIP-2 (or LLaVA), can be selected. This model is a pre-trained large model for image encoding, which achieves accurate description of image features by aligning with text. Based on this pre-trained model, fine-tuning is performed using the same meteorological disaster dataset to enhance the model's ability to infer the correlation between disaster image features and disaster types.

[0102] In some embodiments, the method proposed in this application may further include:

[0103] The following are obtained: a numerical adaptive multimodal model to be trained, an image adaptive multimodal model to be trained, meteorological element training data, a disaster image training set, disaster labels, and disaster text description labels, wherein the meteorological element training data and the disaster image training set have a temporal correlation relationship;

[0104] The numerically adapted multimodal model to be trained is used to identify meteorological element training data and disaster image training set to obtain the fourth meteorological identification result;

[0105] The image-adaptive multimodal model to be trained is used to identify the disaster image training set, and the fifth meteorological identification result is obtained.

[0106] Based on the fourth meteorological identification results and disaster labels, calculate the time-series causal inference loss;

[0107] Based on the fifth meteorological identification result and the disaster text description label, calculate the visual semantic understanding loss;

[0108] The temporal causal reasoning loss and the visual semantic understanding loss are fused to obtain the fusion loss;

[0109] The numerical adaptive multimodal model and the image adaptive multimodal model to be trained are adjusted using the fusion loss to obtain the numerical adaptive multimodal model and the image adaptive multimodal model.

[0110] The models to be trained refer to two initialized multimodal large models, one focusing on numerical time series (such as the fine-tuned version of GPT-4V) and the other focusing on image vision (such as the fine-tuned version of BLIP-2), which have not yet been trained in this step.

[0111] Meteorological element training data refers to structured time-series numerical sequences (temperature, humidity, wind speed, etc.).

[0112] Disaster image training set refers to images / video frames collected in the same time period and area as meteorological element data.

[0113] Disaster labels refer to the type of disaster corresponding to each sample (such as heavy rainfall, dense fog, etc.), which are used for classification and supervision.

[0114] Disaster text description tags refer to textual descriptions of disaster images (such as "visibility less than 100 meters, roads covered by dense fog"), which are used to generate monitoring.

[0115] Temporal correlation refers to the strict temporal alignment between meteorological element data and image data (e.g., meteorological observations and image captures at the same time or within the same time window), ensuring that the model can learn cross-modal causal relationships.

[0116] In some embodiments, a fourth meteorological recognition result is obtained by using a numerically adapted multimodal model to be trained to identify meteorological element training data and disaster image training set. This can refer to inputting meteorological element training data and disaster image training set into the numerically adapted multimodal model to be trained to obtain the fourth meteorological recognition result. The fourth meteorological recognition result may include all content generated by the numerically adapted multimodal model to be trained during the training process (such as attention weights, prior weights, data features, disaster prediction probabilities, etc.).

[0117] In some embodiments, using an image-adaptive multimodal model to be trained to identify a disaster image training set and obtain a fifth meteorological identification result can refer to inputting the disaster image training set into the image-adaptive multimodal model to be trained to obtain the fifth meteorological identification result. The fifth meteorological identification result may include all content generated by the image-adaptive multimodal model during training (e.g., attention weights, prior weights, data features, disaster description information, disaster prediction probability, etc.).

[0118] In some embodiments, the step "calculating the time-series causal inference loss based on the fourth meteorological identification result and the disaster label" may include:

[0119] Based on the fourth meteorological identification results and the disaster labels, the causal order is calculated to predict the loss;

[0120] Obtain the attention weights and prior weights of the numerically adapted multimodal model to be trained;

[0121] The disaster classification loss is calculated based on attention weights and prior weights.

[0122] Based on meteorological element training data and fourth meteorological identification results, calculate the cross-entropy loss of the first disaster classification.

[0123] The temporal causal inference loss is obtained by fusing the causal order prediction loss, disaster classification loss, and multi-class cross-entropy loss.

[0124] In some embodiments, the causal order prediction loss can be calculated based on the fourth meteorological identification result and the disaster label according to the following formula:

[0125]

[0126] in, The true meteorological element values ​​that are being concealed. MSE refers to the mean squared error function, which is the model's predicted value. This loss forces M1 to understand the evolution trend of meteorological elements, laying the foundation for subsequent causal inference of disasters.

[0127] In some embodiments, disaster classification loss can be calculated based on attention weights and prior weights according to the following formula:

[0128]

[0129] in, These are the attention weights for the model at time t. Prior weights (e.g., the weight should be higher for the time frame 5-15 minutes before a disaster). Disaster classification models can encourage M1 to focus on temporal causal windows that are strongly correlated with disaster occurrence.

[0130] In some embodiments, the first disaster classification cross-entropy loss can be calculated based on meteorological element training data and the fourth meteorological identification result according to the following formula:

[0131]

[0132] Where B is the number of samples in the current training batch, and C is the total number of disaster categories (e.g., heavy precipitation, dense fog, extremely strong winds, hail, heavy snowfall, landslides, etc., the value of C is set according to actual business needs). Let the i-th sample be the one-hot encoding of the true label. If this sample belongs to class c, then... Otherwise, it is 0. This represents the probability of class c for the i-th sample data predicted by the M1 model.

[0133] In some embodiments, the causal order prediction loss, disaster classification loss, and multi-class cross-entropy loss can be fused according to the following formula to obtain the temporal causal inference loss:

[0134]

[0135] in, This refers to the loss in temporal causal reasoning.

[0136] In some embodiments, the step "calculating visual semantic understanding loss based on the fifth meteorological identification result and the disaster text description label" may include:

[0137] Based on the fifth meteorological identification results and disaster text description labels, the image text contrast loss is calculated;

[0138] Extract disaster text description information from the fifth meteorological identification results, and calculate the image description generation loss based on the fifth meteorological identification results and disaster text description information;

[0139] Based on the disaster image training set and the fifth meteorological identification results, the cross-entropy loss of the second disaster classification is calculated;

[0140] Obtain a positive sample training set of the disaster image training set, and calculate the visual feature discrimination loss based on the positive sample training set and the fifth meteorological recognition results;

[0141] The image-text contrast loss, image description generation loss, second hazard classification cross-entropy loss, and visual feature discrimination loss are fused together to obtain the visual semantic understanding loss.

[0142] In some embodiments, for each sample, image feature v is associated with the corresponding disaster text description label. This constitutes a positive alignment with other label text features in the same batch. By constructing negative pairs, the image-text contrast loss can then be calculated using the following formula:

[0143]

[0144] Where k represents the index of the negative sample. The temperature information is a hyperparameter used to adjust the smoothness of the loss. The value of M2 is typically 0.07 or 0.1. This loss ensures that M2 can map the visual features of the image to the correct semantic category space, improving the ability to distinguish disaster types.

[0145] In some embodiments, disaster text description information can be extracted from the fifth meteorological identification result. Based on the fifth meteorological identification result and the disaster text description information, the image description generation loss is calculated, as follows:

[0146]

[0147] Where S is the total length of the generated disaster text description information (i.e., the number of words or tags). This refers to the s-th word in the description text. Let v be the image visual feature vector, representing all words generated up to the s-th word. The loss function sums the negative log probabilities over all time steps; minimizing this loss is equivalent to maximizing the likelihood of generating the entire correct descriptive text. In this way, the model learns to predict the next most likely word at each generation, based on the image content and the already generated text, thus generating a complete, fluent disaster description consistent with the image content (e.g., "Visibility less than 100 meters, roads covered by dense fog").

[0148] In some embodiments, the second disaster classification cross-entropy loss can be calculated based on the disaster image training set and the fifth and sixth meteorological recognition results. The calculation method for the cross-entropy loss of the second disaster category can be referenced from that for the first disaster category.

[0149] In some embodiments, a positive sample training set of the disaster image training set can be obtained. Based on the positive sample training set and the fifth meteorological recognition result, the visual feature discrimination loss is calculated, as follows:

[0150]

[0151] in, This refers to visual feature discrimination loss. Visual feature discrimination loss can enhance the cohesion of visual features for similar disasters. For example, fog images from different angles and under different lighting conditions will have their features clustered together. Furthermore, it can distinguish visually similar disasters. For instance, both fog and heavy snowfall reduce visibility and are easily confused; this loss forces the model to differentiate between these two types of features, significantly reducing the false positive rate. Additionally, it can improve the discrimination ability for disasters with small sample sizes. For example, for disasters with low occurrence frequency (such as hail), through comparative learning between similar samples, the model can more effectively utilize limited positive samples, improving recognition performance.

[0152] In some embodiments, the image-text contrast loss, image description generation loss, second hazard classification cross-entropy loss, and visual feature discrimination loss can be fused to obtain the visual semantic understanding loss:

[0153]

[0154] The semantic understanding loss model incorporates four tail constraints: alignment, generation, contrast, and classification. This allows the model to simultaneously satisfy multiple supervision signals, resulting in more discriminative and generalizable features.

[0155] In some embodiments, the step of "fusing the temporal causal reasoning loss and the visual semantic understanding loss to obtain a fusion loss" may include:

[0156] Based on the fourth and fifth meteorological identification results, the cross-modal contrast loss is calculated;

[0157] Based on the fourth and fifth meteorological identification results, the mutual information maximization loss is calculated;

[0158] The cross-modal contrast loss, mutual information maximization loss, temporal causal reasoning loss, and visual semantic understanding loss are weighted and fused to obtain the fusion loss.

[0159] In some embodiments, the features output by M1 in the same disaster sample can be... (Temporal causal characteristics) and characteristics of M2 output (Visual semantic features) are used as positive pairs, and cross-modal features of different disaster samples are used as negative pairs. The cross-modal contrastive loss is calculated based on the following formula:

[0160]

[0161] This loss brings M1 and M2 closer in their understanding of the same disaster event, making their characteristics spatially similar and facilitating subsequent fusion.

[0162] In some embodiments, the mutual information maximization loss can be calculated according to the following formula:

[0163]

[0164] in, It is the mutual information between two feature vectors, used to measure the amount of information they contain together. The Hilbert-Schmidt independence criterion is used to measure the statistical dependence between two variables. A larger value indicates a stronger correlation, while a value of 0 indicates independence. This is the regularization coefficient, used to control the strength of the redundancy penalty. Specifically, by adjusting the regularization coefficient... With proper configuration, the features of M1 and M2 can share core disaster information while retaining unique cues of their respective modalities. By maximizing mutual information, the two features are aligned as much as possible, facilitating subsequent cross-modal attention fusion. By minimizing HSIC, redundancy between the two is reduced, feature duplication is avoided, and the model is forced to retain modality-specific information (e.g., M1 retains temporal causal trends, and M2 retains spatial visual details).

[0165] In some embodiments, the cross-modal contrastive loss, mutual information maximization loss, temporal causal reasoning loss, and visual semantic understanding loss are weighted and fused to obtain the fusion loss. Specifically, the fusion loss can be obtained according to the following formula:

[0166]

[0167] Through this loss design, M1 can learn causal dependencies in temporal data (through mask prediction and attention regularization), M2 can learn fine-grained visual semantics (through image-text comparison and description generation), and the alignment loss ensures that the feature spaces of the two can complement and merge, ultimately achieving a synergistic effect of "temporal causal reasoning + visual semantic understanding".

[0168] In some embodiments, a fusion loss is used to adjust the numerically adapted multimodal model and the image-adapted multimodal model to be trained, respectively, to obtain a numerically adapted multimodal model and an image-adapted multimodal model. For example, the fusion loss is used to adjust the model parameters of the numerically adapted multimodal model to be trained to obtain a numerically adapted multimodal model. As another example, the fusion loss is used to adjust the parameters of the image-adapted multimodal model to be trained to obtain an image-adapted multimodal model.

[0169] 105. The results of the first meteorological disaster identification and the second meteorological disaster identification are fused and judged to obtain the third meteorological disaster identification result.

[0170] In some embodiments, the step of "fusion and judgment of the first meteorological disaster identification result and the second meteorological disaster identification result to obtain the third meteorological disaster identification result" may include:

[0171] The confidence levels of the first and second meteorological disaster identification results are transformed to obtain the first confidence distribution corresponding to the first meteorological disaster identification result and the second confidence distribution corresponding to the second meteorological disaster identification result.

[0172] Based on a pre-defined fusion judgment strategy, the first confidence distribution and the second confidence distribution are judged and processed to obtain the third meteorological disaster identification result.

[0173] The confidence level is defined as follows: for the output vector based on two models... Each component represents the output value for a specific weather hazard, which is then converted into a confidence distribution using the Softmax function.

[0174]

[0175] We obtain data with a confidence level between 0 and 1. For example, the confidence level for heavy precipitation is 95%, and the confidence level for fog is 3%.

[0176] In some embodiments, the step "based on a preset fusion judgment strategy, performing judgment processing on the first confidence distribution and the second confidence distribution to obtain the third meteorological disaster identification result" may include:

[0177] Based on the preset fusion judgment strategy, the first confidence distribution and the second confidence distribution are compared with the preset confidence threshold respectively;

[0178] When the first confidence distribution is less than the preset confidence threshold and the second confidence distribution is less than the preset confidence threshold, the first confidence distribution and the second confidence distribution are weighted and fused to obtain the third meteorological disaster identification result;

[0179] When there is a fourth confidence distribution in the first confidence distribution or the second confidence distribution that is greater than or equal to the preset confidence threshold, the fourth confidence distribution is used as the third meteorological disaster identification result.

[0180] For example, a confidence threshold (preferably 80%) is set. If the confidence of the inference result of a single model is greater than or equal to the threshold, it is directly used as the preliminary judgment result. If none of them reach the threshold, a weighted voting method is used (the weights of M1 and M2 are adjusted according to the model training effect, preferably 0.5 each) to select the disaster type with the highest weighted confidence as the preliminary judgment result. If the inference results of the two models are consistent and the confidence is greater than or equal to the threshold, the result is directly confirmed to improve the reliability of the inference.

[0181] 106. Perform a second verification on the third meteorological disaster identification results to obtain the target meteorological identification results.

[0182] In some embodiments, the step "performing a secondary verification of the third meteorological disaster identification result to obtain the target meteorological identification result" may include:

[0183] Obtain a disaster feature matching rule base;

[0184] The third meteorological disaster identification result is matched and verified with the temporal depth features of meteorological elements based on the disaster feature matching rule base, and the third meteorological disaster identification result is matched and verified with the visual depth features of the image based on the disaster feature matching rule base;

[0185] When the matching verification passes, the third meteorological disaster identification result is used as the target meteorological identification result.

[0186] The disaster feature matching rule base is a pre-built knowledge base that stores the judgment criteria for various meteorological disasters. The rule base is constructed based on operational specifications such as the "National Meteorological Disaster Collection and Reporting Technical Specifications" and expert experience. Each rule typically includes elements such as disaster type (e.g., heavy precipitation, dense fog, extreme winds, landslides), meteorological element matching conditions, image visual matching conditions, and logical combinations. For example, meteorological element matching conditions could include: extreme winds requiring wind speed ≥ 17.2 m / s; heavy precipitation requiring hourly precipitation ≥ 16 mm (or a higher threshold); and dense fog requiring visibility < 1 km. For instance, meteorological element matching conditions could include: dense fog images should show reduced visibility and blurred outlines; landslide images should show cracks, landslides, or deposits; and extreme wind images should show visual evidence such as fallen trees and swaying objects. Logical combinations could include: some disasters may require multiple conditions to be met simultaneously (e.g., "wind speed ≥ 17.2 m / s and fallen trees in the image"), or flexible rules such as "one of the meteorological element conditions or image conditions."

[0187] In some embodiments, the third meteorological disaster identification result can be matched and verified with the temporal depth features of meteorological elements based on a disaster feature matching rule base, and the third meteorological disaster identification result can be matched and verified with the visual depth features of images based on the disaster feature matching rule base. For example, the rule base defines the judgment conditions of meteorological elements for "extreme wind" disasters (such as "maximum wind speed ≥ 17.2 m / s in the past 10 minutes"). The system starts from... The wind speed feature value is obtained either by inverse kinematics or directly, and it is then determined whether the threshold is met. If it is, the meteorological element matching is successful; otherwise, it fails.

[0188] The matching verification is considered successful if both branches match successfully (or at least one branch matches successfully according to preset rules, or a weighted score is used). If any key condition is not met, the verification fails.

[0189] When the matching verification passes, the system confirms that the third meteorological disaster identification result has logical consistency and is supported by multi-source evidence, and outputs it as the target meteorological identification result. This result can be directly used in subsequent automatic image recognition units (adding recognition boxes and labels), or output to the operational early warning system.

[0190] In some embodiments, the method proposed in this application may further include:

[0191] When the matching verification fails, the temporal depth features of the target meteorological elements and the visual depth features of the target image are obtained. The temporal depth features of the target meteorological elements include the temporal depth features of meteorological elements that do not match the disaster feature matching rule base, and the visual depth features of the target image include the visual depth features of the image that do not match the disaster feature matching rule base.

[0192] The numerical adaptation multimodal model is adjusted using the temporal depth features of the target meteorological elements and the results of the third meteorological disaster identification, resulting in an adjusted numerical adaptation multimodal model.

[0193] The image-adaptive multimodal model is adjusted using the visual depth features of the target image and the results of third-party meteorological disaster identification to obtain the adjusted image-adaptive multimodal model;

[0194] The adjusted numerical adaptation multimodal model is used to identify and process the meteorological element feature sequence, and the adjusted image adaptation multimodal model is used to identify and process the meteorological disaster image dataset, so as to obtain the target meteorological identification result.

[0195] Among them, the temporal depth features of the target meteorological elements refer to the specific meteorological element features that cause the verification to fail. For example, if the model outputs "extreme wind" but the actual wind speed feature value is far below the threshold, then the wind speed feature and its temporal variation pattern are extracted; or if the model outputs "heavy precipitation" but the precipitation feature value is zero, then this contradictory feature is extracted. These features do not meet the requirements of the rule base.

[0196] The visual depth features of the target image refer to the specific visual features of the image that cause the verification to fail. For example, if the model outputs "heavy fog", but the image does not show typical features such as reduced visibility or blurred outlines, but instead shows features of heavy snowfall such as falling snowflakes, then these conflicting visual features should be extracted.

[0197] Then, these mismatched features and their corresponding incorrect outputs can be used as new training samples (negative samples) to perform mini-batch gradient updates on the numerically adapted multimodal model. For example, an auxiliary loss function, such as contrastive loss or triplet loss, can be designed to force the model to distance the incorrect sample from the incorrect disaster type in the feature space, while simultaneously narrowing the distance to the correct disaster type. Similarly, these image features that lead to verification failures and the model's incorrect outputs can be used as negative samples to fine-tune the image-adapted multimodal model. The backward gradient of contrastive learning or classification loss can be used to penalize the model's misclassification of these visual features.

[0198] Then, the adjusted numerically adapted multimodal model can be used to re-identify the original meteorological element feature sequence; at the same time, the adjusted image-adaptive multimodal model can be used to re-identify the original meteorological disaster image dataset. Here, "re-identification" can be understood as: using the updated model to re-extract features, and then generating new disaster type judgments again through the feature fusion module and collaborative reasoning module.

[0199] Through this embodiment, the system can proactively learn from failed verification cases and automatically correct model parameters without manual intervention. This significantly improves the model's robustness and adaptability in real-world business environments. Furthermore, fine-tuning using only the small batch of features that caused the failures avoids the high cost of full retraining and enables rapid response to newly emerging misjudgment patterns. The re-identified target meteorological results, after adjustment, undergo secondary verification through a rule base, ensuring the business compliance and multi-source consistency of the output results, thereby meeting the high standards required for actual meteorological monitoring.

[0200] In this embodiment, meteorological element feature sequences and meteorological disaster image datasets can be obtained. Multimodal feature extraction is performed on the meteorological element feature sequences and meteorological disaster image datasets to obtain the temporal depth features of the meteorological element feature sequences and the visual depth features of the images corresponding to the meteorological disaster image datasets. The temporal depth features of the meteorological elements and the visual depth features of the images are fused to obtain multimodal comprehensive features. Numerical adaptation multimodal models and image adaptation multimodal models are used to identify the multimodal comprehensive features, respectively, to obtain the first meteorological disaster identification result corresponding to the numerical adaptation multimodal model and the second meteorological disaster identification result corresponding to the image adaptation multimodal model. The first meteorological disaster identification result and the second meteorological disaster identification result are fused and judged to obtain the third meteorological disaster identification result. The third meteorological disaster identification result is then subjected to secondary verification to obtain the target meteorological identification result. This application's deep fusion method of meteorological element temporal features and image visual features based on an attention mechanism solves the feature fusion problem of incomplete meteorological observation data and achieves efficient association of multimodal features. In addition, this application constructs a multimodal large model collaborative reasoning mechanism, adopts a result fusion rule of parallel reasoning of meteorological and visual adaptive dual multimodal large models and confidence threshold, takes into account the reasoning ability of numerical and visual features, and improves the accuracy of disaster identification.

[0201] To better implement the meteorological disaster identification method based on a multimodal large model provided in this application, one embodiment also provides a meteorological disaster identification device based on a multimodal large model, which can be integrated into a computer device. The meanings of the terms used are the same as in the above-described meteorological disaster identification method based on a multimodal large model, and specific implementation details can be found in the descriptions in the method embodiments.

[0202] In one embodiment, a meteorological disaster identification device based on a multimodal large model is provided. This device can be integrated into a computer device, such as... Figure 3 As shown, the meteorological disaster identification device based on a multimodal large model includes: an acquisition unit 201, a multimodal feature extraction unit 202, a fusion unit 203, an identification unit 204, a judgment unit 205, and a verification unit 204, as detailed below:

[0203] Acquisition unit 201 is used to acquire meteorological element feature sequences and meteorological disaster image datasets;

[0204] The multimodal feature extraction unit 202 is used to perform multimodal feature extraction on the meteorological element feature sequence and the meteorological disaster image dataset to obtain the meteorological element temporal depth features corresponding to the meteorological element feature sequence and the image visual depth features corresponding to the meteorological disaster image dataset.

[0205] The fusion unit 203 is used to fuse the temporal depth features of the meteorological elements and the visual depth features of the image to obtain multimodal integrated features;

[0206] The identification unit 204 is used to identify the multimodal integrated features using a numerically adapted multimodal model and an image-adapted multimodal model respectively, to obtain a first meteorological disaster identification result corresponding to the numerically adapted multimodal model and a second meteorological disaster identification result corresponding to the image-adapted multimodal model;

[0207] The judgment unit 205 is used to perform result fusion judgment on the first meteorological disaster identification result and the second meteorological disaster identification result to obtain a third meteorological disaster identification result;

[0208] The verification unit 206 is used to perform a second verification on the third meteorological disaster identification result to obtain the target meteorological identification result.

[0209] In some embodiments, the determining unit 205 may include:

[0210] The confidence conversion subunit is used to convert the confidence of the first meteorological disaster identification result and the second meteorological disaster identification result respectively, so as to obtain the first confidence distribution corresponding to the first meteorological disaster identification result and the second confidence distribution corresponding to the second meteorological disaster identification result;

[0211] The judgment subunit is used to perform judgment processing on the first confidence distribution and the second confidence distribution based on a preset fusion judgment strategy to obtain the third meteorological disaster identification result.

[0212] In some embodiments, the determining subunit may include:

[0213] The comparison module is used to compare the first confidence distribution and the second confidence distribution with a preset confidence threshold based on the preset fusion judgment strategy.

[0214] The weighted fusion module is used to perform weighted fusion on the first confidence distribution and the second confidence distribution when the first confidence distribution is less than the preset confidence threshold and the second confidence distribution is less than the preset confidence threshold, so as to obtain the third meteorological disaster identification result;

[0215] The update module is used to use the fourth confidence distribution as the third meteorological disaster identification result when there is a fourth confidence distribution in the first confidence distribution or the second confidence distribution that is greater than or equal to the preset confidence threshold.

[0216] In some embodiments, the verification unit 206 may include:

[0217] The first acquisition subunit is used to acquire the disaster feature matching rule base;

[0218] The verification subunit is used to perform matching verification between the third meteorological disaster identification result and the temporal depth features of meteorological elements based on the disaster feature matching rule library, and to perform matching verification between the third meteorological disaster identification result and the visual depth features of the image based on the disaster feature matching rule library;

[0219] An update subunit is used to use the third meteorological disaster identification result as the target meteorological identification result when the matching verification passes.

[0220] In some embodiments, the verification unit 206 may further include:

[0221] The second acquisition subunit is used to acquire the temporal depth features of the target meteorological element and the visual depth features of the target image when the matching verification fails. The temporal depth features of the target meteorological element include the temporal depth features of meteorological elements that do not match the disaster feature matching rule base, and the visual depth features of the target image include the visual depth features of the image that do not match the disaster feature matching rule base.

[0222] The first adjustment subunit is used to adjust the numerical adaptation multimodal model using the temporal depth features of the target meteorological elements and the third meteorological disaster identification results, so as to obtain the adjusted numerical adaptation multimodal model.

[0223] The second adjustment subunit is used to adjust the image-adaptive multimodal model using the visual depth features of the target image and the third meteorological disaster identification result, to obtain the adjusted image-adaptive multimodal model;

[0224] The identification subunit is used to identify the meteorological element feature sequence using the adjusted numerical adaptation multimodal model, and to identify the meteorological disaster image dataset using the adjusted image adaptation multimodal model, so as to obtain the target meteorological identification result.

[0225] In some embodiments, the apparatus proposed in this application may further include:

[0226] The second acquisition unit is used to acquire the numerical adaptive multimodal model to be trained, the image adaptive multimodal model to be trained, meteorological element training data, disaster image training set, disaster labels, and disaster text description labels, wherein the meteorological element training data and the disaster image training set have a temporal correlation relationship;

[0227] The second identification unit is used to identify the meteorological element training data and the disaster image training set using the numerical adaptive multimodal model to be trained, and to obtain the fourth meteorological identification result.

[0228] The third identification unit is used to identify the disaster image training set using the image-adaptive multimodal model to be trained, and obtain the fifth meteorological identification result;

[0229] The first calculation unit is used to calculate the time-series causal inference loss based on the fourth meteorological identification result and the disaster label;

[0230] The second calculation unit is used to calculate the visual semantic understanding loss based on the fifth meteorological identification result and the disaster text description label;

[0231] The loss fusion unit is used to fuse the temporal causal reasoning loss and the visual semantic understanding loss to obtain a fusion loss;

[0232] An adjustment unit is used to adjust the numerical adaptive multimodal model and the image adaptive multimodal model to be trained using the fusion loss, respectively, to obtain the numerical adaptive multimodal model and the image adaptive multimodal model.

[0233] In some embodiments, the first computing unit includes:

[0234] The first calculation subunit is used to calculate the causal order prediction loss based on the fourth meteorological identification result and the disaster label;

[0235] The weight acquisition subunit is used to acquire the attention weights and prior weights of the numerically adapted multimodal model to be trained.

[0236] The second calculation subunit is used to calculate the disaster classification loss based on the attention weight and prior weight;

[0237] The third calculation subunit is used to calculate the first disaster classification cross-entropy loss based on the meteorological element training data and the fourth meteorological identification result;

[0238] The first fusion subunit is used to fuse the causal order prediction loss, the disaster classification loss, and the multi-class cross-entropy loss to obtain the temporal causal inference loss.

[0239] In some embodiments, the second computing unit includes:

[0240] The fourth calculation subunit is used to calculate the image-text contrast loss based on the fifth meteorological identification result and the disaster text description label;

[0241] The fifth calculation subunit is used to extract disaster text description information from the fifth meteorological identification result, and calculate the image description generation loss based on the fifth meteorological identification result and the disaster text description information;

[0242] The sixth calculation subunit is used to calculate the second disaster classification cross-entropy loss based on the disaster image training set and the fifth meteorological recognition result;

[0243] The seventh calculation subunit is used to obtain the positive sample training set of the disaster image training set, and calculate the visual feature discrimination loss based on the positive sample training set and the fifth meteorological recognition result;

[0244] The second fusion subunit is used to fuse the image-text contrast loss, the image description generation loss, the second disaster classification cross-entropy loss, and the visual feature discrimination loss to obtain the visual semantic understanding loss.

[0245] In some embodiments, the loss fusion unit includes:

[0246] The eighth calculation subunit is used to calculate the cross-modal contrast loss based on the fourth meteorological identification result and the fifth meteorological identification result;

[0247] The ninth calculation subunit is used to calculate the mutual information maximization loss based on the fourth meteorological identification result and the fifth meteorological identification result;

[0248] The third fusion subunit is used to perform weighted fusion of the cross-modal contrast loss, the mutual information maximization loss, the temporal causal reasoning loss, and the visual semantic understanding loss to obtain the fusion loss.

[0249] The aforementioned meteorological disaster identification device based on a multimodal large model can improve the accuracy of meteorological disaster identification based on a multimodal large model.

[0250] This application also provides a computer device, which may include a terminal or a server. For example, the computer device may serve as a meteorological disaster identification terminal based on a multimodal large model, and the terminal may be a mobile phone, tablet computer, etc.; or the computer device may serve as a server, such as a meteorological disaster identification server based on a multimodal large model. Figure 4 As shown, it illustrates the structural diagram of the terminal involved in the embodiments of this application, specifically:

[0251] The computer device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 4The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0252] The processor 401 is the control center of the computer device, connecting various parts of the computer device through various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user page, and application programs, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.

[0253] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0254] The computer device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0255] The computer device may also include an input unit 404, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0256] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the computer device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:

[0257] Obtain meteorological element feature sequences and meteorological disaster image datasets;

[0258] Multimodal feature extraction is performed on the meteorological element feature sequence and the meteorological disaster image dataset to obtain the meteorological element temporal depth features corresponding to the meteorological element feature sequence and the image visual depth features corresponding to the meteorological disaster image dataset;

[0259] The temporal depth features of the meteorological elements and the visual depth features of the image are fused to obtain multimodal comprehensive features;

[0260] The multimodal integrated features are identified using a numerically adapted multimodal model and an image-adaptive multimodal model, respectively, to obtain the first meteorological disaster identification result corresponding to the numerically adapted multimodal model and the second meteorological disaster identification result corresponding to the image-adaptive multimodal model.

[0261] The first meteorological disaster identification result and the second meteorological disaster identification result are fused and judged to obtain the third meteorological disaster identification result;

[0262] The third meteorological disaster identification result is verified a second time to obtain the target meteorological identification result.

[0263] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0264] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the above embodiments.

[0265] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0266] Therefore, embodiments of this application also provide a storage medium storing a computer program that can be loaded by a processor to execute the steps in any of the meteorological disaster identification methods based on multimodal large models provided in embodiments of this application. For example, the computer program can execute the following steps:

[0267] Obtain meteorological element feature sequences and meteorological disaster image datasets;

[0268] Multimodal feature extraction is performed on the meteorological element feature sequence and the meteorological disaster image dataset to obtain the meteorological element temporal depth features corresponding to the meteorological element feature sequence and the image visual depth features corresponding to the meteorological disaster image dataset;

[0269] The temporal depth features of the meteorological elements and the visual depth features of the image are fused to obtain multimodal comprehensive features;

[0270] The multimodal integrated features are identified using a numerically adapted multimodal model and an image-adaptive multimodal model, respectively, to obtain the first meteorological disaster identification result corresponding to the numerically adapted multimodal model and the second meteorological disaster identification result corresponding to the image-adaptive multimodal model.

[0271] The first meteorological disaster identification result and the second meteorological disaster identification result are fused and judged to obtain the third meteorological disaster identification result;

[0272] The third meteorological disaster identification result is verified a second time to obtain the target meteorological identification result.

[0273] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0274] Since the computer program stored in the storage medium can execute the steps in any of the meteorological disaster identification methods based on multimodal large models provided in the embodiments of this application, it can achieve the beneficial effects that any of the meteorological disaster identification methods based on multimodal large models provided in the embodiments of this application can achieve, as detailed in the preceding embodiments, and will not be repeated here.

[0275] The foregoing has provided a detailed description of a meteorological disaster identification method, apparatus, computer equipment, and storage medium based on a multimodal large model provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A meteorological disaster identification method based on a multimodal large model, characterized in that, include: Obtain meteorological element feature sequences and meteorological disaster image datasets; Multimodal feature extraction is performed on the meteorological element feature sequence and the meteorological disaster image dataset to obtain the meteorological element temporal depth features corresponding to the meteorological element feature sequence and the image visual depth features corresponding to the meteorological disaster image dataset; The temporal depth features of the meteorological elements and the visual depth features of the image are fused to obtain multimodal comprehensive features; The multimodal integrated features are identified using a numerically adapted multimodal model and an image-adaptive multimodal model, respectively, to obtain the first meteorological disaster identification result corresponding to the numerically adapted multimodal model and the second meteorological disaster identification result corresponding to the image-adaptive multimodal model. The first meteorological disaster identification result and the second meteorological disaster identification result are fused and judged to obtain the third meteorological disaster identification result; The third meteorological disaster identification result is verified a second time to obtain the target meteorological identification result.

2. The method according to claim 1, characterized in that, The step of fusing the first meteorological disaster identification result and the second meteorological disaster identification result to obtain the third meteorological disaster identification result includes: The confidence levels of the first meteorological disaster identification result and the second meteorological disaster identification result are transformed to obtain the first confidence level distribution corresponding to the first meteorological disaster identification result and the second confidence level distribution corresponding to the second meteorological disaster identification result. Based on a preset fusion judgment strategy, the first confidence distribution and the second confidence distribution are judged and processed to obtain the third meteorological disaster identification result.

3. The method according to claim 2, characterized in that, The process of judging the first confidence distribution and the second confidence distribution based on a preset fusion judgment strategy to obtain the third meteorological disaster identification result includes: Based on the preset fusion judgment strategy, the first confidence distribution and the second confidence distribution are compared with the preset confidence threshold respectively; When the first confidence distribution is less than the preset confidence threshold and the second confidence distribution is less than the preset confidence threshold, the first confidence distribution and the second confidence distribution are weighted and fused to obtain the third meteorological disaster identification result; When there is a fourth confidence distribution in the first confidence distribution or the second confidence distribution that is greater than or equal to the preset confidence threshold, the fourth confidence distribution is used as the third meteorological disaster identification result.

4. The method according to claim 1, characterized in that, The secondary verification of the third meteorological disaster identification result to obtain the target meteorological identification result includes: Obtain a disaster feature matching rule base; Based on the disaster feature matching rule base, the third meteorological disaster identification result is matched and verified with the temporal depth features of meteorological elements, and based on the disaster feature matching rule base, the third meteorological disaster identification result is matched and verified with the visual depth features of the image; When the matching verification passes, the third meteorological disaster identification result is used as the target meteorological identification result.

5. The method according to claim 4, characterized in that, The method further includes: When the matching verification fails, the temporal depth features of the target meteorological elements and the visual depth features of the target image are obtained. The temporal depth features of the target meteorological elements include the temporal depth features of meteorological elements that do not match the disaster feature matching rule base, and the visual depth features of the target image include the visual depth features of the image that do not match the disaster feature matching rule base. The numerical adaptation multimodal model is adjusted using the temporal depth features of the target meteorological elements and the third meteorological disaster identification results to obtain the adjusted numerical adaptation multimodal model. The image-adaptive multimodal model is adjusted using the visual depth features of the target image and the third meteorological disaster identification result to obtain the adjusted image-adaptive multimodal model; The adjusted numerical adaptive multimodal model is used to identify the meteorological element feature sequence, and the adjusted image adaptive multimodal model is used to identify the meteorological disaster image dataset to obtain the target meteorological identification result.

6. The method according to claims 1 to 5, characterized in that, Before using a numerically adapted multimodal model and an image-adapted multimodal model to identify the multimodal integrated features and obtain the first meteorological disaster identification result corresponding to the numerically adapted multimodal model and the second meteorological disaster identification result corresponding to the image-adapted multimodal model, the method includes: The following are obtained: a numerical adaptive multimodal model to be trained, an image adaptive multimodal model to be trained, meteorological element training data, a disaster image training set, disaster labels, and disaster text description labels, wherein the meteorological element training data and the disaster image training set have a temporal correlation relationship; The numerical adaptive multimodal model to be trained is used to identify the meteorological element training data and the disaster image training set to obtain a fourth meteorological identification result; The image-adaptive multimodal model to be trained is used to identify the disaster image training set to obtain the fifth meteorological identification result; Based on the fourth meteorological identification result and the disaster label, calculate the time-series causal inference loss; Based on the fifth meteorological identification result and the disaster text description label, calculate the visual semantic understanding loss; The temporal causal reasoning loss and the visual semantic understanding loss are fused to obtain the fusion loss; The fusion loss is used to adjust the numerical adaptive multimodal model and the image adaptive multimodal model to be trained, respectively, to obtain the numerical adaptive multimodal model and the image adaptive multimodal model.

7. The method according to claim 6, characterized in that, The calculation of time-series causal inference loss based on the fourth meteorological identification result and the disaster label includes: Based on the fourth meteorological identification result and the disaster label, calculate the causal order to predict the loss; Obtain the attention weights and prior weights of the numerically adapted multimodal model to be trained; Based on the attention weights and prior weights, calculate the disaster classification loss; Based on the meteorological element training data and the fourth meteorological identification result, calculate the first disaster classification cross-entropy loss; The temporal causal inference loss is obtained by fusing the causal order prediction loss, the disaster classification loss, and the multi-class cross-entropy loss.

8. The method according to claim 6, characterized in that, The calculation of visual semantic understanding loss based on the fifth meteorological identification result and the disaster text description label includes: Based on the fifth meteorological identification result and the disaster text description label, calculate the image text contrast loss; Extract disaster text description information from the fifth meteorological identification result, and calculate image description generation loss based on the fifth meteorological identification result and the disaster text description information; Based on the disaster image training set and the fifth meteorological recognition result, the second disaster classification cross-entropy loss is calculated; Obtain a positive sample training set of the disaster image training set, and calculate the visual feature discrimination loss based on the positive sample training set and the fifth meteorological recognition result; The image-text contrast loss, the image description generation loss, the second disaster classification cross-entropy loss, and the visual feature discrimination loss are fused together to obtain the visual semantic understanding loss.

9. The method according to any one of claims 7 or 8, characterized in that, The step of fusing the temporal causal reasoning loss and the visual semantic understanding loss to obtain a fusion loss includes: Based on the fourth and fifth meteorological identification results, calculate the cross-modal contrast loss; Based on the fourth and fifth meteorological identification results, calculate the mutual information maximization loss; The cross-modal contrast loss, the mutual information maximization loss, the temporal causal reasoning loss, and the visual semantic understanding loss are weighted and fused to obtain the fusion loss.

10. A meteorological disaster identification device based on a multimodal large model, characterized in that, include: The acquisition unit is used to acquire meteorological element feature sequences and meteorological disaster image datasets; A multimodal feature extraction unit is used to extract multimodal features from the meteorological element feature sequence and the meteorological disaster image dataset to obtain the meteorological element temporal depth features corresponding to the meteorological element feature sequence and the image visual depth features corresponding to the meteorological disaster image dataset. The fusion unit is used to fuse the temporal depth features of the meteorological elements and the visual depth features of the image to obtain multimodal integrated features; The identification unit is used to identify the multimodal integrated features using a numerically adapted multimodal model and an image-adapted multimodal model respectively, to obtain a first meteorological disaster identification result corresponding to the numerically adapted multimodal model and a second meteorological disaster identification result corresponding to the image-adapted multimodal model; The judgment unit is used to perform result fusion judgment on the first meteorological disaster identification result and the second meteorological disaster identification result to obtain the third meteorological disaster identification result; The verification unit is used to perform a secondary verification on the third meteorological disaster identification result to obtain the target meteorological identification result.