Multi-modal Spatiotemporal Compensation Method and Device for Advection Fog Warning
By introducing multimodal space-time compensation technology into the marine mass fog early warning method, using parallel neural networks and cross-modal information interaction, the problem of difficulty in achieving accurate early warning in single modal data is solved, and the forecast accuracy and security are improved.
Patent Information
- Application Number
- CN202310088520.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-09
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-02-09
AI Technical Summary
The existing marine fog warning methods mainly rely on single modal data, making it difficult to effectively utilize the information in multimodal data, resulting in low prediction accuracy and difficult to achieve real-time accurate early warning.
A multimodal space-time compensation cluster fog early warning method is designed. By building a parallel neural network, combining time-domain nonlinear encoding of meteorological elements and airspace convolution of ocean images, the encoding and fusion of multimodal data is realized. Using cross-modal information interaction and modal interleaving guidance classifier based on semantic matching, we learn the shared features of cross-modal mass fog to improve prediction accuracy.
By extracting multimodal characteristics of mass fog with high resolution, the difficulty of discrimination is reduced, the forecasting accuracy is improved, and more accurate marine mass fog warning is achieved, ensuring the safety and efficiency of maritime operations and traffic.
Smart Images

Figure CN116299773B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of marine patchy fog warning, and particularly to a patchy fog warning method and device with multi-modal spatio-temporal compensation. Background Art
[0002] Marine patchy fog is a marine atmospheric phenomenon with relatively complex formation causes. Its formation, development and dissipation conform to the laws of atmospheric dynamics and microphysics, and are closely related to meteorological elements such as atmospheric liquid water concentration, air flow movement, and temperature change. It has characteristics such as rapid occurrence, strong regionality, and difficulty in prediction and forecasting. [1] As a kind of fog, when marine patchy fog occurs, it can reduce the visibility of the sea area to less than 1 km, seriously affecting activities such as offshore operations, maritime military activities, navigation, fishery, and nearshore engineering, and posing potential safety hazards to ship navigation. The problem of marine patchy fog prediction has received extensive attention from the academic community and related industries. Improving the accuracy of marine patchy fog prediction can not only avoid economic losses caused by sudden encounters with patchy fog during marine operations, but also play a role in ensuring the safety of people's lives and property.
[0003] The forecasting methods of marine patchy fog mainly include two categories: traditional numerical model forecasting and neural network-based methods. Numerical model forecasting relies on technical means such as sensors and satellites for monitoring, and depends on expert knowledge to establish a numerical meteorological model. The actual data collected from the atmosphere by meteorological detection instruments is used as the initial atmospheric field. [2] Under certain initial value and boundary value conditions, the equations are solved to predict the atmospheric motion state and weather phenomena in a future period.
[0004] In recent years, machine learning has been widely applied in the field of environmental science. Researchers have tried to apply machine learning methods to the prediction of marine patchy fog. For example, based on the hourly meteorological observation data of Xilian Island Station in Lianyungang from 2014 to 2018. [3] Combined with the statistical analysis of sea fog events and meteorological element characteristics, a method for establishing a meteorological element prediction model based on the machine learning C4.5 algorithm is proposed; the hourly meteorological element observation data is converted into time series data, and the long short-term memory network (LSTM) is used to encode the sequence data, and the relationship between the meteorological element time series data is obtained through supervised training. [4] .
[0005] The fog prediction method based on traditional numerical models relies on expert knowledge to establish a numerical meteorological model. The calculation process is complex and it is difficult to implement accurate real-time prediction. The method based on neural networks can quickly learn the occurrence pattern of fog and issue early warnings through iterative training on historical meteorological data. However, most methods only focus on single-modal data, ignoring the potential contribution of the combination of multi-modal data to the prediction of atmospheric visibility, and rarely explore the rich spatio-temporal connotations of multi-modal features and the significance of multi-level and multi-stage feature interactions. To make up for the above deficiencies, the present invention proposes a fog warning method with multi-modal spatio-temporal compensation. Summary of the Invention
[0006] The present invention provides a fog warning method and device with multi-modal spatio-temporal compensation. The present invention comprehensively analyzes historical monitored meteorological data and ocean images, and can predict the probability of fog occurrence in a specific future time in the sea area. When the predicted probability is greater than the threshold, a warning will be issued, and the operators can take corresponding measures in time to reduce the harm caused by the sudden drop in visibility. See the following description for details:
[0007] A fog warning method with multi-modal spatio-temporal compensation, the method includes:
[0008] Construct a parallel neural network including a time-domain non-linear encoding part of meteorological elements and a spatial-domain convolution part of ocean images. The former includes 4 fully connected layers and activation functions, and the latter includes 4 serial convolution blocks. Based on the parallel network, multi-modal data are respectively encoded. The meteorological element observation data are input into the time-domain non-linear encoding part of meteorological elements to generate data encoding, and the satellite remote sensing is input into the spatial-domain convolution part of ocean images to generate image encoding;
[0009] Calculate the relative increment of the data encoding and the image encoding at adjacent times as the change rate, and take the time with a large change rate as the fog formation time;
[0010] Use the cross-modal information interaction part based on semantic matching to fuse the image encoding containing the visual space information of fog with the data encoding containing the time-series information of meteorological elements to achieve spatio-temporal feature compensation;
[0011] Construct a modality interleaving guided classifier to perform cross-modal information interaction at the classifier level, align the predicted probability distributions of the same category objects under different modalities, and learn the common features of cross-modal fog;
[0012] According to the probability distribution output by the modality interleaving guided classifier, if the predicted fog probability is greater than the threshold, it is determined that fog will occur in the sea area after a specific time, and a warning signal is issued.
[0013] A multi-modal spatio-temporal compensation fog warning device, the device comprising: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method steps described in any one of the claims.
[0014] The beneficial effects of the technical solution provided by the present invention are as follows:
[0015] 1. According to the self-data attributes of different modalities, the present invention designs a network structure in which the spatial convolution part of the ocean image and the time-domain non-linear coding part of the meteorological elements are parallel. The former processes the ocean image and extracts visual features such as the color, brightness, and texture of the fog; the latter processes the meteorological element grid data and explores the complex non-linear time-series relationship between the meteorological elements and the formation of the fog. Both contain spatial and time information, meeting the need for further spatio-temporal feature compensation of the fog.
[0016] 2. According to the characteristics of the fast formation speed of the fog, the present invention proposes a relative increment feedback mechanism, calculates the relative increment between adjacent moments as the change rate, and the moment with a large change rate is more likely to be the boundary of the fog formation. At this time, the features are of great value for fog warning, and the calculated relative increment will be fed back to the weights of the multi-modal feature interaction and fusion.
[0017] 3. Considering the attribute differences between different modalities, the present invention designs a cross-modal information interaction part based on semantic matching, controls the fusion weights through relative increments, and implements hierarchical multi-stage fusion of fog image features and meteorological element time-series coding to achieve dynamic adaptive interaction. The different modality branches obtain information compensation from another modality, expanding the information dimension and enhancing the discriminability of the fog features.
[0018] 4. The present invention proposes a modality interleaving guided classifier. Based on the feature embedding space generated by the feature extractor, the image modality coding and the grid data modality coding are used as the parameters of the modality interleaving guided classifier. By reducing the difference between the probability distribution of the meteorological element coding under the guidance of the image modality and the probability distribution of the image features under the guidance of the grid modality, the cross-modal common fog features are learned.
[0019] The fog warning method with multi-modal spatio-temporal compensation proposed by the present invention extracts high-resolution multi-modal fog features, reduces the discrimination difficulty, improves the prediction accuracy, and ensures the efficiency and safety of offshore operations and transportation through timely warning intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flowchart of a fog warning method with multi-modal spatio-temporal compensation;
[0021] Figure 2 is a network structure diagram used in a fog warning method with multi-modal spatio-temporal compensation;
[0022] Figure 3 Schematic diagram for spatio-temporal compensation of multi-modal features for the cross-modal information interaction part. Specific implementation manner
[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes in detail the implementation manners of the present invention.
[0024] In order to make full use of the value of different modal data for predicting marine radiation fog, the embodiments of the present invention introduce a parallel network in the network structure, including: a meteorological element time-domain non-linear encoding part and a marine image spatial-domain convolution part. Different parts correspond to different modal data inputs. On this basis, considering that there is not only intra-modal correlation but also inter-modal information compensation between the two inputs, the embodiments of the present invention design a cross-modal information interaction part based on semantic matching on the basis of the parallel network to achieve spatio-temporal compensation of different modal features; a classifier is promoted by modal interleaving to facilitate the learning of fog-invariant features in cross-modal information. Due to the inter-modal information interaction of the compensated features, the feature hierarchy is richer and the resolution is higher, improving the accuracy of marine radiation fog prediction.
[0025] Embodiment 1
[0026] A radiation fog warning method with multi-modal spatio-temporal compensation, see Figure 1 , the method includes the following steps:
[0027] 101: Obtain the time series of meteorological elements and satellite remote sensing images in the forecast area, and correspond the two modal data through time information to construct a multi-modal original data set (G, I), where G represents the time series of meteorological elements organized in grid data form, and I represents the satellite remote sensing image;
[0028] 102: Normalize the grid data and the satellite remote sensing image to solve the problem of inconsistent dimensions of different modal data;
[0029] As different modal data, the grid data and the satellite remote sensing image have their own physical ranges. At the same time, the differences in the actual meanings of meteorological elements within the grid data modal also lead to differences in the value ranges. For example: the range of the RGB image is 0-255, and the range of the grid data is related to the specificity of the meteorological acquisition equipment and the actual meaning of the meteorological data. The normalization operation maps the numerical range of the data to the same range to avoid the adverse induction of inconsistent data dimensions on the neural network.
[0030] 103: Construct a parallel neural network that includes a time-domain non-linear encoding part for meteorological elements and a spatial-domain convolution part for ocean images. The former includes 4 fully-connected layers and activation functions, and the latter includes 4 serial convolution blocks. Encode multi-modal data based on the parallel network. Input meteorological element observation data into the time-domain non-linear encoding part of meteorological elements to generate data encoding, and input satellite remote sensing into the spatial-domain convolution part of ocean images to generate image encoding;
[0031] In the prior art, single-modal data is mostly used for predicting advection fog. However, the information dimension covered by single-modal data is limited. In the embodiments of the present invention, image-modal data is introduced to make up for it. Images are a data modality that can independently judge atmospheric visibility and can also be regarded as the representation of various meteorological elements acting on the visual space at the same time. The information of the two modalities is both independent and complementary to each other.
[0032] 104: Calculate the relative increment of data encoding and image encoding at adjacent times as the change rate. The moment with a large change rate tends to be the boundary of advection fog formation. At this time, the features reflect the development law of advection fog and have important value for advection fog early warning. The calculated relative increment will be fed back to the weights of multi-modal feature compensation;
[0033] 105: Use the cross-modal information interaction part based on semantic matching to fuse the image encoding containing the visual space information of advection fog and the data encoding containing the time-series information of meteorological elements to achieve spatio-temporal feature compensation;
[0034] In the embodiments of the present invention, the image encoding extracted by the spatial-domain convolution operation of ocean images generates compensation features through semantic matching to realize the transfer of spatial information to the time-domain non-linear encoding part. The relative increment calculated in step 103 participates in the weight ratio distribution of the transfer process. The larger the relative increment, the stronger the advection fog semantics in the image features, and the greater the weight of the image features. The fed-back relative increment changes with time, dynamically controlling the weights of image features participating in the fusion at different times, and realizing the self-adaptation of the spatio-temporal feature interaction operation to the development process of advection fog.
[0035] Integrate the time-series features of meteorological elements extracted by the time-domain non-linear encoding operation into the spatial-domain convolution part of ocean images in stages. The change rate of meteorological elements calculated in step 103 participates in the multi-modal spatio-temporal feature fusion. The moment with a large change rate of meteorological elements reveals the meteorological law of the development and change of advection fog, and a larger weight is assigned to guide the neural network to learn. Solve the problem that the feature fusion operation in the prior art is too simple and the multi-modal and multi-level features are not fully utilized, and give play to the advantages of realizing spatio-temporal information compensation in multi-modal.
[0036] 106: Construct a modality-interleaved guiding classifier to perform cross-modal information interaction at the classifier level, and learn the common features of cross-modal advection fog by aligning the predicted probability distributions of the same category objects in different modalities.
[0037] 107: According to the probability distribution output by the modal interleaving guidance classifier, if the predicted probability of advection fog is greater than the threshold, it is determined that advection fog will occur after a specific time, and a warning signal is issued so that relevant measures can be taken in a timely manner. If the predicted probability of advection fog is lower than the threshold, no warning is given.
[0038] In summary, through the above steps 101-107, the embodiment of the present invention designs a parallel network structure adapted to multi-modal data and a novel cross-modal information interaction operation, enhancing the discriminability of features and improving the accuracy of advection fog prediction.
[0039] Embodiment 2
[0040] The following further introduces the solution in Embodiment 1 in combination with specific examples and calculation formulas, as detailed in the following description:
[0041] 201: Obtain the time series of meteorological elements and satellite remote sensing images in the forecast area, and correspond the two types of modal data through time information to construct a multi-modal original data set (G, I). G represents the time series of meteorological elements organized in grid data form, and I represents the satellite remote sensing image;
[0042] 202: Normalize the grid data and satellite remote sensing images to solve the problem of inconsistent dimensions of different modal data;
[0043] Due to different acquisition devices and their own attributes, the sizes and distribution ranges of different modal data are inconsistent. If no preprocessing is performed, the network will tend to be biased towards modal data with larger absolute values. Therefore, it is necessary to convert data of different specifications into the same specification to solve the problem of inconsistent dimensions. Specifically, the embodiment of the present invention adopts the Z-Score normalization method, based on the mean μ and variance σ 2 to normalize the time series data. Specifically:
[0044]
[0045] where x represents the time series data, and ∈ is a constant close to zero set to maintain numerical stability, represents the normalized time series data.
[0046] 203: Construct a parallel neural network including a time-domain non-linear encoding part of meteorological elements and a spatial-domain convolution part of ocean images. Based on the parallel network, encode the multi-modal data respectively. Input the meteorological element observation data into the time-domain non-linear encoding part of meteorological elements to generate data encoding, and input the satellite remote sensing image into the spatial-domain convolution part of ocean images to generate image encoding;
[0047] When designing the specific details of the network, such as the number of neurons and the size of the convolutional kernel, it is only necessary to ensure that the number of neurons in the fully connected layer at the corresponding stage of the parallel network is the same as the number of output feature channels of the convolutional block. One case of this embodiment will be described below. Build a time-domain non-linear encoding part of meteorological elements including 4 fully connected layers and activation functions. Each fully connected layer has 64, 128, 256, and 512 neurons. The time-domain non-linear encoding network of meteorological elements processes meteorological element grid data. Build a spatial-domain convolution part of ocean images containing 4 convolutional blocks. The number of output feature channels in the 4 stages is the same as the number of neurons in the corresponding stage of the time-domain non-linear encoding. The spatial-domain convolution network of ocean images extracts visual features such as the color, brightness, and texture of the fog patches.
[0048] Due to the different attributes of the input data, the features extracted by the time-domain non-linear encoding part of meteorological elements and the spatial-domain convolution part of ocean images respectively contain information of their respective data modalities. The time-domain non-linear encoding part of meteorological elements encodes the time series of meteorological elements. Through training, it can learn the temporal variation law of meteorological elements related to the formation of fog patches, which belongs to the information in the time dimension. For example, the temporal variation law of relative humidity before the occurrence of fog patches; the spatial-domain convolution part of ocean images extracts features from the real-time sea images. Through training, it can learn the visual spatial features of fog patches, which belongs to the information in the spatial dimension. For example, the visual features such as color, brightness, and texture corresponding to different visibility levels.
[0049] 204: Calculate the relative increment of the encoded meteorological element data and the image encoding at adjacent times as the change rate. The moment with a large change rate tends to be the boundary moment of the formation of fog patches. At this time, the features reflect the development law of fog patches, which is of great value for fog patch warning. The calculated relative increment will be fed back to the weights of the multi-modal feature fusion;
[0050] In the spatial-domain convolution part of ocean images, according to the characteristics that the generation and dissipation of fog patches are relatively fast, the ocean image features are separated based on the spectral characteristics, and irrelevant features are filtered in the frequency domain to couple the features at critical moments of fog patches. Specifically, the boundary features at the moment of fog patch formation are embedded into the spectral peak space through the spectral mapping function, and the intermediate features in the continuous state of fog patches are embedded into the stationary space.
[0051]
[0052] Among them, respectively represent the ocean image features corresponding to the current and the previous moments, are the high-frequency mapping function and the low-frequency mapping function respectively. After taking the difference between adjacent image features, the high-frequency and low-frequency features are separated to obtain the boundary features and intermediate features of fog patches respectively. After coupling in the frequency domain, the increment matrix D is obtained through the inverse mapping Φ t, the dimension of the incremental matrix is the same as that of the image features. After obtaining the feature incremental matrix, the relative increment α is further calculated as follows:
[0053]
[0054] Here The larger the logarithmic part, the more obvious the change in the image features, the greater the possibility that the current moment is when the advection fog is forming, and the greater the contribution of the corresponding image features to the advection fog warning.
[0055] In the time-domain non-linear coding part, the differential of the meteorological element data coding with respect to time is calculated to obtain the time change rate β of the meteorological elements, and the calculation is as follows:
[0056]
[0057] represents the coding of the i-th type of meteorological element sequence data, such as relative humidity, K represents the number of meteorological element categories participating in the prediction, t represents the corresponding time series. The larger the change rate, the greater the change amplitude of the meteorological elements in a short time, the greater the possibility that the current moment is the boundary of the advection fog formation, and the greater the contribution of the corresponding meteorological element coding to the advection fog warning.
[0058] 205: Use the cross-modal information interaction part based on semantic matching to fuse the image coding containing the visual space information of the advection fog with the data coding containing the time series information of the meteorological elements to achieve spatio-temporal feature compensation;
[0059] The function of the cross-modal information interaction part is to achieve two-way interaction and information transfer of features, including the transfer of time series information from the grid data modality to the image modality and the transfer of spatial visual information from the image modality to the grid data modality. Due to the large differences in the private attributes of the modalities, there is a gap. Therefore, the cross-modal information interaction part needs to consider the semantic matching relationship between multiple modalities.
[0060] Specifically, the ocean image spatial domain convolution part embeds the image features into a high-dimensional space. The high-dimensional feature channel domain contains rich semantics, and different channels represent abstractions from different perspectives. The image features are segmented in the channel domain to obtain That is: The obtained feature segments are respectively used to measure the semantic distance from the meteorological element coding, and the generation of compensated features between modalities is guided by the semantic matching matrix. Specifically:
[0061]
[0062] Among them, is the p-th segmented segment of the image feature, F T represents the meteorological element coding, ⊙ represents the matrix Hadamard product operation, respectively represent F T The mean value. Formula 5 calculates the semantic matching coefficient between the p-th segmentation segment of the image feature and the meteorological element encoding.
[0063] Furthermore, based on the semantic matching coefficient, the feature fusion between modalities is controlled. The meteorological element encoding absorbs more image features with high correlation and high matching degree as cross-modal information compensation, and discards or does not consider cross-modal information with a large semantic gap, obtaining the compensated meteorological element encoding.
[0064]
[0065] Among them, represents the semantic matching vector between the image feature obtained by vertical splicing and the meteorological element encoding, p represents the image coding weight conversion operation, which fuses the semantic matching vector as the weight with the image coding to generate the image modality compensation feature. represents the feature map output by the marine image spatial domain convolutional network at the i-th stage. In order to suppress the influence of irrelevant information in the low-level visual features, the output of the shallow network is not dimensionally matched, so i takes 2 and 3.
[0066] Through the cross-modal information interaction part, a single branch of the parallel network realizes the perception of the information flow of the other branch. At the same time, since the information compensation is obtained based on semantic matching and will not bring noise interference to its own information flow, the fog space feature of the image modality and the temporal feature of the grid data modality are better fused.
[0067] 206: Construct a modality interleaving guided classifier to perform cross-modal information interaction at the classifier level, and learn the common features of cross-modal fog by aligning the predicted probability distributions of the same category objects under different modalities.
[0068] Specifically, the modality interleaving guided classifier consists of a meteorological element encoding classification part and an image encoding classification part. The meteorological element encoding classification part of the modality interleaving guided classifier is controlled by using the image modality feature, and the parameters of the image encoding classification part are replaced by randomly permuting the meteorological data encoding, and the modality-independent fog invariant features are explored by reducing the difference in the feature distributions of different modalities.
[0069]
[0070] Among them, represents the probability distribution of the meteorological element encoding under the guidance of using the image modality as a metric, represents the probability distribution of the image encoding under the guidance of using the grid data modality as a metric, f g and f iFeature extractors representing the grid data modality and the image modality respectively, where I and G represent the inputs of the image modality and the grid data modality. R represents a random permutation operation, which aims to break the arrangement order of meteorological elements organized by humans, so that the embedding space of the grid data modality also changes, enhancing the diversity of the cross-modal guidance method.
[0071] The cross-modal guidance classifier is different from ordinary classifiers and needs to be constrained to learn stable cross-modal features of advection fog. By reducing the distance of the output probability distributions of the cross-modal guidance classifier, the predicted probability distributions of the same
[0072] category objects in different modalities are aligned to learn the common features of cross-modal advection fog, specifically:
[0073]
[0074] Here, L div represents the probability distribution difference, represents the set of different cross-modal ways considering multi-modal and within-modal permutation methods, represents the number of elements in the set, represents the predicted probability distributions of the same category objects under different cross-modal guidance methods.
[0075] The cross-modal guidance classifier is a classifier constructed based on the feature embedding space generated by the feature extractor, using the image modality encoding and the grid data modality encoding as parameters, constraining the probability distribution difference, and guiding the learning of cross-modal advection fog invariant features through cross-modal information.
[0076] 207: According to the probability distribution predicted by the cross-modal guidance classifier, if the predicted probability of advection fog occurrence is greater than the threshold, it is determined that advection fog will occur after a specific time, and a warning signal is issued for timely taking relevant measures. If the predicted probability of advection fog is lower than the threshold, no warning is given.
[0077] The advection fog prediction method based on multi-modal spatio-temporal compensation proposed in the embodiments of the present invention makes full use of the rich information contained in the meteorological element time series data modality and the image modality, breaks through the limitation that the traditional numerical model prediction method can only use numerical modality data for calculation, and solves the problem that it is difficult to explain complex meteorological laws with single-modal data.
[0078] Only a few of the existing methods adopt multi-modal data input, and only simple fusion is carried out at the final stage, without making full use of the rich multi-level spatio-temporal characteristics of advection fog. According to the specific data attributes of different modalities, the embodiments of the present invention design an ocean image spatial domain convolution part and a meteorological element time domain non-linear coding part. The former processes ocean images and extracts visual features such as the color, brightness, and texture of advection fog, and the latter processes meteorological element grid data to explore the complex non-linear temporal relationship between meteorological elements and the formation of advection fog.
[0079] Furthermore, based on semantic matching, multi-modal multi-stage spatio-temporal feature interaction is implemented through a cross-modal information interaction part. The spatial features of advection fog extracted by the ocean image spatial domain convolution part are hierarchically integrated into the meteorological element time domain non-linear coding, and the temporal features extracted by the meteorological element time domain non-linear coding part are phased into the ocean image spatial domain convolution operation. With the help of the feedback relative increment, the multi-modal feature fusion weight is controlled to achieve dynamic adaptive cross-modal information interaction. The information compensation after semantic matching expands the information dimension, enhances the discriminability of advection fog features, and improves the accuracy of ocean advection fog prediction.
[0080] Embodiment 3
[0081] An advection fog warning device with multi-modal spatio-temporal compensation, the device includes: a processor and a memory, and program instructions are stored in the memory. The processor calls the program instructions stored in the memory to make the device execute the following method steps:
[0082] Construct a parallel neural network including a meteorological element time domain non-linear coding part and an ocean image spatial domain convolution part. The former includes 4 fully connected layers and activation functions, and the latter includes 4 serial convolution blocks. Based on the parallel network, multi-modal data are respectively encoded. Meteorological element observation data are input into the meteorological element time domain non-linear coding part to generate data coding, and satellite remote sensing is input into the ocean image spatial domain convolution part to generate image coding;
[0083] Calculate the relative increment of data coding and image coding at adjacent times as the change rate, and take the time with a large change rate as the advection fog formation time;
[0084] Use the cross-modal information interaction part based on semantic matching to fuse the image coding containing the visual spatial information of advection fog and the data coding containing the temporal information of meteorological elements to achieve spatio-temporal feature compensation;
[0085] Construct a modality interleaved guiding classifier to perform cross-modal information interaction at the classifier level, align the prediction probability distributions of the same category objects under different modalities, and learn the common features of cross-modal advection fog;
[0086] According to the probability distribution output by the modal staggered guidance classifier, if the predicted probability of advection fog is greater than the threshold, it is determined that advection fog will occur after a specific time, and a warning signal is issued.
[0087] The relative increment of the image encoding is:
[0088] The boundary features at the moment of advection fog formation are embedded into the spectral peak space through the spectral mapping function, and the intermediate features during the continuous process of advection fog are embedded into the stationary space;
[0089]
[0090] where respectively represent the ocean image features corresponding to the current and the previous moments, are the high-frequency mapping function and the low-frequency mapping function respectively. After taking the difference between adjacent image features for high-low frequency feature separation, the advection fog boundary features and intermediate features are obtained respectively. After coupling in the frequency domain, the increment matrix D is obtained through the inverse mapping Φ t , and the dimension of the increment matrix is the same as that of the image features; after obtaining the feature increment matrix, calculate the relative increment α:
[0091]
[0092] where the spatio-temporal feature compensation is:
[0093] The image features are segmented in the channel domain to obtain that is: The obtained feature segments are respectively measured for semantic distance with the meteorological element encoding, and the generation of inter-modal compensation features is guided by the semantic matching vector:
[0094]
[0095] where is the p-th segmented segment of the image feature, F T represents the meteorological element encoding, ⊙ represents the matrix Hadamard product operation, respectively represent F T mean value;
[0096] The compensated meteorological element encoding
[0097]
[0098] where represents D p The semantic matching vector of the image features obtained by vertical splicing and the meteorological element encoding, represents the image encoding weight conversion operation, and the semantic matching vector is used as the weight to fuse with the image encoding to generate the image modal compensation feature; Denote the feature map output by the marine image airspace convolutional network at the i-th stage.
[0099] The constructed modal interleaving guided classifier performs cross-modal information interaction at the classifier level, aligns the predicted probability distributions of the same category objects under different modalities, and learns the common features of cross-modal group fog, specifically:
[0100]
[0101] Among them, Denote the probability distribution of meteorological element encoding under the guidance of taking the image modality as a metric, Denote the probability distribution of image encoding under the guidance of taking the grid data modality as a metric, f g and f i Respectively denote the feature extractors of the grid data modality and the image modality, and I, G denote the inputs of the image modality and the grid data modality;
[0102]
[0103] L div Denote the difference in probability distribution, Denote the set of different modal interleaving methods considering multi-modal and intra-modal permutation methods, Denote the number of elements in the set, Denote the predicted probability distributions of the same category objects under different interleaving guidance methods.
[0104] It should be noted here that the device descriptions in the above embodiments correspond to the method descriptions in the embodiments, and the embodiments of the present invention will not be elaborated here.
[0105] The execution subjects of the above-mentioned processor and memory can be devices with computing functions such as a computer, a single-chip microcomputer, a microcontroller, etc. In specific implementation, the embodiments of the present invention do not limit the execution subject, and it is selected according to the needs in actual applications.
[0106] Data signals are transmitted between the memory and the processor through a bus, and the embodiments of the present invention will not elaborate on this.
[0107] Based on the same inventive concept, the embodiments of the present invention also provide a computer-readable storage medium. The storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute the method steps in the above embodiments.
[0108] The computer-readable storage medium includes but is not limited to flash memory, hard disk, solid-state drive, etc.
[0109] It should be noted here that the description of the readable storage medium in the above embodiments corresponds to the description of the methods in the embodiments, and the embodiments of the present invention will not be elaborated here.
[0110] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part.
[0111] The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium or a semiconductor medium, etc.
[0112] In the embodiments of the present invention, except for those with special specifications for the models of each device, the models of other devices are not limited, as long as the devices can perform the above functions.
[0113] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.
[0114] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-modal spatiotemporal compensation fog early warning method, characterized in that: The method includes: Construct a parallel neural network including a time-domain non-linear coding part of meteorological elements and a spatial-domain convolutional part of ocean images. The former includes 4 fully connected layers and activation functions, and the latter includes 4 serial convolutional blocks. Based on the parallel network, multi-modal data are respectively encoded. The meteorological element observation data are input into the time-domain non-linear coding part of meteorological elements to generate data encoding, and the satellite remote sensing is input into the spatial-domain convolutional part of ocean images to generate image encoding; Calculate the relative increment of data encoding and image encoding at adjacent times as the change rate, and take the time with a large change rate as the time when the radiation fog forms; Use the cross-modal information interaction part based on semantic matching to fuse the image encoding containing the visual space information of the radiation fog and the data encoding containing the time-series information of meteorological elements to achieve spatio-temporal feature compensation; Construct a modality interleaving guided classifier to perform cross-modal information interaction at the classifier level, align the predicted probability distributions of the same category objects under different modalities, and learn the common features of cross-modal radiation fog; According to the probability distribution output by the modality interleaving guided classifier, if the predicted probability of radiation fog is greater than the threshold, it is determined that radiation fog will occur after a specific time, and a warning signal is issued; Among them, the relative increment of the image encoding is: Embed the boundary features at the time of radiation fog formation into the spectral peak space through a spectral mapping function, and embed the intermediate features during the continuous process of radiation fog into the stationary space; Among them, respectively represent the ocean image features corresponding to the current and previous moments, are the high-frequency mapping function and the low-frequency mapping function respectively. After taking the difference between adjacent image features, high and low frequency feature separation is performed to obtain the fog cluster boundary features and intermediate features respectively. After coupling in the frequency domain, the increment matrix D is obtained through the inverse mapping Φ t , the dimension of the increment matrix is the same as that of the image feature; after obtaining the feature increment matrix, calculate the relative increment α:
2. The multi-modal spatio-temporal compensation-based warning method for patchy fog according to claim 1, wherein, The spatio-temporal feature compensation is: The image features are segmented in the channel domain to obtain That is: The obtained feature segments are respectively used to measure the semantic distance from the meteorological element coding, and the generation of cross-modal compensation features is guided by the semantic matching vector: Among them, is the p-th segmented fragment of the image feature, F T represents the meteorological element code, and ⊙ represents the matrix Hadamard product operation, respectively represent F T mean value; Compensated meteorological element coding Among them, represents D p the semantic matching vector of the image features obtained by vertical splicing and the meteorological element coding, represents the image coding weight conversion operation, which uses the semantic matching vector as the weight to fuse with the image coding to generate the image modality compensation feature; represents the feature map output by the marine image spatial domain convolutional network in the i-th stage; represents the coding of the i-th type of meteorological element sequence data.
3. A multi-modal spatio-temporal compensation method for fog warning according to claim 1, characterized in that, The construction of the modality interleaving guided classifier to perform cross-modal information interaction at the classifier level, align the predicted probability distributions of the same category objects under different modalities, and learn the common features of cross-modal radiation fog specifically is: Among them, represents the probability distribution of meteorological element coding guided by the image modality as a metric, represents the probability distribution of image coding guided by the grid data modality as a metric, f g and f i respectively represent the feature extractors of the grid data modality and the image modality, I and G represent the inputs of the image modality and the grid data modality; R represents the random permutation operation; L div represents the probability distribution difference represents a set of different modality interleaving methods considering multi-modal and intra-modal permutation methods represents the number of elements in the set represents the predicted probability distribution of objects of the same category under different interleaving guidance methods 4. A multi-modal spatio-temporal compensation device for group fog warning, characterized in that, The device includes: a processor and a memory. Program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method steps described in any one of claims 1-3.
Citation Information
Patent Citations
System and method for early warning agglomerate fog on extra-large bridge
CN111341118A
Intelligent typhoon probability forecasting method and device based on multi-modal data fusion
CN115271181A