Lightweight multi-modal foundation cloud picture recognition and prediction method

By building a lightweight network framework, combining pre-trained models and multi-module feature fusion technology, the existing technology's foundation cloud map recognition and prediction problems in complex scenarios and computing resources are solved, and efficient and accurate recognition and prediction effects are achieved.

CN119992190AActive Publication Date: 2025-05-13NORTHEAST DIANLI UNIVERSITY +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510069994.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and predict significant targets in foundation cloud maps in complex scenarios and limited computing resources, especially when the targets and background contrast are low.

Method used

By building a lightweight network framework, a pre-trained model is used to combine edge guidance module, semantic fusion module and spectrum transformation module to perform feature encoding, fusion and decoding, and through supervised learning of optimization strategies such as loss functions and truth-value maps, a lightweight foundation cloud graph recognition and prediction model is trained.

Benefits of technology

With limited computing resources, efficient identification and prediction of foundation cloud maps are achieved, and target recognition accuracy is improved in complex scenarios and low contrast conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992190A_ABST
    Figure CN119992190A_ABST
Patent Text Reader

Abstract

The invention discloses an identification and prediction method for a lightweight multi-modal foundation cloud picture, and relates to the technical field of computer vision. The method comprises the following steps: importing an input image into a pre-training model, and carrying out model training iteration according to an optimization strategy to obtain a trained lightweight foundation cloud picture recognition and prediction model; the optimization strategy comprises the step of performing supervised learning by utilizing a loss function and a truth value graph; the input image comprises multi-modal feature data of a foundation cloud picture; and predicting a to-be-detected image by using the lightweight foundation cloud picture recognition and prediction model, and determining a final output prediction image. According to the method, the recognition and prediction of the foundation cloud picture can be realized by constructing a lightweight network framework under the condition that the computing power resources are limited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method for recognizing and predicting a lightweight multi-modal ground-based cloud image. Background Art

[0002] With the rapid development of renewable energy, wind power projects have become an important part of the global energy industry. The ground-based cloud image saliency recognition and prediction technology plays an important role in wind power forecasting projects, which is closely related to the widespread application of wind power generation technology, the demand for power system stability, and the development of surveying and mapping technology and artificial intelligence. In wind power forecasting projects, ground-based cloud image saliency detection technology is mainly used for cloud layer identification and classification. By performing saliency detection on ground-based cloud images, different types of clouds can be accurately identified, such as cumulus, stratus, cirrus, etc. Different types of clouds have different effects on the generation and transmission of wind energy, so accurate identification of cloud types is crucial to improving the accuracy of wind power forecasting.

[0003] The saliency recognition and prediction technology of ground-based cloud images can be divided into two mainstream methods: traditional methods and deep learning methods. Among them, traditional methods mainly use heuristic priors or intuitive human perception to detect and divide salient target areas, and obtain saliency maps that describe attributes such as bright colors, textures, strong contrast, and compactness. However, the saliency maps generated by traditional methods are not ideal for complex or multi-target images, and cannot be applied to new scenes and complex targets, and have strong limitations. Existing deep learning methods usually require a lot of computing resources to improve detection performance, which will bring new challenges when applied to devices with limited computing power.

[0004] The current mainstream methods have the following shortcomings: first, due to the large differences in structural morphology and boundary details of salient targets, it is difficult to identify and depict target details in complex scenes; second, the information flow provided by semantic information and target details in traditional architectures is limited and cannot cover the entire salient target, which easily leads to incomplete predictions; and traditional architectures find it difficult to capture temporal target details and robust salient information that are feature invariant under certain transformations, which reduces the ability to identify and predict targets in complex noisy scenes; finally, low contrast between targets and backgrounds is common in ground-based cloud images, and misjudgment of feature categories will increase the difficulty of identifying and locating target boundaries under low contrast, thereby affecting the accuracy of predictions.

[0005] Therefore, current mainstream methods cannot ensure the detection accuracy of targets when computing resources are limited. Summary of the invention

[0006] The purpose of the present invention is to provide a lightweight multimodal ground-based cloud image recognition and prediction method, which can realize the recognition and prediction of ground-based cloud images by constructing a lightweight network framework when computing power resources are limited.

[0007] To achieve the above object, the present invention provides the following solutions:

[0008] A lightweight multi-modal ground-based cloud image recognition and prediction method, comprising:

[0009] Acquire an input image; the input image includes multimodal feature data of a ground-based cloud image;

[0010] Constructing a pre-training model; the pre-training model includes a U-shaped connected encoding part and a decoding part; the decoding part includes an edge guidance module, a semantic fusion module and a spectrum conversion module; the input ends of the edge guidance module and the semantic fusion module are both connected to the encoding part, and the output ends are both connected to the spectrum conversion module;

[0011] The input image is imported into the encoding part for feature encoding, the encoded feature data is input into the decoding part for feature fusion and decoding, and the model training iteration is performed according to the optimization strategy to obtain a trained lightweight ground-based cloud image recognition and prediction model; the optimization strategy includes using a loss function and a truth map for supervised learning;

[0012] The lightweight ground-based cloud image recognition and prediction model is used to predict the image to be tested, and the predicted image to be finally output is determined.

[0013] Optionally, the process of importing the input image into the encoding part for feature encoding specifically includes:

[0014] The input image is encoded using MobileNet as a baseline to obtain encoded first-stage features, second-stage features, third-stage features, fourth-stage features and fifth-stage features; wherein the first-stage features and the second-stage features are used to describe location information and target details; and the third-stage features, the fourth-stage features and the fifth-stage features are used to describe abstract semantic information.

[0015] Optionally, the process of inputting the encoded feature data into the decoding part for feature fusion and decoding specifically includes:

[0016] Inputting the first-stage features and the second-stage features into the edge-guided module for operation to obtain spatial features of different receptive fields, and performing element-by-element addition / subtraction, Hadamard product operation and channel cascade strategy on the spatial features to obtain fine-grained target detail features;

[0017] Performing semantic compression and dilated pyramid pooling processing on the fifth stage features to generate convolution kernels with different expansion rates for the semantic fusion module;

[0018] Based on the convolution kernels with different expansion rates, the third stage features and the fourth stage features are input into the semantic fusion module for operation to obtain multi-scale features of semantic information, and a depth-separable convolution operation and a channel cascade strategy are performed on the multi-scale features to obtain rich semantic information features;

[0019] The fine-grained target detail features and the rich semantic information features are input into the spectrum transformation module, the spatial domain features are combined with the frequency domain features, and multi-scale feature fusion and decoding are performed through channel cascading and step-by-step fusion transformation operations.

[0020] Optionally, the first-stage features and the second-stage features are input into the edge guidance module for calculation, and the specific process includes:

[0021] A multi-scale convolution operation is performed on the first-stage features and the second-stage features to obtain spatial features of different receptive fields; the formula of the multi-scale convolution operation is:

[0022]

[0023] in, and They represent the first-stage characteristics and the second-stage characteristics respectively, and f i 1 and f i 2 They represent the generated spatial features, Ψ represents the bilinear interpolation operation, C 2i-1 (i=1,…,4) represents a multi-scale convolution kernel, DSConv(*,*) represents a depth-wise separable convolution operation;

[0024] Based on the spatial feature f1 1 , f1 2 and Through convolution operation, average pooling operation and element-by-element subtraction operation, fusion features and prediction features are obtained. The mathematical formula is expressed as:

[0025]

[0026] Among them, g1 and g2 represent the generated fusion features, and Respectively represent the generated prediction features, Conv(*,ε i ) represents the convolution calculation with a kernel size of 1×1. represents the average pooling operation; ρ represents the PReLU function;

[0027] In the space feature f3 1 , f3 2 , The channel cascade strategy, Hadamard product operation and element-by-element addition operation are used to integrate the spatial detail feature information on the fusion features g1 and g2 to obtain the fine-grained target detail features. The mathematical formula is expressed as:

[0028]

[0029] D out =CCS(a1,a2)

[0030] Among them, × represents the Hadamard product operation, σ represents the Sigmoid function, represents two-dimensional batch normalization, D out represents fine-grained target detail features, CCS represents the channel cascade strategy, h1 and h2 represent the corresponding intermediate layer features, and a1 and a2 represent the corresponding dense spatial features.

[0031] Optionally, based on the convolution kernels with different dilation rates, the third stage features and the fourth stage features are input into the semantic fusion module for operation, and the specific process includes:

[0032] Performing a dynamic deep convolution operation with a dilation rate on the third stage features and the fourth stage features to obtain multi-scale features and The calculation formula is:

[0033]

[0034] Among them, DConv represents the dynamic depth convolution operation, ω and υ represent the convolution kernel matrix, r i (r i =1,2,3) represents the expansion ratio, and They represent the third stage characteristics and the fourth stage characteristics respectively;

[0035] Use element-wise addition and point-wise convolution operations to multi-scale features and Perform integration to generate features y1 and y2. The calculation formula is:

[0036]

[0037] Among them, + represents element-by-element addition operation, Ψ represents bilinear interpolation operation, PConv is point-by-point convolution, DSConv represents depth-wise separable convolution, αi and β i All represent convolution kernels (i=1,2);

[0038] The channel cascade strategy is performed on the features y1 and y2 to fuse the feature information and obtain the feature h with rich semantic information. out , the calculation formula is:

[0039] h out =CCS(y1,y2)

[0040] Among them, CCS represents the channel cascade strategy.

[0041] Optionally, the loss function in the optimization strategy is constructed based on a first loss function and a second loss function; the first loss function is composed of a binary cross entropy loss function and an intersection-over-union loss function; and the second loss function adopts an L2 loss function.

[0042] Optionally, the first loss function and the second loss function specifically include:

[0043] Define the binary cross entropy loss function And the intersection-over-union loss function for:

[0044]

[0045] Among them, p(x,y)∈[0,1] is the predicted probability of the image pixel point (x,y), and g(x,y)∈[0,1] is the true value label of the image pixel point (x,y);

[0046] Construct the first loss function of the i-th stage:

[0047] Construct the second loss function: Among them, L mse represents the mean square error loss, ψ is the parameter correction linear unit, and They represent the prediction features generated in the edge guidance module respectively.

[0048] Optionally, the loss function in the optimization strategy is expressed as:

[0049]

[0050] Among them, α=1 and β=0.5.

[0051] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0052] The present invention discloses a method for recognizing and predicting a lightweight multimodal ground-based cloud image, the method comprising importing an input image into a pre-trained model, and performing model training iteration according to an optimization strategy to obtain a trained lightweight ground-based cloud image recognition and prediction model; the optimization strategy comprises using a loss function and a truth map for supervised learning; the input image comprises multimodal feature data of the ground-based cloud image; the lightweight ground-based cloud image recognition and prediction model is used to predict the image to be tested, and the predicted image outputted finally is determined. The present invention can realize the recognition and prediction of ground-based cloud images by constructing a lightweight network framework when computing power resources are limited. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0054] Figure 1 It is a schematic diagram of the process flow of the lightweight multi-modal ground-based cloud image recognition and prediction method of the present invention;

[0055] Figure 2 This is a comparison chart of the significance prediction results using the method of the present invention and other methods in this embodiment. DETAILED DESCRIPTION

[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0057] The purpose of the present invention is to provide a lightweight multimodal ground-based cloud image recognition and prediction method, which can realize the recognition and prediction of ground-based cloud images by constructing a lightweight network framework when computing power resources are limited.

[0058] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0059] like Figure 1As shown, the present invention provides a lightweight multimodal ground-based cloud image recognition and prediction method. Considering the differences in multi-level feature spaces, a multi-feature fusion network is first designed as the baseline for saliency reasoning; in order to effectively aggregate fine-grained visual features and abstract semantic information while ensuring the consistency of feature space, an edge guidance module and a semantic fusion module are proposed and integrated into the multi-feature fusion network; in order to enhance the robustness of saliency information, a spectrum transformation module combining time domain and frequency domain features is proposed to highlight saliency target details and encourage saliency prediction. In terms of optimization settings, a saliency prediction result optimization strategy is proposed to remove interference information and noise in complex images. This strategy uses binary cross entropy loss and intersection-over-union loss to calculate the total pixel loss, aiming to optimize the global structure of the input image and improve the accuracy of saliency prediction. Mean square error loss is also introduced to further express enhanced target details. In terms of result comparison, the experimental results of the method proposed in the present invention on four ground-based cloud image datasets show that the method of the present invention is superior to other existing advanced methods. The specific steps are:

[0060] An input image is obtained; the input image includes multimodal feature data of a ground-based cloud image.

[0061] Constructing a pre-training model; the pre-training model includes a U-shaped connected encoding part and a decoding part; the decoding part includes an edge guidance module, a semantic fusion module and a spectrum conversion module; the input ends of the edge guidance module and the semantic fusion module are both connected to the encoding part, and the output ends are both connected to the spectrum conversion module;

[0062] The input image is imported into the encoding part for feature encoding, the encoded feature data is input into the decoding part for feature fusion and decoding, and the model training iteration is performed according to the optimization strategy to obtain a trained lightweight ground-based cloud image recognition and prediction model; the optimization strategy includes using a loss function and a truth map for supervised learning;

[0063] The lightweight ground-based cloud image recognition and prediction model is used to predict the image to be tested, and the predicted image to be finally output is determined.

[0064] As a specific implementation method, the overall structure is as follows Figure 1 As shown in Figure 2. MobileNet is used as a baseline in the image encoding stage. Features are encoded in a top-down and bottom-up manner based on a U-shaped architecture, where E1-E5 represent low-level features to high-level features, respectively. E1 and E2 have rich location information and target details, which are more conducive to obtaining salient boundaries. As high-level features, E3, E4, and E5 contain abstract semantic information, which is more conducive to understanding target categories.

[0065] First, the high-level feature E5 is used to generate convolution kernels with different dilation rates for the semantic fusion module; secondly, E3 and E4 are used as the input of this module to obtain multi-scale features with abstract semantic information, and depth-wise separable convolution operations and channel cascade strategies are used on the multi-scale features to obtain richer semantic information; thirdly, features E1 and E2 are used as the input of the edge-guided module to obtain spatial features with different receptive fields, and element-wise addition / subtraction, Hadamard product operations and channel cascade strategies are performed on the spatial features to generate output features with fine-grained object details. Fourth, the features containing rich semantic information and fine-grained object details are used as the input of the spectrum transformation module, and the spatial domain features are combined with the frequency domain features, which helps to obtain a more comprehensive object representation. Then, multi-scale feature fusion and decoding are performed through channel cascade and step-by-step fusion transformation. Finally, we use the loss function to fully supervise the output saliency prediction of each level of the architecture and generate the final ground-based cloud map saliency map, which is supervised in real time by the saliency truth map.

[0066] The lightweight multimodal ground-based cloud image recognition and prediction framework proposed in the present invention includes a semantic fusion module, an edge guidance module and a spectrum transformation module. The specific calculation methods and formulas of each module are as follows.

[0067] 1. Semantic Fusion Module

[0068] Step 1: Input features and Perform dynamic depthwise convolution with dilation rate to obtain features and The detailed calculation formula is as follows:

[0069]

[0070] Among them, DConv represents the dynamic depth convolution operation, ω and υ represent the convolution kernel matrix, r i is the expansion ratio (r i =1,2,3), and They represent the third stage characteristics and the fourth stage characteristics respectively.

[0071] Step 2: Use element-wise addition and point-wise convolution operations to integrate multi-scale features and Thus generating features y1 and y2.

[0072]

[0073] Where + represents element-by-element addition, Ψ represents bilinear interpolation, PConv represents point-by-point convolution, DSConv represents depth-wise separable convolution, and α iand β i Both represent convolution kernels (i=1,2).

[0074] Step 3: Perform channel cascade strategy on features y1 and y2 to fuse feature information and finally output h out It is expressed as:

[0075] h out =CCS(y1,y2) (3)

[0076] Among them, CCS represents the channel cascade strategy.

[0077] 2. Channel Cascading Strategy

[0078] Step 1: Perform scaled harmonic mapping and matrix multiplication operations on the input features x1 and x2∈{N,C,H,W} to obtain the activated feature tensor z.

[0079]

[0080] where · represents matrix multiplication, f i (i=1,2) represents the scale harmonic map, and τ represents the softmax function.

[0081] Step 2: Use inverse scale harmonic mapping on features y1, y2 and z to fuse and reconstruct spatial features, thereby obtaining features u1 and u2. The mathematical expression is shown in equation (5).

[0082] u1=f1 -1 (y1·z),u2=f2 -1 (z·y2) (5)

[0083] Among them, f i -1 (i=1,2) represents the inverse scale harmonic mapping.

[0084] Step 3: Element-wise addition, channel convolution operations, and depth-wise separable convolution operations are performed to generate the output features g out .

[0085]

[0086] Among them, DSConv(*,*) represents the depth-wise separable convolution operation, and ω, υ, and μ represent the convolution kernel matrices.

[0087] 3. Edge Guidance Module

[0088] Step 1: Multi-scale convolution operations are performed on the input features and Thus generating spatial features f with different receptive fields i 1and f i 2 .

[0089]

[0090] Where Ψ is the bilinear interpolation algorithm, C 2i-1 (i=1,…,4) represents a multi-scale convolution kernel.

[0091] Step 2: Obtain the fused features g1 and g2 and the predicted features through convolution operation, average pooling operation and element-by-element subtraction operation and Its mathematical formula is as follows:

[0092]

[0093] Among them, g1 and g2 represent the generated fusion features, and Respectively represent the generated prediction features, Conv(*,ε i ) represents the convolution calculation with a kernel size of 1×1. represents the average pooling operation; ρ represents the PReLU function.

[0094] Step 3: In the space feature f3 1 , f3 2 , The channel cascade strategy, Hadamard product operation and element-by-element addition operation are used to integrate the spatial detail feature information on the fusion features g1 and g2 to obtain the spatial aggregation feature D out It is expressed as:

[0095]

[0096] Among them, × represents the Hadamard product operation, σ represents the Sigmoid function, represents two-dimensional batch normalization, D out represents fine-grained target detail features, CCS represents the channel cascade strategy, h1 and h2 represent the corresponding intermediate layer features, and a1 and a2 represent the corresponding dense spatial features.

[0097] 4. Spectrum transformation module

[0098] Step 1: For the input feature f in Perform convolution operation to obtain hierarchical features X G and X L .

[0099] X G =DSConv(f in ,ω),XL =DSConv(f in ,υ) (10)

[0100] Where ω and υ represent the convolution kernel of size 3×3.

[0101] Step 2: Spectral transformation and convolution operation are applied to feature X G , to obtain the transformation feature X GG and X GL .

[0102] X GG =ST(X G ), X GL =DSConv(X G ,θ) (11)

[0103] Where ST represents the Fourier transform module, θ∈R 1×1×C .

[0104] Step 3: Multi-region convolution operations are performed on the level feature X L On the top, we get the spatial domain feature X LG and X LL as follows:

[0105] X LG =MR.C(X L ,m),X LL =MR.C(X L ,n) (12)

[0106] Among them, MR.C represents the multi-region convolution operation module, and m and n both represent convolution kernels.

[0107] Step 4: The multi-channel output features are integrated through element-by-element addition and nonlinear transformation to obtain the fused feature X OG and X OL It is expressed as shown in formula (13).

[0108] X OG =δ(X GG +X LG ), X OL =δ(X GL +X LL ) (13)

[0109] Where δ represents the ReLU function.

[0110] Step 5: Multi-channel features X OG and X OL is cascaded and compressed to enrich the pixel-level data representation, which is more conducive to the positioning and detection of the target, and finally outputs the feature f out It is expressed as the following mathematical formula.

[0111] f out =DSConv(τ(Concate([X OG ,X OL ],dim=1)),θ) (14)

[0112] Where τ represents channel compression and θ is a convolution kernel of size 1×1.

[0113] 5. Multi-region convolution operation module

[0114] Step 1: For the input feature x in Perform scale-harmony tensor partitioning to generate small-scale features U L , U R , L L and L R .

[0115] U L ,U R ,L L ,L R =ψ(x in ) (15)

[0116] Where ψ represents the scale adjustment and tensor partition, x in ∈{N,C,H,W},

[0117] Step 2: In Feature U L , U R , L L and L R Convolution operation is performed on the y-axis to obtain multi-region features x i (i=1,…,4).

[0118]

[0119] where ω i (i=1,…,4) represents the convolution kernel matrix.

[0120] Step 3: Use feature concatenation and convolution operations to perform multi-region feature i Reconstruct and fuse, output feature x out It is expressed as follows:

[0121]

[0122] Concate(*, dim=3) means concatenation at the width of the tensor, and Concate(*, dim=2) means concatenation at the height of the tensor. is the convolution kernel matrix, and x out ∈{N,C,H,W}.

[0123] 6. Fourier Transform Module

[0124] Step 1: Input feature y in Perform Fourier transform and feature reshaping to obtain frequency domain features and

[0125]

[0126] Where FFT stands for Fourier transform, ρ is the feature reshaping, θ i (i=1,2) represents a convolution kernel of size 1×1.

[0127] Step 2: Frequency Domain Features and are merged and reshaped to create fused frequency domain features, and an inverse Fourier transform is applied to obtain features z in the time domain.

[0128]

[0129] Here, · represents matrix multiplication operation, κ represents structural transformation, and IFFT represents inverse Fourier transform.

[0130] Step 3: Element-wise addition and depth-wise separable convolution operations are used to integrate the time domain features to obtain the output feature y out It can be expressed as follows:

[0131] y out =DSConv(y1+y2+z,θ) (20)

[0132] Where DSConv(*,θ) represents a function with parameters θ∈R 1×1×C Depthwise separable convolution.

[0133] Optimizing strategies for results

[0134] In order to further optimize the saliency prediction results and remove the interference information and noise in the image, this scheme proposes a new loss function, which consists of two parts: binary cross entropy loss Sum intersection loss Able to calculate the loss for each pixel, Aims to optimize the global structure of the input image. These two components complement each other, thus improving the accuracy of saliency prediction. and The definition at stage i is as follows:

[0135]

[0136]

[0137] Among them, p(x,y)∈[0,1] is the predicted probability of pixel point (x,y), and g(x,y)∈[0,1] is the true value label of pixel point (x,y).

[0138] Therefore, the total loss function of the i-th stage is As shown in formula (23):

[0139]

[0140] Then, L2 loss is used to further express the enhanced target details, and the function L2 is expressed as:

[0141]

[0142] Where L mse represents the mean square error loss, ψ is the parameter correction linear unit, and They represent the prediction features generated in the edge guidance module respectively.

[0143] Finally, the total loss L total It is expressed by mathematical formula (25):

[0144]

[0145] Where α=1 and β=0.5.

[0146] 3.2.2 Results of the technical solution of the present invention (or utility model)

[0147] In this embodiment, the model training is performed on the SWINYSEG dataset with additional annotations, and the model evaluation is performed on the HBMSEG, SWIMSEG, SWINSEG and SHWIMSEG datasets. The information of each training and test dataset is as follows: the SWINYSEG dataset contains 6768 sky and cloud images with 10 types of cloud information, all of which are used in the training phase; the HBMSEG dataset contains 11,000 complex images, which are used in the testing phase; the SWIMSEG dataset contains 1013 test images of sky and cloud patches; the SHWIMSEG dataset contains 156 complex test images, most of which have low, medium and high exposure features; the SWINSEG dataset contains 115 night images of the sky and clouds.

[0148] In this embodiment, the comparison results of F-measure, MAE and S-measure of the prediction technology of the present invention and other technologies on the SWIMSEG, SWINSEG, SHWIMSEG, and HBMSEG datasets are shown in Table 1. From the results in Table 1, we can find that the prediction technology of the present invention shows good performance on almost all ground cloud image databases in terms of the three scoring indicators, which also shows the effectiveness and usability of the method of the present invention. Specifically, the prediction technology of the present invention achieves results of 0.900, 0.931, 0.711 and 0.539 respectively when using the S-measure metric to calculate the scores on the four datasets. In addition, in most cases, the method proposed in this scheme outperforms other existing lightweight state-of-the-art methods, such as HVPNet and SAMNet, which also proves the reliability of the technology of this scheme. In particular, when the S-measure metric (F-measure metric) is used to calculate the scores on the SWIMSEG, SWINSEG, SHWIMSEG, and HBMSEG datasets, the performance of our method is better than SAMNet by 7.6% (6.1%), 11.1% (6.4%), 4.9% (4.3%), and 19.8% (17.1%), respectively, which greatly reduces the error.

[0149] Table 1 Comparison of F-measure, S-measure and MAE of the proposed method with other methods on four data sets

[0150]

[0151] This embodiment also provides a visual comparison of the saliency prediction of ground cloud images using the technology of the present invention and other existing state-of-the-art methods on four datasets using some challenging scenes (including high exposure (1st and 2nd rows), dark environment (3rd and 4th rows), clumping clouds (5th and 7th rows) and reticular clouds (6th and 8th rows), as shown in Figure 1. Figure 2As shown, it can be concluded that the method of the present invention exhibits more continuous and smoother target detection and segmentation performance, which provides a clear description of the target morphology. Specifically, the first row provides a cloud image in a high-exposure scene. The method of the present invention can clearly identify and describe the details of the clouds in the high-exposure area and achieve satisfactory results. In the seventh row, despite the presence of large clouds, most methods cannot clearly locate and segment the entire prominent cloud area. Compared with the fifth row, the reticular cloud image in the eighth row has a target cloud color that is very similar to the sky background, making it more challenging to locate and segment the reticular cloud, mainly due to its more dispersed and complex distribution; from the visualization results, it can be seen that the method of the present invention can accurately locate the cloud and segment its shape, while detecting fewer false pixels and noise. At the same time, the method of the present invention shows effectiveness in suppressing background noise in a dark environment and accurately locating and segmenting significant cloud targets. As shown in the third and fourth rows, the image depicts the distribution of reticular clouds in a dark environment (the fourth row). When the distribution is more dispersed and complex, the method proposed by the present invention can also accurately identify the location of the cloud image target.

[0152] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0153] This article uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only used to help understand the core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A lightweight multi-modal ground-based cloud image recognition and prediction method, characterized in that: include: Acquire an input image; the input image includes multimodal feature data of a ground-based cloud image; Constructing a pre-training model; the pre-training model includes a U-shaped connected encoding part and a decoding part; the decoding part includes an edge guidance module, a semantic fusion module and a spectrum conversion module; the input ends of the edge guidance module and the semantic fusion module are both connected to the encoding part, and the output ends are both connected to the spectrum conversion module; The input image is imported into the encoding part for feature encoding, the encoded feature data is input into the decoding part for feature fusion and decoding, and the model training iteration is performed according to the optimization strategy to obtain a trained lightweight ground-based cloud image recognition and prediction model; the optimization strategy includes using a loss function and a truth map for supervised learning; The lightweight ground-based cloud image recognition and prediction model is used to predict the image to be tested, and the predicted image to be finally output is determined.

2. The method for recognizing and predicting lightweight multimodal ground-based cloud images according to claim 1, characterized in that: The process of importing the input image into the encoding part for feature encoding specifically includes: The input image is encoded using MobileNet as a baseline to obtain encoded first-stage features, second-stage features, third-stage features, fourth-stage features and fifth-stage features; wherein the first-stage features and the second-stage features are used to describe location information and target details; and the third-stage features, the fourth-stage features and the fifth-stage features are used to describe abstract semantic information.

3. The method for recognizing and predicting lightweight multimodal ground-based cloud images according to claim 2, characterized in that: The process of inputting the encoded feature data into the decoding part for feature fusion and decoding specifically includes: Inputting the first-stage features and the second-stage features into the edge-guided module for operation to obtain spatial features of different receptive fields, and performing element-by-element addition / subtraction, Hadamard product operation and channel cascade strategy on the spatial features to obtain fine-grained target detail features; Performing semantic compression and dilated pyramid pooling processing on the fifth stage features to generate convolution kernels with different expansion rates for the semantic fusion module; Based on the convolution kernels with different expansion rates, the third stage features and the fourth stage features are input into the semantic fusion module for operation to obtain multi-scale features of semantic information, and a depth-separable convolution operation and a channel cascade strategy are performed on the multi-scale features to obtain rich semantic information features; The fine-grained target detail features and the rich semantic information features are input into the spectrum transformation module, the spatial domain features are combined with the frequency domain features, and multi-scale feature fusion and decoding are performed through channel cascading and step-by-step fusion transformation operations.

4. The method for recognizing and predicting lightweight multimodal ground-based cloud images according to claim 3 is characterized in that: The first-stage features and the second-stage features are input into the edge guidance module for calculation, and the specific process includes: A multi-scale convolution operation is performed on the first-stage features and the second-stage features to obtain spatial features of different receptive fields; the formula of the multi-scale convolution operation is: in, and They represent the first stage characteristics and the second stage characteristics respectively. and They represent the generated spatial features, Ψ represents the bilinear interpolation operation, C 2i-1 (i=1,…,4) represents a multi-scale convolution kernel, DSConv(*,*) represents a depth-wise separable convolution operation; Based on spatial features and Through convolution operation, average pooling operation and element-by-element subtraction operation, fusion features and prediction features are obtained. The mathematical formula is expressed as: Among them, g1 and g2 represent the generated fusion features, and Respectively represent the generated prediction features, Conv(*,ε i ) represents the convolution calculation with a kernel size of 1×1. represents the average pooling operation; ρ represents the PReLU function; Features in space The channel cascade strategy, Hadamard product operation and element-by-element addition operation are used to integrate the spatial detail feature information on the fusion features g1 and g2 to obtain the fine-grained target detail features. The mathematical formula is expressed as: <h2 style=";text-align:left;direction:ltr">D<h2 style=";text-align:left;direction:ltr"> out <h2 style=";text-align:left;direction:ltr"> =CCS(a1,a2) Among them, × represents the Hadamard product operation, σ represents the Sigmoid function, represents two-dimensional batch normalization, D out represents fine-grained target detail features, CCS represents the channel cascade strategy, h1 and h2 represent the corresponding intermediate layer features, and a1 and a2 represent the corresponding dense spatial features.

5. The method for recognizing and predicting lightweight multimodal ground-based cloud images according to claim 3, characterized in that: Based on the convolution kernels with different expansion rates, the third stage features and the fourth stage features are input into the semantic fusion module for operation, and the specific process includes: Performing a dynamic deep convolution operation with a dilation rate on the third stage features and the fourth stage features to obtain multi-scale features and The calculation formula is: Among them, DConv represents the dynamic depth convolution operation, ω and υ represent the convolution kernel matrix, r i (r i =1,2,3) represents the expansion ratio, and They represent the third stage characteristics and the fourth stage characteristics respectively; Use element-wise addition and point-wise convolution operations to multi-scale features and Perform integration to generate features y1 and y2. The calculation formula is: Among them, + represents element-by-element addition operation, Ψ represents bilinear interpolation operation, PConv is point-by-point convolution, DSConv represents depth-wise separable convolution, α i and β i All represent convolution kernels (i=1,2); The channel cascade strategy is performed on the features y1 and y2 to fuse the feature information and obtain the feature h with rich semantic information. out , the calculation formula is: h out =CCS(y1,y2) Among them, CCS represents the channel cascade strategy.

6. The method for recognizing and predicting lightweight multimodal ground-based cloud images according to claim 1, characterized in that: The loss function in the optimization strategy is constructed based on a first loss function and a second loss function; the first loss function is composed of a binary cross entropy loss function and an intersection-over-union loss function; and the second loss function adopts an L2 loss function.

7. The method for recognizing and predicting lightweight multimodal ground-based cloud images according to claim 6, characterized in that: The first loss function and the second loss function specifically include: Define the binary cross entropy loss function And the intersection-over-union loss function for: Among them, p(x,y)∈[0,1] is the predicted probability of the image pixel point (x,y), and g(x,y)∈[0,1] is the true value label of the image pixel point (x,y); Construct the first loss function of the i-th stage: Construct the second loss function: Among them, L mse represents the mean square error loss, ψ is the parameter correction linear unit, and They represent the prediction features generated in the edge guidance module respectively.

8. The method for recognizing and predicting lightweight multimodal ground-based cloud images according to claim 7, characterized in that: The loss function in the optimization strategy is expressed as: Among them, α=1 and β=0.5.

Citation Information

Patent Citations

  • Medical image gland segmentation method

    CN116563315A

  • Landslide image segmentation method based on multilayer feature information fusion

    CN118261926A

  • RGB-D underwater saliency target detection method based on semantic guidance fusion

    CN118570623A

  • Image segmentation method of mirror image semantic segmentation lightweight insight network based on double contrast knowledge distillation

    CN118941788A

  • Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment

    WO2024230038A1