Full-solar-surface sun image quality classification method and system based on deep learning
By constructing a deep learning model that integrates Res2NetBasicBlock and ViTAttention modules, the problem of low efficiency in the quality assessment of full-surface solar images was solved, achieving efficient and automated image classification and improving the model's generalization performance and classification accuracy.
Patent Information
- Application Number
- CN202511042034.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-14
AI Technical Summary
Existing methods for assessing the quality of full-plane solar images rely on manually designed physical features and thresholds, resulting in low efficiency. Furthermore, deep learning models exhibit complexity and instability in cloud detection, making it difficult to balance the generation and discrimination processes, leading to decreased generalization performance.
We employ a deep learning-based approach to construct a model that integrates the Res2Net BasicBlock module and the ViTAttention module. This model performs image quality classification through multi-scale convolution and global attention mechanisms, including dataset construction, model training, and hyperparameter optimization. We use a label-smooth cross-entropy loss function and dynamic learning rate adjustment.
It achieves efficient and automated full-surface solar image quality classification, improves data processing efficiency, eliminates the need for manual inspection, and has excellent classification performance and generalization ability.
Smart Images

Figure CN120953666A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system for classifying the quality of full-plane solar images based on deep learning, belonging to the field of astronomical image classification technology. Background Technology
[0002] Ground-based observation is a crucial tool in solar physics research, helping to reveal solar activity patterns and provide early warnings of space disasters. However, ground-based observation equipment is inevitably obscured by clouds between the Earth and the Sun, leading to image contamination and reduced image quality. Therefore, these contaminated images must be screened and classified. The usual practice is manual screening. However, modern ground-based observation equipment generates massive amounts of data. For example, the SolarFull-disk Multi-layer Magnetograph (SFMM) can capture solar images at second-level intervals, acquiring several terabytes of Hα solar image data each month, of which approximately 10% are obscured by clouds to varying degrees. Therefore, relying on manual assessment of image usability for solar image data acquired by SFMM is impractical; a reasonable and efficient image quality assessment method is needed for evaluation and screening to ensure high-quality data.
[0003] Current methods for assessing the quality of full-plane solar images are mainly traditional, relying on manually designed physical features and corresponding thresholds for differentiation. However, the selection of thresholds is largely influenced by human experience, increasing the need for manual intervention and significantly reducing data processing efficiency. To overcome the limitations of traditional methods, the use of deep learning techniques for full-plane solar image quality assessment has been gradually explored. Existing deep learning detection methods use manually labeled data to train adversarial neural networks for cloud detection. However, their shortcomings lie in the high complexity and instability of adversarial training, making it difficult for the model to balance the generation and discrimination processes in practical applications. Compared to other deep learning methods, it is more susceptible to hyperparameter sensitivity and local optimization limitations, resulting in a significant decrease in generalization performance. Therefore, providing high-quality methods and approaches for full-plane solar image quality classification remains an urgent problem to be solved. Summary of the Invention
[0004] This invention provides a deep learning-based method for classifying the quality of full-disk solar images, which can be used to better classify the quality of full-disk solar images captured by SFMM without further manual inspection, thereby improving data processing efficiency.
[0005] The technical solution of this invention is:
[0006] According to a first aspect of the present invention, a deep learning-based method for classifying the quality of full-surface solar images is provided, comprising:
[0007] The acquired full-surface solar images are classified to obtain full-surface solar images of multiple quality categories with balanced proportions, and a dataset is constructed.
[0008] Each image in the dataset is scaled and cropped in turn to obtain a dataset of a preset size; the dataset of the preset size is then randomly divided into a training set, a validation set, and a test set according to a preset ratio.
[0009] A deep learning model is constructed based on the concatenated head network, backbone network, and output layer network.
[0010] The deep learning model was trained and its hyperparameters optimized based on the training and validation sets, and the model that performed best on the test set was saved as the full-surface solar image quality classification model.
[0011] The full-plane solar image to be classified is used as input to the full-plane solar image quality classification model to obtain the full-plane solar image quality classification result.
[0012] Furthermore, the acquired full-surface solar images were categorized into three quality classes based on the degree of cloud cover: clean, relatively light, and relatively severe.
[0013] Furthermore, the deep learning model is specifically as follows:
[0014] Build the Res2NetBasicBlock module;
[0015] Construct the ViTAttention module;
[0016] A composite module is constructed based on the Res2NetBasicBlock module and the ViTAttention module;
[0017] Convolution, Batch Normalization (BN), ReLU, and max pooling are used as the head network of the deep learning model. The backbone network consists of four concatenated composite modules, with downsampling operations introduced after the first three composite modules. The output layer network is built after the fourth composite module. The output layer network obtains the output through adaptive average pooling, flattening, and fully connected layer operations.
[0018] Furthermore, the construction of the composite module is specifically as follows: a composite module is constructed based on two stacked Res2NetBasicBlock modules and a ViTAttention module connected in series after the two stacked Res2NetBasicBlock modules.
[0019] Furthermore, the Res2NetBasicBlock module includes:
[0020] Channel decomposition: The input feature map is evenly divided into 4 sub-feature maps along the channel dimension;
[0021] Hierarchical convolution: Except for the first sub-feature map, the nth sub-feature map is added element-wise to the output of the (n-1)th sub-feature map, followed by a 3×3 group convolution, and then a 1×1 convolution is added to restore the channel interaction capability to obtain the output of the nth sub-feature map; n=2,3,4;
[0022] Feature fusion: Features from all branches of the hierarchical convolution are concatenated and merged to generate a high-order representation with multi-scale fusion; then, the concatenated features are interacted and fused across channels through convolution and BN to serve as the main path output; finally, residual connections are introduced to access the main path output.
[0023] Furthermore, the ViTAttention module includes:
[0024] Location encoding: Learnable location encoding is fused with input features through bilinear interpolation, preserving prior knowledge of two-dimensional spatial structure;
[0025] Multi-head attention: A global attention mechanism is introduced to calculate the global pixel correlation matrix, and a progressive multi-head mechanism is used. The module will divide the feature map into multiple independent subspaces according to the number of progressive heads from the channel dimension.
[0026] Feature reconstruction: The multi-head attention output is fused through convolution to preserve important spatial response regions.
[0027] Furthermore, the deep learning model training strategy is as follows:
[0028] Learning rate adjustment strategy: Use dynamic learning rate, with an initial learning rate set at 3×10. -4 The learning rate is updated based on the macro-average F1 score on the validation set. If the score does not improve for three consecutive epochs, the learning rate is decayed with a coefficient of 0.5, with a lower limit of 1×10. -6 ;
[0029] Loss function selection: Introducing label-smoothed cross-entropy loss and setting a smoothing coefficient λ, the theoretical lower bound expression of the loss function is established as follows:
[0030] According to a second aspect of the present invention, a deep learning-based full-surface solar image quality classification system is provided, comprising:
[0031] The first construction module is used to classify the acquired full-surface solar images to obtain full-surface solar images of multiple quality categories with balanced proportions, and to construct a dataset;
[0032] The first acquisition module is used to perform scaling and center cropping operations on each image in the dataset in sequence to obtain a dataset of a preset size; the dataset of the preset size is then randomly divided into a training set, a validation set, and a test set according to a preset ratio.
[0033] The second building module is used to construct a deep learning model based on the concatenated head network, backbone network, and output layer network.
[0034] The training and selection module is used to train the deep learning model and optimize its hyperparameters based on the training and validation sets, and save the model that performs best on the test set as the full-surface solar image quality classification model.
[0035] The second acquisition module is used to take the full-surface solar image to be classified as input to the full-surface solar image quality classification model and obtain the full-surface solar image quality classification result.
[0036] The beneficial effects of this invention are:
[0037] This invention proposes a deep learning model architecture that integrates multi-scale convolution and global attention mechanisms. The head network uses convolution to achieve coarse-grained capture of low-level semantic features. The core modules of the model include the Res2NetBasicBlock module and the ViTAttention module. The Res2NetBasicBlock module adopts a hierarchical processing architecture, achieving multi-scale feature extraction through channel splitting, balancing computational efficiency and model performance. The ViTAttention module introduces a global attention module, with the number of attention heads increasing progressively with network depth, thus establishing a feature enhancement mechanism of "local convolutional perception - global attention association". The network's end uses a lightweight classification module, achieving compressed mapping of high-dimensional features to the class space through channel projection and spatial pooling, ultimately outputting the class probability distribution via a fully connected layer. This architecture, through a stage-cascaded convolution-attention collaborative design, integrates global feature association capabilities while maintaining the local feature extraction capabilities of convolution, constructing a feature extraction mechanism that combines multi-scale local perception and spatial global association reasoning. By combining the advantages of convolutional neural networks and attention mechanisms, this invention achieves excellent quality classification performance on images captured by SFMM, eliminating the need for further manual inspection. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of the structure of the deep learning model proposed in this invention.
[0039] Figure 2 This is a schematic diagram of the Res2NetBasicBlock structure.
[0040] Figure 3 This is a schematic diagram of the ViTA attention structure.
[0041] Figure 4 A clean image of the sun taken by SFMM.
[0042] Figure 5 An image of the sun with slight cloud cover, taken by SFMM.
[0043] Figure 6 An image of the sun severely affected by clouds, captured by SFMM.
[0044] Figure 7 This is a graph showing the loss function during training of the deep learning model proposed in this invention.
[0045] Figure 8 This is the classification performance of the deep learning model proposed in this invention on the test set.
[0046] Figure 9 This is the classification performance of UNet on the test set.
[0047] Figure 10 This is the classification performance of ResNet on the test set. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other.
[0049] Example 1: As Figures 1-10 As shown, according to a first aspect of the present invention, a deep learning-based method for classifying the quality of full-surface solar images is provided, comprising:
[0050] Step 1: Classify the acquired full-surface solar images to obtain three categories of full-surface solar images with equal proportions, and construct a dataset. The three categories of full-surface solar images are: Category 1 full-surface solar images, Category 2 full-surface solar images, and Category 3 full-surface solar images.
[0051] For example, the acquired full-surface solar images are divided into three categories based on the degree of cloud cover: clean, relatively light, and relatively severe.
[0052] Step 2: Perform scaling and center cropping operations on each image in the dataset in sequence to obtain a dataset of a preset size; randomly divide the dataset of the preset size into a training set, a validation set, and a test set according to a preset ratio;
[0053] Step 3: Based on the concatenated head network, backbone network, and output layer network, construct a deep learning model; further, this includes:
[0054] Step 3.1: Construct the Res2NetBasicBlock module;
[0055] refer to Figure 2 The Res2NetBasicBlock module aims to implement a hierarchical residual multi-scale feature fusion mechanism, specifically including:
[0056] 1. Channel decomposition: The input feature map is evenly divided into 4 sub-feature maps along the channel dimension;
[0057] 2. Hierarchical Convolution: Except for the first sub-feature map, the nth sub-feature map is element-wise added to the output of the (n-1)th sub-feature map, followed by 3×3 groups of convolutions (each group only processes local channels, thereby reducing the number of parameters and computation, and enhancing the diversity of features, but may lead to information isolation between channels), and then a 1×1 convolution is added afterward to restore the channel interaction capability, thus forming a "local feature progressive enhancement" structure; n=2,3,4;
[0058] 3. Feature Fusion: Features from all branches are concatenated to generate a high-order representation with multi-scale fusion. A 1×1 convolution is used to perform cross-channel interaction and fusion of the concatenated features, enhancing their expressive power while maintaining consistency between the output channel number and the module settings, ensuring compatibility with the input of the residual connections.
[0059] Input features Figure X When passing through the Res2NetBasicBlock module, 1×1 convolution and batch normalization are first used to achieve channel alignment while maintaining spatial dimension invariance (H×W) (input / output dimension matching and residual connection compatibility), as shown in the following formula (assuming batch size is B and input channel number is C). in The number of output channels is C out (Spatial dimension is H×W, number of groups s=4 is the scale decomposition number):
[0060] Y = BN(Conv) 1×1 (x));
[0061] Where BN stands for Batch Normalization, used to stabilize inter-layer output and the network training process; Conv1×1 is a 1×1 convolution, which converts the input channels from C... in Adjust to C out Keep the spatial dimension H×W unchanged.
[0062] The feature map is then uniformly segmented along the channel dimension. This step still preserves the complete spatial information (H×W) of each sub-feature, as shown in the formula below:
[0063]
[0064] After the above operations, the input feature map has been divided into four channel sub-maps while maintaining the same spatial dimension. At this point, a progressive feature fusion mechanism is used to gradually aggregate local contextual information from adjacent scales. The formula is shown below:
[0065] Y1 = X1;
[0066]
[0067] in, GroupedConv represents the addition of features. 3×3 It's a grouped convolution, where channels are grouped and then independently convolved with 3×3 to reduce computational cost, and then processed by Conv... 1×1 The ability to interact with channels is restored. Y2 integrates Y1 and X2, corresponding to fine-grained texture (cloud edges), while Y4 accumulates information from the first three layers, corresponding to large-scale structure (overall cloud distribution). The entire process embodies hierarchical information integration from local to global, breaking down high-dimensional features into multiple low-dimensional sub-features, enabling the network to learn multi-scale features in parallel.
[0068] Subsequently, Y1, Y2, Y3, and Y4 are aggregated into a unified space through feature concatenation, preserving unique information at each scale while eliminating scale bias. The formula is shown below:
[0069]
[0070] Here, Concat is used for feature concatenation, and 1×1 convolution is used to perform cross-channel interaction and fusion on the concatenated features, thereby enhancing the expressive power of the features and adjusting the number of channels.
[0071] Finally, a residual connection is introduced and linked to the main path output to ensure that the network retains at least the original input performance (solving the degradation problem). The formula is shown below:
[0072] Output=ReLU(BN(Z)+Proj(X));
[0073] Where Proj(X) represents the residual connectivity feature.
[0074] Step 3.2: Construct the ViTAttention module. See the specific structure below. Figure 3 This module is inserted between feature extraction layers to enhance the model's spatial correlation modeling capabilities through a global attention mechanism. Specific implementation includes:
[0075] 1. Location Encoding: Learnable location encoding is fused with input features through bilinear interpolation, preserving prior knowledge of the two-dimensional spatial structure;
[0076] 2. Multi-head attention: A global attention mechanism is introduced to calculate the global pixel correlation matrix, and a progressive multi-head mechanism is used. The module will divide the feature map into multiple independent subspaces according to the progressive head number (1-2-2-4) from the channel dimension.
[0077] 3. Feature Reconstruction: Multi-head attention outputs are fused through 1×1 convolution to preserve important spatial response regions.
[0078] When the input data feature map passes through the ViTAttention module, spatial location information is first added to the feature map to compensate for the lack of explicit location awareness in traditional CNNs. The method used is bilinear interpolation. The formula is shown below:
[0079]
[0080] Among them, X pos X is the encoded feature map, and P is the input feature map; P is the learnable parameter with an initial dimension of 1. Extend to the target dimension through interpolation; Represents a bilinear interpolation function;
[0081] The location-encoded features are then mapped to the query, key, and value semantic spaces, forming the basis for self-attention computation. The formula is shown below:
[0082]
[0083] in, W represents the convolution operation. qkv For weight projection, Chunk means efficiently mapping the output of a single convolution to the three semantic spaces Q / K / V.
[0084] Subsequently, attention weights are constructed using query-key similarity metrics to establish a correlation between model learning and the spatial location of the input feature map. If there is an input sequence... (L = H × W is the sequence length, D is the number of channels), and the calculation process is as follows:
[0085]
[0086] Where Q = XW Q K = XW K V = XW V The model uses a linear projection. It employs a progressive head configuration (1-2-2-4) to retain more attention heads during the deep decoding phase to capture global semantic patterns.
[0087] Then, attention weights are used to multiply the value matrices to achieve adaptive fusion of contextual information. The formula is shown below:
[0088]
[0089] Among them, O attn N represents the attention output. h HW represents the attention head dimension (number), HW represents the spatial location dimension, i.e., the length of the input feature map after flattening out its spatial resolution, and D represents the attention head dimension. h Denotes the subspace dimension processed by each attention head, in D h In 3D space, the eigenvector at each location is a convex combination of its eigenvectors and the eigenvectors at all locations. The elements A in matrix A... i,j This indicates the strength of the correlation between position i and position j;
[0090] Finally, the multi-head attention results are fused back into the standard feature map format to maintain compatibility with the CNN architecture, allowing the module output to be used as inter-layer input and passed to the Res2NetBasicBlock module. The formula is shown below:
[0091] O reshape =Reshape(O attn );
[0092]
[0093] Among them, O reshape W represents the recombination feature map. proj For projective convolution kernel, O proj The final output feature map, This is a convolution operation.
[0094] Step 3.3: Construct the deep learning model proposed in this invention, the detailed structure of which is as follows: Figure 1 As shown, a deep learning model is built upon the implementation of the two modules mentioned above. The model's head network first uses 7x7 convolutions, batch normalization (BN), ReLU, and max pooling to downsample the input image and extract basic features, thereby reducing computational complexity and expanding the receptive field for subsequent use.
[0095] A sequential hierarchical structure is constructed after the head, feeding the input image into the Res2NetBasicBlock module. Multi-scale image feature information is extracted through parallel convolutions at the channel scale, and another identical module is then connected after this module. Stacking two Res2NetBasicBlock modules maintains parameter efficiency while achieving deep fusion of multi-scale features through a hierarchical residual structure, balancing computational efficiency and model expressiveness. This design balances network depth and complexity, enhancing gradient flow through dual residual paths and providing a foundation for efficient feature extraction. The ViTAttention module is then connected to this structure, and the two work together to establish a "local convolutional perception-global attention association" feature enhancement mechanism, capable of simultaneously modeling local texture anomalies and global intensity distribution patterns of cloud cover to identify complex and varied cloud cover patterns. The backbone of the model is formed by four stacked composite modules, with a downsampling operation (3*3 convolution) introduced after the first three composite modules.
[0096] Finally, the output layer network is built. The feature map is compressed to a 1×1 spatial dimension and the channel information is preserved by adaptive average pooling (AdaptiveAvgPool2d). After flattening, the 512-dimensional feature is mapped to the number of target categories by a fully connected layer (Linear). The final output is the probability of the image belonging to each category. The image will be classified into the category with the highest probability.
[0097] Step 4: Model Training and Hyperparameter Optimization. The model is trained using the training set, and then hyperparameters are optimized based on the model's performance on the validation set. Finally, the model that performs best on the test set is saved. The model training strategy is as follows:
[0098] Learning rate adjustment strategy: Use dynamic learning rate, with an initial learning rate set at 3×10. -4 The learning rate is updated based on the macro-average F1 score on the validation set. If the score does not improve for three consecutive epochs, the learning rate is decayed with a coefficient of 0.5, with a lower limit of 1×10. -6 ;
[0099] Loss function selection: Introducing label-smoothed cross-entropy loss and setting a smoothing coefficient λ, the theoretical lower bound expression of the loss function is established as follows:
[0100] Step 5: Input the full-surface solar image to be classified into the full-surface solar image quality classification model to obtain the full-surface solar image quality classification result.
[0101] According to a second aspect of the present invention, a deep learning-based full-displacement solar image quality classification system is provided, comprising: a first construction module for classifying acquired full-displacement solar images to obtain full-displacement solar images of multiple quality categories with balanced proportions, and constructing a dataset; a first acquisition module for sequentially scaling and center-cropping each image in the dataset to obtain a dataset of a preset size; randomly dividing the dataset of the preset size into a training set, a validation set, and a test set according to a preset ratio; a second construction module for constructing a deep learning model based on a concatenated head network, backbone network, and output layer network; a training and selection module for training and hyperparameter optimization of the deep learning model based on the training set and validation set, and saving the model that performs best on the test set as the full-displacement solar image quality classification model; and a second acquisition module for inputting the full-displacement solar image to be classified as the full-displacement solar image quality classification model to obtain the full-displacement solar image quality classification result. Each module in the above-mentioned deep learning-based full-displacement solar image quality classification system can be implemented entirely or partially through software, hardware, or a combination thereof. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or they can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above modules.
[0102] Example 2: Figures 1-10 As shown, the optional implementation steps of the deep learning-based full-disk solar image quality classification method proposed in this invention are described below, using real Hα-band full-disk solar images. Specifically, they include:
[0103] S1. A total of 4350 full-plane solar images taken by SFMM in October 2023 and May-June 2024 were selected as the dataset (image size is 4096×4096). Among them, there are 1450 clean solar images, 1450 solar images with slight cloud impact, and 1450 images with severe cloud impact. That is, there are 1450 images in each of the three categories of clean, relatively slight, and relatively severe. Figure 4 , Figure 5 , Figure 6 Samples of the three different types of images mentioned above are shown in turn.
[0104] To avoid underfitting because the model cannot fully learn the features of images of a certain category due to too little data of that category, the proportion of the three categories of images is set to 1:1:1 when constructing the dataset.
[0105] S2. Scale each image to 256×256 and crop it to 224×224 in the center. That is, the size of all data after processing is 224×224 when it is used as input to the network model. Then, randomly divide the dataset of size 224×224 into training set, validation set and test set according to the proportion of 70%, 20% and 10%. By randomly dividing it, the interference of time sequence on the data can be avoided.
[0106] The above technical solution performs a scaling operation on the original image. This operation scales the image from 4096×4096 to a specific 256×256. This operation reduces computation and speeds up training without affecting the accuracy of quality classification. Furthermore, a center cropping operation is introduced on top of the scaling to reduce interference from useless features.
[0107] S3. Construct a deep learning model based on the concatenated head network, backbone network, and output layer network;
[0108] S4. Model Training and Hyperparameter Optimization. The model is trained using the training set, then hyperparameters are optimized based on the model's performance on the validation set. Finally, the model that performs best on the test set is saved. The model training strategy is as follows:
[0109] Learning rate adjustment strategy: Use dynamic learning rate, with an initial learning rate set at 3×10. -4 The learning rate is updated based on the macro-average F1 score on the validation set. If the score does not improve for three consecutive epochs, the learning rate is decayed with a coefficient of 0.5, with a lower limit of 1×10. -6 .
[0110] Loss function selection: Introducing label smoothing cross-entropy loss to avoid the model's overconfidence in the training labels and to improve the model's performance in classification tasks, the formula is shown below:
[0111]
[0112] Where K represents the total number of categories of solar images across the entire solar surface; p i This represents the probability distribution of the model output, measuring the degree of match between the model's predictions and the smoothed labels; y i The probability distribution of the label is represented by 1-λ;
[0113] Set the smoothing coefficient λ = 0.1. According to the formula, the theoretical lower bound of the loss function for perfect prediction by the model is:
[0114]
[0115] The loss function graph of the deep learning model training proposed in this invention is shown below. Figure 7As shown, the loss function nearly converged after about 50 epochs of model training.
[0116] S5. Model Evaluation. The model's performance is measured using the classification results on the test set. Several key evaluation metrics commonly used in image classification are selected: Precision, Total Precision, Recall, F1 Score, and Macro F1 to demonstrate and evaluate the model's performance.
[0117] Figure 8 This demonstrates the performance of the best model on the validation set on the test set. For the clean solar image category (clear), the model's Precision, Recall, and F1 Score are 0.979, 0.966, and 0.972, respectively; for the solar image category with slight cloud cover (less), the model's performance on these three metrics is 0.947, 0.979, and 0.963; and for the solar image category with severe cloud cover (more), the model's performance on these three metrics is 1.000, 0.979, and 0.990. Further based on... Figure 8 The confusion matrix yielded a model with Total Precision and Macro F1 scores of 0.9747 and 0.975, respectively.
[0118] To evaluate the model's performance, UNet and ResNet were trained as comparative experiments, with their performance varying from [previous data point to previous data]. Figure 9 (UNet) and Figure 10 (ResNet) demonstration. The performance of UNet is as follows: for clean solar images, the model's Precision, Recall, and F1 Score are 0.960, 0.986, and 0.973, respectively; for solar images slightly affected by clouds, the model's performance on these three metrics is 0.952, 0.952, and 0.952; for solar images severely affected by clouds, the model's performance on these three metrics is 0.993, 0.966, and 0.979. Further based on... Figure 9 The confusion matrix yielded the model's Total Precision and Macro F1 score of 0.9678 and 0.968, respectively. The ResNet performance was as follows: for clean solar images, the model's Precision, Recall, and F1 Score were 0.953, 0.979, and 0.966, respectively; for solar images slightly affected by clouds, the model's performance on these three metrics was 0.951, 0.945, and 0.948; and for solar images severely affected by clouds, the model's performance on these three metrics was 0.972, 0.952, and 0.962. Further based on... Figure 10The confusion matrix yielded a Total Precision and Macro F1 score of 0.9586 and 0.959 for the model, respectively. Overall, the model of this invention has the best classification performance.
[0119] This result demonstrates that the deep learning model proposed in this invention performs excellent classification of Hα-band full-surface solar images captured by SFMM based on cloud cover. Its fully automated processing and high accuracy eliminate the need for subsequent manual screening. This model can then be used for quality classification of other Hα-band full-surface solar images to be classified.
[0120] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A deep learning-based method for classifying the quality of full-surface solar images, characterized in that, include: The acquired full-surface solar images are classified to obtain full-surface solar images of multiple quality categories with balanced proportions, and a dataset is constructed. Each image in the dataset is sequentially scaled and cropped to obtain the dataset at a preset size. The dataset of a preset size is randomly divided into a training set, a validation set, and a test set according to a preset ratio; A deep learning model is constructed based on the concatenated head network, backbone network, and output layer network. The deep learning model was trained and its hyperparameters optimized based on the training and validation sets, and the model that performed best on the test set was saved as the full-surface solar image quality classification model. The full-plane solar image to be classified is used as input to the full-plane solar image quality classification model to obtain the full-plane solar image quality classification result.
2. The deep learning-based full-surface solar image quality classification method according to claim 1, characterized in that, The acquired full-surface solar images were categorized into three quality classes based on the degree of cloud cover: clean, relatively light, and relatively severe.
3. The deep learning-based full-surface solar image quality classification method according to claim 1, characterized in that, The deep learning model is specifically as follows: Build the Res2NetBasicBlock module; Construct the ViTAttention module; A composite module is constructed based on the Res2NetBasicBlock module and the ViTAttention module; Convolution, Batch Normalization (BN), ReLU, and max pooling are used as the head network of the deep learning model. The backbone network consists of four concatenated composite modules, with downsampling operations introduced after the first three composite modules. The output layer network is built after the fourth composite module. The output layer network obtains the output through adaptive average pooling, flattening, and fully connected layer operations.
4. The deep learning-based full-surface solar image quality classification method according to claim 3, characterized in that, The composite module is constructed by using two stacked Res2NetBasicBlock modules and a ViTAttention module connected in series with the two stacked Res2NetBasicBlock modules.
5. The deep learning-based full-surface solar image quality classification method according to claim 3, characterized in that, The Res2NetBasicBlock module includes: Channel decomposition: The input feature map is evenly divided into 4 sub-feature maps along the channel dimension; Hierarchical convolution: Except for the first sub-feature map, the nth sub-feature map is added element-wise to the output of the (n-1)th sub-feature map, followed by a 3×3 group convolution, and then a 1×1 convolution is added to restore the channel interaction capability to obtain the output of the nth sub-feature map; n=2,3,4; Feature fusion: Features from all branches of the hierarchical convolution are concatenated and merged to generate a high-order representation with multi-scale fusion; then, the concatenated features are interacted and fused across channels through convolution and BN to serve as the main path output; finally, residual connections are introduced to access the main path output.
6. The deep learning-based full-surface solar image quality classification method according to claim 3, characterized in that, The ViTAttention module includes: Location encoding: Learnable location encoding is fused with input features through bilinear interpolation, preserving prior knowledge of two-dimensional spatial structure; Multi-head attention: A global attention mechanism is introduced to calculate the global pixel correlation matrix, and a progressive multi-head mechanism is used. The module will divide the feature map into multiple independent subspaces according to the number of progressive heads from the channel dimension. Feature reconstruction: The multi-head attention output is fused through convolution to preserve important spatial response regions.
7. The deep learning-based full-surface solar image quality classification method according to claim 1, characterized in that, The deep learning model training strategy is as follows: Learning rate adjustment strategy: Use dynamic learning rate, with an initial learning rate set at 3×10. -4 The learning rate is updated based on the macro-average F1 score on the validation set. If the score does not improve for three consecutive epochs, the learning rate is decayed with a coefficient of 0.5, with a lower limit of 1×10. -6 ; Loss function selection: Introducing label-smoothed cross-entropy loss and setting a smoothing coefficient λ, the theoretical lower bound expression of the loss function is established as follows:
8. A deep learning-based full-surface solar image quality classification system, characterized in that, include: The first construction module is used to classify the acquired full-surface solar images to obtain full-surface solar images of multiple quality categories with balanced proportions, and to construct a dataset; The first acquisition module is used to perform scaling and center cropping operations on each image in the dataset in sequence to obtain the dataset at a preset size. The dataset of a preset size is randomly divided into a training set, a validation set, and a test set according to a preset ratio; The second building module is used to construct a deep learning model based on the concatenated head network, backbone network, and output layer network. The training and selection module is used to train the deep learning model and optimize its hyperparameters based on the training and validation sets, and save the model that performs best on the test set as the full-surface solar image quality classification model. The second acquisition module is used to take the full-surface solar image to be classified as input to the full-surface solar image quality classification model and obtain the full-surface solar image quality classification result.