Rapid flow axis automatic extraction method based on cross pseudo-supervision and semi-supervision
Through the cross-pseudo-supervised semi-supervised deep learning method and the eight neighborhood connection algorithm for the rapids center axis point, the problems of low efficiency and poor generalization capabilities of the traditional rapids shaft extraction method are solved, and efficient and accurate automatic extraction of rapids shafts is achieved, suitable for meteorological monitoring and aviation safety.
Patent Information
- Application Number
- CN202411818362.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the traditional manual drawing method of rapid axes is low efficiency, has large errors and is subjective. The existing automation methods lack generalization capabilities under complex wind farm conditions, and the full supervision learning depends too heavily on the labeled data, making it difficult to meet the needs of efficient and accurate automatic extraction.
A deep learning method based on cross-pseudo-supervised semi-supervised is adopted, combined with Swin-Unet network and consistency learning, and a small amount of labeled data and a large amount of unlabeled data are used to generate pseudo-labels. The model is trained through the cross-pseudo-supervised module, and the automatic extraction of the rapid axis is combined with the eight neighborhood connection algorithm of the rapid center axis point.
It improves the accuracy and robustness of the rapid axes identification, reduces the dependence on labeled data, and realizes efficient, low-cost and safe automated identification. It is suitable for meteorological monitoring and aviation safety, and supports real-time detection and dynamic updates.
Smart Images

Figure CN120339784A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of meteorological technology, and particularly relates to a method for automatically extracting jet axes based on cross pseudo-supervision semi-supervision. Background Art
[0002] Atmospheric jets refer to narrow strong wind belts with high wind speeds in the atmosphere, usually appearing in the middle and upper layers of the atmosphere, and have an important impact on weather systems and surface weather. In weather forecasting, the jet axis is a key indicator used to represent the center line of the jet, which helps meteorologists more intuitively understand the weather situation and make predictions. The traditional drawing of the jet axis mainly relies on manual operation, and forecasters manually draw the jet axis through the MICAPS system. This manual method is not only inefficient but also easily affected by human subjective factors, resulting in large errors. To improve the automation level, researchers have tried to automatically extract the jet axis from radar data and grid wind field data. Existing methods mainly rely on mathematical models and algorithms, such as merging algorithms, polynomial fitting, key point detection, and threshold methods. These methods reduce subjectivity to a certain extent and improve the automation level, but they are still difficult to achieve high-precision extraction when facing complex atmospheric wind fields, especially in the case of jet bifurcation or merger. In addition, the stability and generalization ability of these methods under different wind field conditions are limited.
[0003] In recent years, deep learning technology has gradually been applied to the meteorological field, especially in the task of automatically extracting jet axes, due to its strong performance in image recognition and pattern classification. Full supervision learning requires a large amount of high-quality labeled data, which makes it have great limitations in practical applications. To overcome the data dependence problem, semi-supervised learning has become an effective solution, which can use a large amount of unlabeled data for model training with less labeled data, improving the generalization ability of the model.
[0004] With the development of semi-supervised learning methods, consistency learning and cross pseudo-supervision techniques have gradually attracted attention. By adding perturbations to unlabeled data and comparing it with the unperturbed data, the robustness and accuracy of the model can be improved. However, existing pseudo-supervision methods usually only use pseudo-labels for auxiliary supervision and do not effectively incorporate them into the training process, limiting the further improvement of model performance. How to make full use of unlabeled data to improve the accuracy and efficiency of automatic jet axis extraction has become the key direction of current technological development. Summary of the Invention
[0005] In view of the above deficiencies in the prior art, the present invention provides a method for automatically extracting jet axes based on cross-pseudo-supervised semi-supervised learning. By combining semi-supervised learning and an eight-neighborhood connection algorithm based on jet center axis points, the present invention solves the problems of low efficiency and poor generalization ability of traditional jet axis extraction methods, while reducing the dependence on labeled data and improving the accuracy and robustness of jet axis recognition.
[0006] To achieve the above object, the technical solution adopted by the present invention is: a method for automatically extracting jet axes based on cross-pseudo-supervised semi-supervised learning, comprising the following steps:
[0007] S1. Preprocess the acquired 11 types of grid wind field data and perform label mask conversion on the manually labeled data;
[0008] S2. Construct a cross-pseudo-supervised semi-supervised deep learning model;
[0009] S3. Use the preprocessing results and label mask conversion results in step S1 to train and evaluate the cross-pseudo-supervised semi-supervised deep learning model, and obtain an image segmentation result according to the evaluated cross-pseudo-supervised semi-supervised deep learning model;
[0010] S4. According to the image segmentation result, use the eight-neighborhood connection method based on jet center axis points to extract the jet axis.
[0011] Further, the S1 includes the following steps:
[0012] S101. Acquire 11 types of grid wind field data;
[0013] S102. Calculate the wind speed and wind direction of each grid point. Among them, according to the quadrant of the wind speed, adjust to determine the accuracy of the direction angle;
[0014] S103. Map the calculated wind speed and wind direction to RGB three-channel data to obtain wind field image coding data, and complete the preprocessing of the wind field processing data. Among them, the R-channel data represents the magnitude of the wind speed mapped to the numerical range of the R-channel of the color picture through coding, and the G-channel data and B-channel data respectively represent the wind direction mapped to the numerical ranges of the G-channel and B-channel of the color picture through coding. The wind field image coding data includes labeled data and unlabeled data;
[0015] S104. Based on the manually labeled data, read the label mask;
[0016] S105. Extract all the connected target area lines in the label mask to obtain the contour of the label mask;
[0017] S106. Randomly screen and transform the contour of the label mask to obtain the retained contour;
[0018] S107. Redraw a new label mask based on the remaining contour, save the new label mask, and complete the label mask conversion process.
[0019] Furthermore, the expression of the wind speed is as follows:
[0020]
[0021] where Speed represents the wind speed, and both U and V represent wind speed components;
[0022] The expression of the wind direction is as follows:
[0023]
[0024] where Direction represents the wind direction.
[0025] Furthermore, mapping the calculated wind speed and wind direction to RGB three-channel data specifically means: using the following formula to map the calculated wind speed and wind direction to RGB three-channel data:
[0026]
[0027] where R (i,j) represents the value of the R channel at the pixel (i, j) position of the color picture mapped by encoding the wind speed magnitude at the position (i, j) of the MICAPS11-class grid wind field data, G (i,j) and B (i,j) respectively represent the values of the G channel and B channel at the pixel (i, j) position of the color picture mapped by encoding the wind direction at the position (i, j) of the MICAPS11-class grid wind field data, Speed (i,j) represents the wind speed magnitude at the position (i, j) of the MICAPS11-class grid wind field data, and Direction (i,j) represents the wind direction at the position (i, j) of the MICAPS11-class grid wind field data.
[0028] Furthermore, the cross pseudo-supervised semi-supervised deep learning model includes:
[0029] The first Swin-Unet network is used to convert the preprocessing result in step S1 and the label mask conversion result into a predicted first jet region mask;
[0030] The second Swin-Unet network is used to convert the preprocessing result in step S1 and the label mask conversion result into a predicted second jet region mask, where the first Swin-Unet network and the second Swin-Unet network adopt different initialization processes;
[0031] The cross pseudo-supervision module is used to obtain the image segmentation result according to the first jet region mask and the second jet region mask by generating the first pseudo-label and the second pseudo-label that mutually serve as supervision signals. Each batch contains both labeled data and unlabeled data. In each training batch, the cross pseudo-supervision semi-supervised deep learning model updates the weights using the labeled data and simultaneously guides the learning of the unlabeled data through the pseudo-labels.
[0032] Furthermore, the structures of the first Swin-Unet network and the second Swin-Unet network are the same, and both include:
[0033] An encoder, which is used to use multiple Swin Transformer modules to divide the preprocessing result of step S1 and the label mask conversion result into non-overlapping windows to extract deep features, where the self-attention mechanism is applied within each window;
[0034] A bridging module, which is used to connect the encoder and the decoder;
[0035] A decoder, which is used to restore the extracted deep features to the same resolution as the input image through layer-by-layer upsampling, where the input image is the preprocessing result of step S1 and the label mask conversion result;
[0036] An output layer, which is used to use a convolutional layer to convert the image output by the decoder into a predicted jet region mask, where the jet region mask includes a first jet region mask and a second jet region mask.
[0037] Furthermore, the expression of the supervised learning loss function of the cross pseudo-supervision semi-supervised deep learning model is as follows:
[0038] L Sup = λL CE + μL Dice
[0039]
[0040] where L Sup represents the supervised learning loss function of the cross pseudo-supervision semi-supervised deep learning model, both λ and μ represent weight coefficients, L CE represents the cross-entropy loss function, L Dice represents the dice loss function, C represents the total number of categories, y i represents the true label, represents the probability of predicting to belong to category i.
[0041] Furthermore, S4 includes the following steps:
[0042] S401. Based on the image segmentation result, find the starting pixel point, where there is only one neighborhood with a color value within the eight-neighborhood around the starting pixel point;
[0043] S402. Traverse the eight-neighborhood of the starting pixel point, preferentially select the adjacent point that is closest to the current point and has not been traversed in a certain direction, and mark it as visited;
[0044] S403. Taking the starting pixel point as the center, determine whether all adjacent points have been traversed. If so, add each adjacent point to the set of central points of the connected jet axes in sequence, and proceed to step S404; otherwise, return to step S402;
[0045] S404. Based on the set of central points of the jet axes, draw the jet lines to complete the extraction of the jet axes.
[0046] Advantages of the present invention:
[0047] (1) Compared with traditional fully supervised learning or manual processing methods, the present invention combines semi-supervised learning and an eight-neighborhood connection algorithm based on the central points of the jet axis to solve the problems of low efficiency and poor generalization ability of traditional jet axis extraction methods, while reducing the dependence on labeled data and improving the accuracy and robustness of jet axis recognition.
[0048] (2) In terms of data security, the present invention integrates the jet axis recognition process and pseudo-label generation entirely within a local deep learning framework, reducing the dependence on cloud computing and data transmission, and avoiding the risk of sensitive data leakage. Since important meteorological data such as high-altitude wind fields are involved in the training process, using local training can ensure that the data does not leave the local environment, improving the data security of the system.
[0049] (3) In semi-supervised learning, the present invention makes full use of a large amount of unlabeled data, and these data do not need to be labeled, transmitted, or manually sorted. It avoids the risk brought by the leakage of traditional meteorological labeled data, further protecting the privacy of the data.
[0050] (4) The present invention combines unlabeled data with a small amount of labeled data through a cross-pseudo-supervision mechanism, greatly reducing the dependence on expensive labeled data. Compared with traditional fully supervised models, it reduces the manual labeling cost and the transmission volume of labeled data, and reduces the demand for labeled data.
[0051] (5) The present invention uses the Swin-Unet network structure to achieve efficient computing through hierarchical convolution and self-attention mechanisms. The Swin-Unet network supports running on edge devices with limited resources, avoiding the high consumption of cloud computing resources, and realizing a lightweight architecture and distributed computing support.
[0052] (6) During the training process, the present invention designs a dynamic consistency loss function, gradually adjusts the consistency weight according to the training progress, reduces unnecessary iterations, improves the training efficiency, thereby reducing the consumption of computing resources and time, and realizes the optimization of dynamic consistency loss.
[0053] (7) The present invention uses two Swin-Unet networks with different initializations to generate pseudo-labels for each other, and improves the generalization ability of the model in complex atmospheric wind fields through pseudo-supervised training. This strategy can effectively reduce the influence of low-quality pseudo-labels, avoid the model falling into local optimal solutions, and realizes the improvement of the model's robustness through cross pseudo-supervision.
[0054] (8) The present invention obtains the jet axis through skeleton extraction technology, and uses semi-supervised learning strategies and a large amount of unlabeled data to significantly improve the generalization ability of the model and the accuracy of jet axis extraction.
[0055] (9) Through key technologies such as cross pseudo-supervision and dynamic consistency optimization, the present invention realizes an efficient, low-cost, and safe jet axis identification scheme. Compared with traditional fully supervised learning models and manual identification methods, the present invention has significant advantages in terms of data security, network resource savings, real-time performance, and generalization ability. These features enable the present invention to achieve more efficient applications in complex atmospheric environments, provide reliable technical support for meteorological monitoring and aviation safety, and in terms of reducing manual intervention in automatic identification, the jet axis identification system of the present invention realizes a fully automated process from data input to model output, without manual participation, and significantly reduces the time of manual intervention. In terms of real-time processing capabilities, the cross pseudo-supervised semi-supervised deep learning model submitted by the present invention can be integrated into the real-time monitoring system of meteorological stations to ensure the timely detection and dynamic update of jet axes, and provide real-time support for aviation flight, weather forecasting, and emergency response. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a flowchart of the method of the present invention.
[0057] Figure 2 is a flow block diagram of the present invention.
[0058] Figure 3 is a block diagram of the cross pseudo-supervised semi-supervised deep learning model in this embodiment.
[0059] Figure 4 is a schematic diagram of extracting the jet axis by using the eight-neighborhood connection method of the jet center axis point in this embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0060] The specific embodiments of the present invention will be described below to facilitate the understanding of those skilled in the art of the present technology. It should be clear, however, that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0061] Embodiment
[0062] Before explaining the present invention, the following terms will be explained:
[0063] In the Swin-Unet network, "Swin" refers to the Swin Transformer, which is a Shifted Window. Swin is the abbreviation of Shifted Window; the Swin Transformer refers to a Transformer model with a Shifted Window; the Transformer is a neural network architecture; therefore, the Swin-Unet network is a network model that uses the Swin Transformer to replace some of the encoding (encoder) and decoding (decoder) structures in Unet.
[0064] Currently in the meteorological field, the manual drawing method of the jet axis is inefficient, has large errors and strong subjectivity, and is difficult to meet the requirements of high efficiency and accuracy. Although some automated methods have tried to extract the jet axis from radar data and wind field data, traditional methods based on mathematical and mechanism analysis have problems of insufficient generalization ability and low recognition accuracy when dealing with complex wind fields, especially in the scenarios of jet divergence and merger, and the effect is poor. In addition, the existing fully supervised learning models have a large demand for high-quality labeled data, which limits their wide promotion in practical applications. Therefore, how to use deep learning models to improve the automation level, accuracy and generalization ability of jet axis extraction, while reducing the dependence on a large amount of labeled data, is the main problem faced by the current technology field.
[0065] The present invention proposes a semi-supervised deep learning model for automatically extracting jet axes based on the cross pseudo-supervision method, aiming to achieve efficient, automated, and accurate extraction of jet axes through improved semi-supervised deep learning techniques. The model improves the robustness and generalization ability of the model by introducing the cross pseudo-supervision (CPS) strategy and generating pseudo-labels using a small amount of labeled data and a large amount of unlabeled data. Specifically, the present invention uses Swin-Unet as the backbone network, combines the consistency learning method with the eight-neighborhood connection algorithm based on the jet center axis points, to ensure that the semi-supervised deep learning model for automatically extracting jet axes can effectively process complex wind field scenarios in jet region recognition. Through this model, accurate extraction of jet axes can be achieved with reduced dependence on labeled data, improving the overall efficiency and recognition effect, and it is widely applicable to jet detection and prediction tasks in the meteorological field.
[0066] In this embodiment, the cross pseudo-supervision (CPS) strategy uses two Swin-Unet networks (networks with different initializations). Even if the inputs are the same, due to different initialization methods (such as Kaiming and Xavier initializations), the two models will produce different prediction results. This structure lays the foundation for consistency learning: the prediction results of different models on the same data should be consistent. In the cross pseudo-supervision (CPS) strategy, the prediction results of the first Swin-Unet network are used as pseudo-labels to supervise the second Swin-Unet network; similarly, the prediction results of the second Swin-Unet network are used as pseudo-labels to supervise the first Swin-Unet network. This way of interactive training is the embodiment of consistency learning: by constraining the learning directions of the two Swin-Unet networks with pseudo-labels, their predictions tend to be consistent, even if some data has no true labels.
[0067] As Figure 1 、 Figure 2 and Figure 3 shown, the present invention provides a method for automatically extracting jet axes based on cross pseudo-supervision semi-supervision, and its implementation method is as follows:
[0068] S1. Preprocess the obtained 11 types of grid wind field data, and perform label mask conversion on the manually labeled data. The implementation method is as follows:
[0069] S101. Obtain 11 types of grid wind field data;
[0070] S102. Calculate the wind speed and wind direction of each grid point. Among them, adjust according to the quadrant of the wind speed to determine the accuracy of the direction angle;
[0071] S103. Map the calculated wind speed and wind direction to RGB three-channel data to obtain the encoded data of the wind field image, completing the preprocessing of the wind field processing data. Among them, the data of the R channel represents the magnitude of the wind speed mapped to the numerical range of the R channel of the color picture through encoding, and the data of the G channel and the B channel respectively represent the wind direction mapped to the numerical ranges of the G channel and the B channel of the color picture. The encoded data of the wind field image includes labeled data and unlabeled data;
[0072] S104. Based on the manually labeled data, read the label mask;
[0073] S105. Extract all the connected target area lines in the label mask to obtain the contour of the label mask;
[0074] S106. Randomly screen and transform the contour of the label mask to obtain the reserved contour;
[0075] S107. According to the reserved contour, redraw and generate a new label mask, and save the new label mask to complete the label mask conversion process.
[0076] In this embodiment, the original data collected and prepared comes from 11 types of grid wind field data in the MICAPS (Meteorological Information Comprehensive Analysis and Processing System) system. The 11 types of grid wind field data include longitude, latitude, U-direction wind speed component, and V-direction wind speed component, representing the vector wind speed information of each point in the atmospheric wind field. By processing these data, wind speed and wind direction information can be extracted as the basic features of the model input.
[0077] In this embodiment, wind field vector decomposition: To convert the vector wind field data into an image data format that can be used by the deep learning model, it is first necessary to calculate the wind speed and wind direction of each grid point. The wind speed (Speed) is calculated by taking the square root of the sum of the squares of the two wind speed components U and V, and the formula is shown in Equation (1):
[0078]
[0079] Among them, Speed represents the wind speed, and both U and V represent wind speed components.
[0080] In this embodiment, the wind direction Direction is calculated by the arctangent function (arctan) of U and V, and Equation (2) is adjusted according to the quadrant of the wind speed to ensure the correctness of the direction angle:
[0081]
[0082] In this embodiment, for the RGB channel conversion: the calculated wind speed and wind direction need to be mapped to RGB three-channel data for use as the input to the deep learning model. The specific mapping formulas are shown in (3), (4), and (5).
[0083] Channel R: represents the magnitude of the wind speed. For grid points with a wind speed value less than 14 m / s, the corresponding value of channel R is set to 0; for grid points with a wind speed value greater than or equal to 14 m / s, the wind speed value is mapped to a pixel value in the range [0, 255] through the following function:
[0084]
[0085] Channels G and B: represent the wind direction. The wind direction only indicates the direction angle and has no magnitude meaning, so it cannot be directly used as an input feature of the model. Therefore, the cosine and sine functions are used to convert the wind direction into coordinates (x, y) on a two-dimensional plane, and then these coordinates are converted into pixel values of channels G and B. The conversion formulas are as follows:
[0086]
[0087] where R (i,j) represents the value of channel R at the pixel position (i, j) of the color picture obtained by encoding the wind speed magnitude at the position (i, j) of the MICAPS-11 grid point wind field data, G (i,j) and B (i,j) respectively represent the values of channel G and channel B at the pixel position (i, j) of the color picture obtained by encoding the wind direction at the position (i, j) of the MICAPS-11 grid point wind field data, Speed (i,j) represents the wind speed magnitude at the position (i, j) of the MICAPS-11 grid point wind field data, and Direction (i,j) represents the wind direction at the position (i, j) of the MICAPS-11 grid point wind field data.
[0088] In this embodiment, the wind speed and wind direction are calculated through formulas. The wind speed data is mapped to the red channel (channel R), and the two components of the wind direction are mapped to the green (channel G) and blue (channel B) channels, forming a standard RGB image data format as the input to the cross pseudo-supervised semi-supervised deep learning model.
[0089] In this embodiment, for the label mask conversion: in an image segmentation task, a label mask is a two-dimensional matrix image used to label different categories or feature regions. It assigns a category label to each pixel so that the cross pseudo-supervised semi-supervised deep learning model can learn the correspondence between features and target regions. In this processing, the present invention performs the following conversion on the label mask:
[0090] Reading the label mask: First, read the label mask as a binary image or a multi-class mask image, where each pixel value corresponds to a label for a different class or region. The white region represents the target of interest, and the black region represents the background.
[0091] Contour detection of the label mask: Use a contour detection algorithm to extract all the connected target region lines in the label mask. In this step, different connected regions in the mask are parsed as independent labels.
[0092] Random screening and transformation: To reduce the number of targets in the label mask or perform downsampling according to specific requirements, the present invention randomly samples the extracted contours.
[0093] Generating a new label mask: According to the retained contours, the present invention redraws and generates a new label mask. Using drawing methods such as drawContours to draw the randomly retained regions on the new mask to obtain a result that is simplified compared to the original mask.
[0094] Visualization and saving: The transformed label mask is saved in PNG format for use in the training and evaluation of the cross pseudo-supervised semi-supervised deep learning model.
[0095] S2. Constructing a cross pseudo-supervised semi-supervised deep learning model; this cross pseudo-supervised semi-supervised deep learning model includes:
[0096] The first Swin-Unet network, which is used to convert the preprocessing result of step S1 and the label mask conversion result into a predicted first jet region mask;
[0097] The second Swin-Unet network, which is used to convert the preprocessing result of step S1 and the label mask conversion result into a predicted second jet region mask, where the first Swin-Unet network and the second Swin-Unet network adopt different initialization processes;
[0098] The cross pseudo-supervised module, which is used to obtain the image segmentation result according to the first jet region mask and the second jet region mask by generating the first pseudo-label and the second pseudo-label that mutually serve as supervision signals. Each batch contains both labeled data and unlabeled data, and in each training batch, the cross pseudo-supervised semi-supervised deep learning model updates the weights using the labeled data and guides the learning of the unlabeled data through the pseudo-labels.
[0099] In this embodiment, a batch refers to the amount of data processed in one update (training iteration) of the cross pseudo-supervised semi-supervised deep learning model. Labeled data: Jet axis data containing true labels, which is used to directly calculate the loss. Unlabeled data: Without true labels, the pseudo-labels generated by another Swin-Unet network are used as supervision signals.
[0100] The structures of the first Swin-Unet network and the second Swin-Unet network are the same, both including:
[0101] An encoder, which uses multiple Swin Transformer modules to divide the preprocessing result of step S1 and the converted result of the label mask into non-overlapping windows to extract deep features, and within each window, a self-attention mechanism is applied;
[0102] A bridging module, which is used to connect the encoder and the decoder;
[0103] A decoder, which restores the extracted deep features to the same resolution as the input image through layer-by-layer upsampling, where the input image is the preprocessing result of step S1 and the converted result of the label mask;
[0104] An output layer, which uses a convolutional layer to convert the image output by the decoder into a predicted jet region mask, where the jet region mask includes a first jet region mask and a second jet region mask.
[0105] In this embodiment, as Figure 2 shown, a cross-pseudo-supervised semi-supervised deep learning model is constructed using the deep learning framework PyTorch. This cross-pseudo-supervised semi-supervised deep learning model includes two Swin-Unet networks with different network initializations and a cross-pseudo-supervision (CPS) part. Both of the two Swin-Unet networks with different initializations have an encoder, a bridging module, a decoder, and an output layer, and the inputs of the two Swin-Unet networks are both the preprocessing result of step S1 and the converted result of the label mask.
[0106] In this embodiment, the two Swin-Unet networks with different network initializations are used to generate intermediate results of labeled data and intermediate results of unlabeled data for cross-pseudo-supervision by the cross-pseudo-supervision CPS.
[0107] In this embodiment, the Swin-Unet network is a deep learning model that combines the Swin Transformer module and the U-Net architecture, mainly used for image segmentation tasks. The Swin Transformer module can better handle complex feature segmentation tasks by leveraging the long-range dependence capture ability of the Transformer and the encoding-decoding structure of the U-Net.
[0108] In this embodiment, the structure of the Swin-Unet network includes the following parts:
[0109] Encoder part: The encoder part is responsible for downsampling the input high-resolution image layer by layer to extract features. Different from the traditional U-Net that uses a Convolutional Neural Network (CNN) for feature extraction, the Swin-Unet network adopts the Swin Transformer module. The Swin Transformer module divides the input image into non-overlapping windows, and applies the self-attention mechanism inside each window. This localized self-attention can effectively reduce the computational cost while retaining local information. As the number of layers increases, the size of the window also gradually expands, so as to capture global dependencies at a higher level.
[0110] Bridging part: The bridging part in the middle of the Swin-Unet network plays a role in connecting the encoder and the decoder. In this part, the features (features represent the numerical representation of the input image, usually the feature maps formed after passing through multiple convolutional layers. These feature maps capture important information in the image, such as edges, textures, shapes, and other important features) are further processed through multiple Swin Transformer modules to ensure that the cross-pseudo-supervised semi-supervised deep learning model can capture the long-range dependencies in the input image while improving the expression ability of the network.
[0111] Decoder part: The decoder part restores the deep features extracted by the encoder to the same resolution as the original image layer by layer, gradually recovering the spatial information. Each layer in the decoder combines the output of the corresponding encoder layer to form skip connections, which can ensure that the cross-pseudo-supervised semi-supervised deep learning model retains sufficient detail information during the upsampling process. Through these skip connections, the decoder part can effectively restore the spatial structure of the jet region.
[0112] Output part: The last layer of the Swin-Unet network uses a convolutional layer to convert the output of the decoder into a predicted jet region mask. The output image, that is, the jet region mask, and each pixel value in the jet region mask represents whether the corresponding pixel belongs to the jet region, usually represented by a binary value: for example, a value of 1 (or 255) indicates that the pixel belongs to the jet region, and a value of 0 indicates that the pixel does not belong to the jet region. Each pixel of this output image represents whether the pixel belongs to the jet region, which provides a basis for subsequent cross-pseudo-supervision.
[0113] In this embodiment, the idea of Cross Pseudo Supervision (CPS) is as follows: Through two Swin-Unet networks with the same structure but different initializations, they respectively generate pseudo labels and use each other as supervision signals. Each Swin-Unet network not only learns to process labeled data but also relies on the pseudo labels of another Swin-Unet network to supervise the learning of unlabeled data, thereby improving the generalization performance.
[0114] It calculates the loss function value by jointly using the output segmentation results and labels of two Swin-Unet networks with different initializations. The loss function includes a supervised loss term and an unsupervised loss term. During the consistency learning process, a cross-pseudo-supervision structure is adopted. One loss term is constructed by "Output 1" and "Pseudo Label 2", and another loss term is constructed by "Output 2" and "Pseudo Label 1". Each batch contains labeled data and unlabeled data. For unlabeled data, the corresponding One-Hot pseudo label is obtained through the Softmax operation and then used as a supervision signal. For labeled data, the outputs of the two neural networks are constrained by the supervised loss. The pseudo labels output by cross-pseudo-supervision are all incorporated into the training set to enhance the generalization ability of the network.
[0115] In this embodiment, in supervised learning, the cross-entropy loss function and the dice loss function are two loss functions commonly used in semantic segmentation tasks. The cross-entropy loss function is used to measure the difference between the class probability distribution predicted by the model and the One-Hot encoding distribution of the true label. Its formula is:
[0116]
[0117] The dice loss function is based on the Dice Coefficient, and its goal is to maximize the overlap between the predicted region and the true region. The formula for the Dice Coefficient is:
[0118]
[0119] The cross-entropy loss function is suitable for processing pixel-level classification tasks, while the dice loss function helps to improve the regional overlap accuracy of the segmentation results, especially performing well on unbalanced datasets. However, using only the cross-entropy loss function or the dice loss function often cannot achieve the best results. Therefore, this patent combines these two loss functions as the loss function for supervised learning.
[0120] L Sup = λL CE + μL Dice (8)
[0121] Where L SupDenote the supervised learning loss function of the cross pseudo-supervised semi-supervised deep learning model. Both λ and μ denote weight coefficients, and L CE denotes the cross-entropy loss function, and L Dice denotes the Dice loss function. C represents the total number of categories, and y i denotes the true label, denotes the probability of predicting to belong to category i, and p i denotes the predicted probability, and p i denotes the true label.
[0122] S3. Use the preprocessing result of step S1 and the label mask conversion result to train and evaluate the cross pseudo-supervised semi-supervised deep learning model, and obtain the image segmentation result according to the evaluated cross pseudo-supervised semi-supervised deep learning model;
[0123] In this embodiment, read the data of the MICPAS 11 classes. The time span of the dataset is from 2019 to 2022. After wind speed vector decomposition and RGB channel conversion, there are 1000 pictures, including the dataset, labels and original data. 10% of the data is divided into the test set, and the remaining is divided into the training set and the validation set according to the ratio of 8:1. The cross pseudo-supervised semi-supervised deep learning model maintains data independence during the training process and effectively evaluates the generalization ability of the model. Use the training dataset to train the cross pseudo-supervised semi-supervised deep learning model. The cross pseudo-supervised semi-supervised deep learning model uses the samples of the test set to adjust the parameter weights during the training process to minimize the loss function.
[0124] In this embodiment, evaluate the cross pseudo-supervised semi-supervised deep learning model: Select the commonly used evaluation metrics in the field of image segmentation to evaluate the performance of the cross pseudo-supervised semi-supervised deep learning model: intersection over union (IoU), Dice coefficient (Dice) and precision (Precision, Pre). These metrics are used to measure the overlap degree, similarity and prediction accuracy between the prediction result and the actual annotation respectively. The intersection over union IoU, also known as the Jaccard or Jacquard coefficient, is the ratio of the intersection and union of the two sets of the true value and the predicted value, which measures the overlap degree between the predicted region and the true region, and mainly focuses on the proportion of the overlapping area between the prediction result and the true label relative to the combined area of the two. The Dice coefficient pays more attention to the similarity between the prediction result and the actual annotation by calculating the intersection size between them. The precision is used to evaluate the prediction accuracy of the cross pseudo-supervised semi-supervised deep learning model, that is, the proportion of the correctly predicted positive samples in all positive prediction samples.
[0125] In this embodiment, the cross pseudo-supervised semi-supervised deep learning model is mainly applied to the automatic identification of the jet axis in the atmospheric wind field, solving the problems of low efficiency, strong subjectivity and large errors in the traditional manual drawing method. Through the semi-supervised learning framework of deep learning, the model uses the cross pseudo-supervision (CPS) strategy to effectively improve the generalization ability and recognition accuracy of the cross pseudo-supervised semi-supervised deep learning model. This technology breaks through the limitations of traditional manual drawing and full-supervised learning in terms of time, manpower and data dependence, bringing profound application value to the fields of atmospheric science and aviation safety. Specific applications of the cross pseudo-supervised semi-supervised deep learning model:
[0126] Meteorological data analysis and jet monitoring: Quickly identify the location and intensity of atmospheric jets, provide accurate weather analysis assistance for forecasters, and improve the accuracy of weather forecasts. It is particularly suitable for analyzing the upper-air wind field and monitoring extreme weather related to jets, such as snowstorms and heavy precipitation.
[0127] Meteorological emergency response and aviation support: Identify the impact of the jet axis on aviation routes, optimize route planning, reduce aircraft buffeting and fuel consumption in strong wind belts. Provide decision-making support for meteorological emergency response and improve the early warning ability in bad weather.
[0128] High-precision weather prediction: The jet axis is an important part of the meteorological system and has an important impact on frontal movement, heavy rain and typhoon paths. The model can automatically extract the jet axis from the wind field, provide accurate upper-air wind field situation maps for meteorological forecasters, and significantly improve the accuracy of short-term and medium-term weather forecasts.
[0129] Automated meteorological system integration: The cross pseudo-supervised semi-supervised deep learning model can be integrated into existing meteorological data processing systems (such as MICAPS), replacing the traditional manual drawing process and improving the automation and real-time performance of data processing.
[0130] S4. According to the image segmentation result, use the eight-neighborhood connection method based on the jet center axis point to extract the jet axis. The implementation method is as follows:
[0131] S401. According to the image segmentation result, find the starting pixel point, where there is only one neighborhood with a color value in the eight-neighborhood around the starting pixel point;
[0132] S402. Traverse the eight-neighborhood of the starting pixel point, preferentially select the adjacent point closest to the current point and not yet traversed in the direction, and mark it as visited;
[0133] S403. Taking the starting pixel point as the center, judge whether all adjacent points have been traversed. If so, add each adjacent point to the connected jet axis center point set in turn and enter step S404. Otherwise, return to step S402;
[0134] S404. Draw the jet stream line based on the central point set of the jet stream axis to complete the extraction of the jet stream axis.
[0135] In this embodiment, as Figure 4 shown, when connecting and drawing the jet stream line for the central point set of the jet stream axis extracted from the skeleton, the default connection method from bottom to top will result in abnormal axes, that is, as shown in the green dashed box in the figure, which is the problem of scattered point connection within the red dashed box after rectangular coordinate transformation. To solve this problem, the present invention proposes an eight-neighborhood connection algorithm based on the central axis points of the jet stream to ensure that the scattered points in the image are connected in an orderly manner according to the spatial order, rather than randomly. The core of this algorithm is: first, find a suitable starting pixel point, that is, there is only one neighborhood with a color value within the eight neighborhoods around this pixel point, which can ensure the uniqueness and orderliness of the connection; after finding the starting point, the algorithm traverses the eight neighborhoods of the pixel points, preferentially selects the adjacent point closest to the current point and unvisited, and marks it as visited; then, continue to use this point as the center and repeat this process, adding each adjacent point to the connected point set in turn.
[0136] In this embodiment, to improve the algorithm efficiency, a direction array is used to represent eight possible neighborhoods (up, down, left, right, and four diagonal directions), so as to quickly access adjacent pixel points. The whole process traverses all pixel points through recursion and loop to ensure that the connection order of the points is consistent with their spatial distribution, effectively avoiding random connection. This method not only simplifies the connection process of the central axis points of the jet stream, but also ensures that the spatial structure remains consistent when converted to a general data format, solving the problem of random connection of scattered points in the image. The effect is as shown in the red solid box and the blue dashed box in the figure.
[0137] In summary, the present invention combines semi-supervised learning and an eight-neighborhood connection algorithm based on the central axis points of the jet stream to solve the problems of low efficiency and poor generalization ability of traditional jet stream axis extraction methods, while reducing the dependence on labeled data and improving the accuracy and robustness of jet stream axis recognition.
Claims
1. A method for automatically extracting jet axes based on cross pseudo-supervision semi-supervision, characterized in that, It includes the following steps: S1. Preprocess the obtained 11 types of grid wind field data, and perform label mask conversion on the manually annotated data; S2. Construct a cross pseudo-supervised semi-supervised deep learning model; S3. Use the preprocessing results and label mask conversion results in step S1 to train and evaluate the cross pseudo-supervised semi-supervised deep learning model, and obtain the image segmentation result according to the evaluated cross pseudo-supervised semi-supervised deep learning model; S4. According to the image segmentation result, use the eight-neighborhood connection method based on the jet center axis point to extract the jet axis.
2. The automatic extraction method of the jet axis based on cross pseudo-supervision semi-supervision according to claim 1, characterized in that The S1 includes the following steps: S101. Obtain 11 types of grid wind field data; S102. Calculate the wind speed and wind direction of each grid point. Among them, according to the quadrant of the wind speed, adjust to determine the accuracy of the direction angle; S103. Map the calculated wind speed and wind direction to RGB three-channel data to obtain wind field image coding data, and complete the preprocessing of the wind field processing data. Among them, the R-channel data represents the magnitude of the wind speed mapped to the numerical range of the R-channel of the color picture through coding, and the G-channel data and B-channel data respectively represent the wind direction mapped to the numerical ranges of the G-channel and B-channel of the color picture. The wind field image coding data includes annotated data and unannotated data; S104. Based on the manually annotated data, read the label mask; S105. Extract all connected target area lines in the label mask to obtain the contour of the label mask; S106. Randomly screen and convert the contour of the label mask to obtain the retained contour; S107. According to the retained contour, redraw and generate a new label mask, and save the new label mask to complete the label mask conversion process.
3. The method for automatically extracting the jet axis based on cross pseudo-supervision semi-supervision according to claim 2, characterized in that, The expression of the wind speed is as follows: Among them, Speed represents the wind speed, and both U and V represent wind speed components; The expression of the wind direction is as follows: Among them, Direction represents the wind direction.
4. The automatic extraction method of jet axis based on cross pseudo-supervision semi-supervised according to claim 3, characterized in that The mapping of the calculated wind speed and wind direction to RGB three-channel data is specifically: use the following formula to map the calculated wind speed and wind direction to RGB three-channel data: Among them, R (i,j) represents the wind speed magnitude at the (i, j) position of the MICAPS 11-class grid wind field data, which is mapped to the R-channel value at the (i, j) position of the color picture pixel through coding. G (i,j) and B (i,j) respectively represent the wind direction at the (i, j) position of the MICAPS 11-class grid wind field data, which are mapped to the G-channel value and B-channel value at the (i, j) position of the color picture pixel through coding. Speed (i,j) represents the wind speed magnitude at the (i, j) position of the MICAPS 11-class grid wind field data, and Direction (i,j) represents the wind direction at the (i, j) position of the MICAPS 11-class grid wind field data.
5. The automatic extraction method of the jet axis based on cross pseudo-supervision semi-supervision according to claim 1, wherein The cross pseudo-supervised semi-supervised deep learning model includes: The first Swin-Unet network is used to convert the preprocessing results and label mask conversion results in step S1 into a predicted first jet area mask; The second Swin-Unet network is used to convert the preprocessing results and label mask conversion results in step S1 into a predicted second jet area mask, where the first Swin-Unet network and the second Swin-Unet network adopt different initialization processes; The cross pseudo-supervised module is used to obtain the image segmentation result according to the first jet area mask and the second jet area mask by generating the first pseudo-label and the second pseudo-label that mutually serve as supervision signals. Among them, each batch contains annotated data and unannotated data, and in each training batch, the cross pseudo-supervised semi-supervised deep learning model updates the weights using the annotated data, and at the same time guides the learning of the unannotated data through the pseudo-label.
6. The automatic extraction method of jet axis based on cross pseudo-supervision semi-supervision according to claim 5, characterized in that The structures of the first Swin-Unet network and the second Swin-Unet network are the same, and both include: An encoder, which uses multiple Swin Transformer modules to divide the preprocessing result of step S1 and the label mask conversion result into non-overlapping windows to extract deep features, where the self-attention mechanism is applied within each window; A bridging module for connecting the encoder and the decoder; A decoder, which restores the extracted deep features to the same resolution as the input image through layer-by-layer upsampling, where the input image is the preprocessing result of step S1 and the label mask conversion result; An output layer, which uses a convolutional layer to convert the image output by the decoder into a predicted jet region mask, where the jet region mask includes a first jet region mask and a second jet region mask.
7. The automatic extraction method of jet axis based on cross pseudo-supervised semi-supervised according to claim 5, characterized in that The expression of the supervised learning loss function of the cross pseudo-supervised semi-supervised deep learning model is as follows: L Sup = λL CE + μL Dice Among them, L Sup represents the supervised learning loss function of the cross pseudo-supervised semi-supervised deep learning model. Both λ and μ represent weight coefficients. L CE represents the cross-entropy loss function, and L Dice represents the dice loss function. C represents the total number of categories, and y i represents the true label, represents the probability of predicting to belong to category i.
8. The automatic extraction method of jet axis based on cross pseudo-supervision semi-supervision according to claim 1, characterized in that S4 includes the following steps: S401. According to the image segmentation result, find the starting pixel point, where there is only one neighborhood with a color value within the eight-neighborhood around the starting pixel point; S402. Traverse the eight-neighborhood of the starting pixel point, preferentially select the adjacent point that is closest to the current point and has not been traversed in the direction, and mark it as visited; S403. Taking the starting pixel point as the center, determine whether all adjacent points have been traversed. If so, add each adjacent point to the connected jet axis center point set in turn and enter step S404. Otherwise, return to step S402; S404. Based on the jet axis center point set, draw the jet line to complete the extraction of the jet axis.
Citation Information
Cited By
River channel torrent area prediction method based on big data acquisition and processing
CN121413821A
Ship rear propeller cavitation recognition and segmentation method based on deep neural network
CN122156647A