A method for classifying sea conditions of an offshore wind farm

Keyframes are extracted through HSV color histogram and Canny edge features, combined with Vision Transformer network and 2D Ranchos interpolation method, the problem of high-resolution sea condition images cannot be directly input, and high-precision and fast sea condition classification are achieved, which is suitable for all-weather sea condition monitoring.

CN119888375BActive Publication Date: 2025-07-29OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510352833.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-29
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

In the existing sea condition classification methods, high-resolution sea condition images cannot be directly input to the existing model, resulting in low credibility of classification results, small and non-standardized data sets, insufficient samples of extreme sea condition, affecting the classification effect.

Method used

Keyframes were extracted using HSV color histogram and Canny edge features to build a sea condition classification model based on Vision Transformer network, and high-resolution sea condition image input was adapted through the 2D Ranchos interpolation method, and wave morphological features were captured using the multi-head self-attention mechanism to construct an equalized data set and fine-tuning training.

Benefits of technology

Real-time sea condition classification is realized all-weather real-time sea condition classification, significantly improving classification accuracy and speed, improving the robustness and applicability of the model, and being able to process high-resolution sea condition images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888375B_ABST
    Figure CN119888375B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for classifying sea conditions of an offshore wind farm, comprising the following steps: S1: extracting key frames from the original marine monocular vision in-situ video; S2: annotating the sea condition levels based on the sea surface features in the key frame images to obtain a sea condition data set; S3: constructing a sea condition classification model; S4: preprocessing the images in the sea condition data set to obtain a model input sequence containing image patch embedding sequences, class vectors, and position encoding vectors; S5: inputting the model input sequence into the sea condition classification model for fine-tuning training until the model converges to obtain a trained sea condition classification model. The present invention constructs a sea condition classification model based on the ViT network, and uses a 2D interpolation method to enable the model to adapt to high-resolution sea condition image inputs. The multi-head self-attention mechanism of the model is used to capture the global features of wave forms, realizing all-weather real-time classification and significantly improving the accuracy, speed, and robustness of sea condition classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of ocean observation, and particularly relates to a method for classifying sea conditions in an offshore wind farm. Background Art

[0002] The sea condition table formulated by the World Meteorological Organization is an important standard for evaluating the current sea conditions in an ocean wind farm. It can describe the wave state in the wind farm, thereby providing a reference for meteorological forecasting and ocean engineering, and effectively optimizing the working window of offshore construction operations.

[0003] With the development of computer vision and image model technologies, the classification task of sea conditions has gradually shifted from analyzing the response data of wave measuring instruments to image classification tasks. The purpose is to invert the true sea condition level within the field of view through information such as the texture, breaking, and white caps on the ocean surface. Regarding the method of realizing sea condition classification through image recognition technology and image classification models, there are also some designs in the prior art. However, the existing recognition and classification methods generally face the following problems:

[0004] (1) The existing data sets are mostly small-scale and non-standardized data sets, with insufficient extreme sea condition samples, which easily lead to bias in model training and low credibility of classification results;

[0005] (2) The original sea condition images usually have a high resolution, while the existing feature extraction and image classification models can accept input images with a low resolution and cannot directly process the original sea condition images. If the size of the original sea condition images is directly reduced and input into the model for recognition and classification, it will seriously affect the wave form in the images, thereby affecting the effect of sea condition classification.

[0006] Therefore, there is an urgent need for a sea condition classification method that can process high-resolution original sea condition images and whose classification accuracy and speed can both meet the requirements of all-weather real-time sea condition classification. Summary of the Invention

[0007] The purpose of the present invention is to solve one of the above technical problems, and provide a method for classifying sea conditions in an offshore wind farm. By fusing the HSV color histogram and Canny edge features to extract key frames, constructing a balanced data set covering multiple levels of sea conditions, constructing a sea condition classification model based on the Vision Transformer network, adapting the model to high-resolution sea condition image input through the 2D Lanczos interpolation method, and using the multi-head self-attention mechanism of the model to capture the global features of the wave form, all-weather real-time classification is realized, and the accuracy, speed, and robustness of sea condition classification are significantly improved.

[0008] To achieve the above purpose, the technical solution adopted by the present invention is:

[0009] A method for classifying sea conditions in an offshore wind farm, comprising the following steps:

[0010] S1: Obtain the original marine monocular vision in-situ video data of the offshore wind farm, and extract key frames from the video;

[0011] S2: Obtain the significant wave height and wind speed of the current sea area when the key frame image is taken, and label the key frames based on the significant wave height and wind speed to obtain a sea condition dataset classified by sea condition level;

[0012] S3: Construct a sea condition classification model based on the Vision Transformer network; the sea condition classification model includes an encoder and a multi-layer perceptron classification head;

[0013] S4: Preprocess the images in the sea condition dataset to obtain a model input sequence that combines image patch embedding sequences, class vectors, and position encoding vectors;

[0014] S5: Input the model input sequence into the sea condition classification model for fine-tuning training, so that the model input sequence continuously performs forward propagation in the encoder to extract the features corresponding to the class vectors, and input the features into the multi-layer perceptron classification head to obtain the classification results of the class vectors. During the training process, adjust the model parameters until the model converges to obtain the trained sea condition classification model;

[0015] S6: Use the trained sea condition classification model for sea condition classification.

[0016] In some embodiments of the present invention, the method for preprocessing the pictures in step S4 includes the following steps:

[0017] S41: Adjust the images in the sea condition dataset to a predetermined size;

[0018] S42: Perform image patch embedding processing on the adjusted images, divide the images into multiple image patches of a fixed size, and generate an image patch embedding sequence;

[0019] S43: Set a learnable embedding vector as the class vector for classification;

[0020] S44: Perform two-dimensional position encoding on each image patch using sine and cosine functions to obtain the original two-dimensional position encoding vector; perform 2D Lanczos interpolation operation on the original two-dimensional position encoding vector, calculate the position encoding of the interpolation positions in each dimension, and merge the position encodings of all dimensions to obtain the complete position encoding vector.

[0021] S45: Connect the image patch embedding sequence, class vector, and complete position encoding vector to generate a model input sequence.

[0022] In some embodiments of the present invention, in step S44, the method of performing two-dimensional position encoding on each image block using sine and cosine functions is as follows:

[0023] For any position in each image block (x,y) , the calculation formula for the position encoding of the even dimensions is:

[0024] ;

[0025] ;

[0026] The calculation formula for the position encoding of the odd dimensions is:

[0027] ;

[0028] ;

[0029] where x and y are the abscissa and ordinate in the current image block respectively, and PE(x, 2i) is the position x encoding in the 2nd i dimension, i and j are indices dependent on the dimension, D is the embedded dimension.

[0030] In some embodiments of the present invention, the formula for calculating the position encoding of the interpolation position in each dimension is:

[0031] ;

[0032] where is the encoding result of the interpolation position at the new position, is the target pixel position of each dimension, is the value of the original two-dimensional position encoding vector at , m and n are the pixel distances traversing the Lanczos window, and are respectively x and y the scaling factors of the two dimensions, is the Lanczos window parameter for controlling the interpolation range, is the Lanczos function;

[0033] where and are calculated as follows:

[0034] ;

[0035] ;

[0036] Among them, W old and H old are respectively the quantities of the original two-dimensional position encoding in the width and height dimensions, W new and H new are respectively the quantities of the position encoding required for the model input image in the width and height dimensions;

[0037] The calculation formula of the Lanczos function is:

[0038] .

[0039] In some embodiments of the present invention, during the fine-tuning training process, Focal Loss the loss function is used to accelerate the convergence of the model. The expression of the Focal Loss function is:

[0040] ;

[0041] Among them, is the predicted probability of the k th class of samples. The sum of the predicted probabilities of all classes is 1. is the One-Hot encoding of the true label, that is, the corresponding to the true class, otherwise it is 0. is the weight of class k . is the focusing coefficient, which is used to amplify the effect.

[0042] In some embodiments of the present invention, it further includes the following steps:

[0043] Calculate the macro-average value of the classification results of the sea state classification model; in this classification result, each classification category corresponds to a sea state level; the macro-average value includes the arithmetic mean of the precision, recall, and F1-Score of each class. The calculation formulas of the macro-average value are respectively:

[0044] ;

[0045] ;

[0046] ;

[0047] Among them, is the arithmetic mean of the precision, is the arithmetic mean of the recall, is the arithmetic mean of the F1-Score;

[0048] Calculate the micro-average of the classification results of the sea state classification model; the micro-average includes precision, recall, and F1-Score calculated based on the sum of TP, FP, and FN for each category. The micro-formulas are as follows:

[0049] ;

[0050] ;

[0051] ;

[0052] where is calculated based on the sum of TP, FP, and FN for each category, is the recall rate calculated based on the sum of TP, FP, and FN for each category, is the F1-Score calculated based on the sum of TP, FP, and FN for each category. TP means correctly predicting a positive sample as positive; TN means correctly predicting a negative sample as negative, FP means incorrectly predicting a negative sample as positive, and FN means incorrectly predicting a positive sample as negative; l i is the number of samples that the model predicts as the i th category and predicts correctly, m i is the number of samples that the model predicts as the i th category, n i is the number of samples that actually belong to the i-th category;

[0053] Calculate the classification accuracy Accuracy for each category in the classification results of the model:

[0054] ;

[0055] Evaluate the classification results of the trained sea state classification model based on the macro-average, micro-average, and classification accuracy of each category.

[0056] In some embodiments of the present invention, the method for extracting key frames from a video specifically includes the following steps:

[0057] S11: Extract the HSV histogram feature vector and canny edge contour feature vector of each frame image of the original marine monocular vision in-situ video;

[0058] S12: Concatenate the HSV histogram feature and canny edge contour feature to obtain a fused feature vector;

[0059] S13: Set a predetermined unit video duration; at every interval of the predetermined unit video duration, use the adaptive mean clustering method to classify each frame image within the unit video duration into corresponding clustering clusters in sequence, and update the clustering centers of each clustering cluster in real time;

[0060] S14: Extract the image frame closest to its respective clustering center as the key frame of each clustering.

[0061] In some embodiments of the present invention, the method for extracting the HSV histogram feature vector of each frame image of the original marine monocular vision in-situ video specifically includes the following steps:

[0062] Convert the image from the RGB domain to the HSV domain and then draw its color histogram; where H represents hue, S represents saturation, and V represents value; in the HSV domain, represent the color component through hue, and compress the three RGB channels into one channel for color segmentation and brightness processing;

[0063] After obtaining the color histogram, perform normalization processing on the histograms of the three dimensions respectively, and connect them into an HSV histogram feature vector.

[0064] In some embodiments of the present invention, the method for extracting the canny edge contour feature vector of each frame image of the original marine monocular vision in-situ video specifically includes the following steps:

[0065] Convert the image to a grayscale image and use a convolution kernel to perform Gaussian filtering to eliminate Gaussian noise and salt-and-pepper noise in the image;

[0066] Use the Sobel operator to calculate the magnitudes S and directions on the x-axis and y-axis, and the calculation formulas are respectively:

[0067] ;

[0068] ;

[0069] where I is the image to be processed, S x is the horizontal Sobel operator, S y is the vertical Sobel operator;

[0070] Use non-maximum suppression to retain the maximum gradient within the region and suppress other extractions. After completing non-maximum suppression, obtain a binary image, where the gray values of non-edge points are all 0, and the gray values of points that may be edges are 255;

[0071] Perform double-threshold screening on the binary image, set a predetermined upper threshold and a lower threshold, identify the pixel points in the image that are greater than the upper threshold as definitely being boundaries, identify the pixel points in the image that are less than the lower threshold as definitely not being boundaries, and identify the pixel points between the upper threshold and the lower threshold as candidates. If they are connected to the boundary, they are retained; otherwise, they are discarded.

[0072] In some embodiments of the present invention, step S13 specifically includes the following steps:

[0073] Perform an initialization operation, randomly select a feature vector as the initial clustering center. For each unselected data point, calculate the minimum distance between it and the existing clustering centers, and select the next clustering center according to the probability distribution of this distance;

[0074] Assign data points, calculate the cosine similarity of each feature vector B with the clustering center A and assign it to the most similar clustering center; the calculation formula of the cosine similarity is:

[0075] ;

[0076] where, A is the clustering center, B is the feature vector;

[0077] Update the clustering center. For each cluster, calculate the average value of the data points within the cluster to obtain a new clustering center;

[0078] Repeat the operations of assigning data points and updating the clustering center until a predetermined maximum number of iterations is reached.

[0079] The beneficial effects of the present invention are as follows:

[0080] 1. The present invention uses the sine and cosine functions and the 2D Lanczos interpolation method to perform position encoding on the extracted key-frame images respectively, realizing that without losing the wave information in the images, the key-frame images are input into the sea state classification model for training at a higher resolution.

[0081] 2. The present invention applies the Vision Transformer network structure to the monocular vision sea state classification task, and realizes the accurate classification of multi-level sea states by constructing a sea state classification model based on the Vision Transformer network structure. Compared with the existing sea state classification models, the classification speed is significantly improved while ensuring the classification accuracy;

[0082] 3. The present invention extracts key frames based on the HSV color histogram, Canny edge detection, and clustering methods, and establishes a multi-level sea state classification dataset based on monocular vision images for training a sea state classification model, enabling the images in the sea state dataset to cover sea surface conditions at different angles and under different illuminations, improving the applicability of the trained sea state classification model at different times, and realizing all-weather sea state classification.

[0083] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures pointed out in the specification, claims, and drawings. Brief Description of the Drawings

[0084] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will describe the specific embodiments of the present invention in detail with reference to the drawings. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0085] Figure 1 It is a flowchart of a sea state classification method for an offshore wind farm;

[0086] Figure 2 It is an algorithm flowchart of the key frame extraction method provided by the embodiment of the present invention;

[0087] Figure 3 It is a schematic diagram of the HSV transformation image of the key frame image provided by the embodiment of the present invention; <s

[0088] Among them, a is the original in-situ monocular vision image of the ocean, and b is the image after HSV domain transformation;

[0089] Figure 4 It is the color histogram corresponding to the HSV transformation image provided by the embodiment of the present invention;

[0090] Among them, a is the hue histogram, b is the saturation histogram, and c is the brightness histogram;

[0091] Figure 5 It is a schematic diagram of the network structure of the sea state classification model provided by the embodiment of the present invention;

[0092] Figure 6 It is a schematic diagram of the principle of the self-attention mechanism provided by the embodiment of the present invention;

[0093] Figure 7 It is a schematic diagram of the 2D interpolation method provided by the embodiment of the present invention;

[0094] Figure 8Schematic diagram of the hardware topology structure of the sea condition classification system for an offshore wind farm provided by an embodiment of the present invention. Detailed implementation manners

[0095] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be described and explained below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without creative efforts belong to the scope of protection of the present application.

[0096] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0097] In the case of no conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0098] The technical solutions of the present invention will be described in detail below in conjunction with specific embodiments and the accompanying drawings of the specification.

[0099] As shown in the appended Figure 1 - appended Figure 8 figures, in a schematic embodiment of a method for classifying sea conditions of an offshore wind farm in the present invention, the sea condition classification method includes the following steps.

[0100] S1: Collect the original marine monocular vision in-situ video data of the offshore wind farm through a monocular vision camera installed on an offshore platform, and extract key frames from the video.

[0101] In some embodiments of the present invention, as shown in the appended Figure 2 figures, the method for extracting key frames from the video specifically includes the following steps:

[0102] S11: Extract the HSV histogram feature vectors of each frame image of the original marine monocular vision in-situ video, and use the Canny edge detection algorithm to extract the canny edge contour feature vectors of each frame image.

[0103] S12: Concatenate the HSV histogram features and the canny edge contour features to obtain a fused feature vector.

[0104] S13: Set a predetermined unit video duration; at each interval of the predetermined unit video duration, use the K-Means++ clustering method to classify each frame image within the unit video duration into corresponding clustering clusters in sequence, and update the clustering centers of each clustering cluster in real time.

[0105] S14: Extract the image frame closest to its respective clustering center as the key frame of each clustering.

[0106] In some embodiments of the present invention, the method for extracting the HSV histogram feature vector of each frame image in the original marine monocular vision in-situ video specifically includes the following steps:

[0107] Convert the image from the RGB domain to the HSV domain and then draw its color histogram; where H represents hue, S represents saturation, and V represents value. In the HSV domain, the color component is represented by hue to make it more in line with human perception of color. This color component includes red, green, and blue components, and the three RGB channels are compressed into one channel for color segmentation and brightness processing, making it less susceptible to brightness changes.

[0108] It should be noted that the color histogram is a global feature of an image. By statistically counting the occurrence frequency of each color in the image, it provides a concise representation of the overall color characteristics of the image. In the case of night, the color histogram becomes a grayscale histogram, ensuring the normal extraction of key frames in night videos.

[0109] In this embodiment, the HSV transformation result of the original marine monocular vision in-situ image and its corresponding color histogram are as shown in the appendix Figure 3 - Appendix Figure 4 As shown, the color distribution in the image can be more clearly seen from the HSV image, and it has a stronger color contrast.

[0110] After obtaining the color histogram, perform normalization processing on the histograms of the three dimensions respectively, and connect them into an HSV histogram feature vector.

[0111] Since in actual operation, while ensuring the hue information, it is necessary to simplify the calculation complexity as much as possible. Therefore, in this embodiment, the number of bins in the hue channel is reduced from 180 to 90, and the number of bins in the other two channels is reduced from 256 to 64, thereby obtaining a feature vector with a smaller dimension and easier to calculate.

[0112] In some embodiments of the present invention, the method for extracting the canny edge contour feature vector of each frame image in the original marine monocular vision in-situ video specifically includes the following steps:

[0113] Convert the image to a grayscale image and use it for The convolution kernel is subjected to Gaussian filtering to eliminate Gaussian noise and salt-and-pepper noise in the image.

[0114] In image processing, since the gradient can well reflect the change of pixels, the greater the gradient change, the greater the difference between adjacent pixels, that is, there exists a regional edge. Therefore, in the Canny edge detection algorithm, the Sobel operator is used to calculate the magnitude S and direction of the gradients on the x-axis and y-axis The calculation formulas are respectively:

[0115] ;

[0116] ;

[0117] where I is the image to be processed, S x is the horizontal Sobel operator, S y is the vertical Sobel operator;

[0118] The Sobel edges obtained through the above calculations are still rough, and their edges are formed by the coincidence of multiple edges. Therefore, non-maximum suppression is further used to retain the maximum gradient within the region and suppress other extractions to achieve more accurate edge detection.

[0119] After non-maximum suppression, a binary image is obtained. Among them, the gray values of non-edge points are all 0, and the gray values of points that may be edges are 255.

[0120] Since the obtained binary image still contains many false edges caused by noise and other reasons, the binary image is further subjected to double-threshold screening. The upper bound and lower bound of the predetermined threshold are set. The pixel points in the image greater than the upper threshold are determined to be definitely boundaries, the pixel points in the image less than the lower threshold are determined to be definitely not boundaries, and the pixel points between the upper threshold and the lower threshold are determined to be candidates. If they are connected to the boundary, they are retained; otherwise, they are discarded.

[0121] In the above exemplary embodiments, the color histogram and the Canny edge operator combine the global color information and local texture information of the current frame, and can perform a more comprehensive evaluation on the image frames in the original marine monocular vision in-situ video.

[0122] In some embodiments of the present invention, the K-Means++ clustering method is used to extract key frames from the original marine monocular vision in-situ video. It should be noted that clustering is an unsupervised learning method in the field of machine learning. Without prior knowledge and labels, it can explore the internal laws and structures in the data, divide the data into multiple categories, and objects within each category have high similarity, while objects in different categories show obvious differences.

[0123] Specifically, step S13 includes the following steps:

[0124] Step 1: Perform an initialization operation. Randomly select a feature vector as the initial clustering center. For each unselected data point, calculate the minimum distance between it and the existing clustering centers, and select the next clustering center according to the probability distribution of this distance.

[0125] Step 2: Assign data points. Use cosine similarity to calculate the similarity between each feature vector B and the clustering center A, and assign it to the most similar clustering center; the calculation formula of cosine similarity is:

[0126] ;

[0127] where A is the clustering center and B is the feature vector.

[0128] Step 3: Update the clustering center. For each cluster, calculate the average value of the data points within the cluster to obtain a new clustering center;

[0129] Step 4: Iterate until the condition is met. Repeat the operations of assigning data points and updating the clustering center, that is, repeat the above Step 2 and Step 3 until the predetermined maximum number of iterations is reached.

[0130] In this embodiment, the predetermined unit video duration is set to five minutes, three clustering centers are set during clustering, the fusion feature vectors of all frames are calculated, and the K-Means++ clustering algorithm is used to calculate the key frames within this unit video duration, that is, in the original marine monocular vision in-situ video, three key frames are extracted every five minutes. Therefore, among the 1272 videos with a duration of one hour obtained, a total of 45,792 key frames are extracted as data samples. After data cleaning and deleting meaningless distorted, overexposed, underexposed, lens flare and other key frames, 42,790 are left for subsequent processing.

[0131] S2: Obtain the significant wave height and wind speed of the current sea area at the moment when the key frame image is taken through the wave radar and wind direction and speed meter installed on the offshore platform respectively, and then label the sea state level of the key frames based on the significant wave height and wind speed to obtain a sea state dataset classified by sea state level. In the dataset, the resolution of each key frame image is .

[0132] In this embodiment, based on the data of the wave radar and the wind direction and speed meter, the upper limit of the significant wave height is set to 293 cm, and the upper limit of the wind speed is set to 12.97 m / s, corresponding to the sea state of level 6 in "Marine Wind Force Grades and Performances". The significant wave height and wind speed in the current sea area are judged, and the sea state of each key frame is evaluated using the wave height and wind speed respectively, and a sea state classification data set containing images of sea states from level 1 to level 6 is made.

[0133] In this embodiment, the specific method for determining the significant wave height and wind speed of the current sea area is as follows: through the image recognition model, based on the sea surface features in the current key frame image, look up the sea surface wind force grade and performance table shown in Table 1 to obtain the wave height grade and wind speed grade. If the two grades are the same, then this grade is the sea state grade of the current key frame; if the two grades are different, then according to the visual features of the current key frame, make an artificial judgment in accordance with the sea state feature description in "Marine Wind Force Grades and Performances".

[0134] Table 1 Sea surface wind force grade and performance

[0135]

[0136] Before determining the significant wave height and wind speed of the current sea area based on the sea surface features in the key frame image, to solve the problem of unbalanced class samples, in some embodiments of the present invention, step S2 further includes the following steps: classify the extracted key frames, and use two methods, majority class sample undersampling and minority class sample data augmentation, to process key frames of different classes.

[0137] Among them, the purpose of majority class sample undersampling is to reduce the number of majority class samples. Specifically, in this embodiment, the key frames with wind speed between 0 - 7.5 m / s and significant wave height between 0 - 100 cm are defined as majority class samples, and they are equally divided into three groups. Clustering operations are performed within each group, and the number of clustering centers in each group is 1 / 3 of the sample number. In this way, the number of majority class samples can be reduced to one-fourth of the original, totaling 14,893 key frames. Using the clustering method can reduce the sample number while minimizing information loss as much as possible.

[0138] The significance of data augmentation lies in generating minority-class samples to increase their proportion in the total samples. At the same time, data augmentation can also enhance the diversity of training data, alleviate overfitting of the network, and improve the generalization ability of the network. In computer vision, data augmentation of images mainly focuses on geometric transformation, color transformation, noise addition, and blurring. However, for monocular ocean images, it is necessary to ensure the integrity of features such as white caps, highlight areas, wave shapes, wave crests, and sea surface colors in the images. In this embodiment, six augmentation methods are selected from the nine commonly used image data augmentation methods shown in Table 2, and each key frame of the minority-class samples is randomly augmented twice, expanding the samples to three times the original, for a total of 15,796 key frames. The six selected augmentation methods are horizontal flipping, adding salt-and-pepper noise, image sharpening, perspective transformation, histogram equalization, and blurring.

[0139] Table 2 Data Augmentation Methods and Effects

[0140]

[0141] In this embodiment, among all 30,689 key frames, the resolution of each key frame is , where 70% is used as the training set, 20% is used as the validation set, and 10% is used as the test set. Among them, there are 6,764 key frames of sea state level 1, accounting for 22.04%; 4,145 key frames of sea state level 2, accounting for 13.51%; 7,072 key frames of sea state level 3, accounting for 23.04%; 4,387 key frames of sea state level 4, accounting for 14.30%; 3,275 key frames of sea state level 5, accounting for 10.67%; 5,046 key frames of sea state level 6, accounting for 16.44%. The basic proportion of data samples at each sea state level is balanced.

[0142] S3: Construct a sea state classification model ViT-L / 32 based on the Vision Transformer network; the structure of this sea state classification model is as shown in the appendix Figure 5 , including an encoder and a multi-layer perceptron classification head.

[0143] It should be noted that the Vision Transformer (ViT) network is inspired by the Transformer deep learning model architecture, splits images into patches, and uses the linear embedding sequence of these image patches as input to train a classification model in a supervised manner based on the self-attention mechanism. Compared with traditional convolutional neural networks, the self-attention mechanism can effectively capture global context information in images, which is very important for sea state classification because the characteristics of the ocean environment are often global. Since the ViT network does not rely on convolutional operations but splits images into patches for processing, it can reduce the bias of the model and improve the sensitivity to different sea state characteristics.

[0144] Among them, the encoder consists of alternating multi-head self-attention layers and multi-layer perceptrons (MLPs). Layer normalization is applied before these two layers, and residual connections are applied after the two layers. Among them, layer normalization converts the activation values of each layer into a distribution with a mean of 0 and a standard deviation of 1, and then scales and offsets the results to keep the input of each layer stable and alleviate the problems of vanishing gradients or exploding gradients. Its calculation formula is:

[0145] .

[0146] Among them, and are the mean and variance respectively, and are the learnable scaling parameter and offset parameter.

[0147] The core of the self-attention mechanism is to convert the two-dimensional feature input into query, key, and value matrices, and then calculate the attention scores based on the matching of key-value pairs, and assign weights to the value matrix according to this weight matrix. Finally, the output after attention weighting is obtained. Its schematic diagram is as shown in Appendix Figure 6 shown, and the specific calculation formula is:

[0148] ;

[0149] .

[0150] Among them, Q h , K h , V h are the query matrix, key matrix, and value matrix respectively. The softmax activation function is used to normalize the weights and convert the weights into a probability distribution, making the self-attention scores more interpretable. d k is the dimension of the key matrix. Dividing by is to prevent from having too large a value, resulting in too small an output value of softmax, thus causing the model to fall into the problem of vanishing gradients.

[0151] Through this operation, the stability of the gradient and the training effect of the model can be enhanced. Secondly, compared with RNN and convolutional kernels, the self-attention mechanism can consider all input features, thus increasing the receptive field of the model for the two-dimensional features input at this moment.

[0152] The multi-head self-attention mechanism contains multiple self-attention layers. For the same input, after passing through different self-attention layers, h linear transformations are performed to generate h sets of queries, keys, and values. The output matrices are concatenated together and then passed through a linear layer to map the concatenated output to the final output space. Its calculation formula is:

[0153] .

[0154] Among them, , is a learnable weight matrix.

[0155] The multi-layer perceptron contains two fully connected layers and is an inverted bottleneck structure. The first fully connected layer increases the feature dimension from D to 4D, and the second fully connected layer restores the dimension. The Gaussian error linear unit (GeLU) is used as the activation function in the middle. Its approximate calculation formula is:

[0156] .

[0157] The residual structure directly adds the input to the output of a certain layer, mainly used to alleviate the vanishing gradient problem in deep networks. In the Transformer encoder, both the multi-head self-attention layer and the feed-forward network layer have residual connections. The output calculation formulas for the two layers are:

[0158] .

[0159] Among them, X is the input feature, Y 1 is the output after being processed by the multi-head self-attention layer, Y 2 is the result after being processed by the entire encoder layer.

[0160] S4: Preprocess the images in the sea state dataset to obtain a model input sequence that combines image patch embedding sequences, class vectors, and position encoding vectors.

[0161] It should be noted that since the resolution of the input images in the conventional Vision Transformer network model during pre-training is , and the resolution of the key frame images in the dataset is , if the sea state monocular images are directly reduced to size, it will seriously affect the wave form in the images, thus affecting the sea state classification effect. And if the sea state monocular images are directly input into the network, it will seriously affect the sea state classification effect of the model. Therefore, it is necessary to preprocess the images in the dataset before sea state classification.

[0162] In some embodiments of the present invention, the method for preprocessing the pictures in the dataset in step S4 includes the following steps:

[0163] S41: Adjust the images in the sea condition dataset to a predetermined size. In this embodiment, the images are scaled down proportionally to a size of 1280×720.

[0164] S42: Perform image patch embedding processing on the adjusted images, divide the images into multiple image patches of a fixed size, and generate an image patch embedding sequence. In this embodiment, block segmentation with a size of 32×32 is used, and the calculation formula for the new number of blocks is:

[0165] 。

[0166] Original number of position encodings, Number of position encodings required for the input image 。

[0167] S43: Set a learnable embedding vector as the category vector for classification.

[0168] S44: Perform two-dimensional position encoding on each image patch using sine and cosine functions to obtain the original two-dimensional position encoding vector; perform a 2D Lanczos interpolation operation on the original two-dimensional position encoding vector, calculate the position encoding of the interpolation positions in each dimension, and combine the position encodings of all dimensions to obtain the complete position encoding vector.

[0169] S45: Concatenate the image patch embedding sequence, the category vector, and the complete position encoding vector to generate the model input sequence.

[0170] In some embodiments of the present invention, in step S44, the method for performing two-dimensional position encoding on each image patch using sine and cosine functions is:

[0171] For any position in each image patch (x,y) , the calculation formula for the position encoding of the even dimension is:

[0172] ;

[0173] ;

[0174] The calculation formula for the position encoding of the odd dimension is:

[0175] ;

[0176] ;

[0177] Where x and y are the abscissa and ordinate in the current image patch respectively, PE(x,2i)is the position x in the 2nd i dimensional encoding, i and j is the dimension-dependent index, D is the embedded dimension.

[0178] In some embodiments of the present invention, in step S44, as shown in the appendix Figure 7 shown, the method for performing the 2D Lanczos interpolation operation specifically includes the following steps:

[0179] Obtain the original two-dimensional position encoding using the sin-cos encoding method.

[0180] Calculate the quantities of the original two-dimensional position encoding and the new position encoding to determine the spatial relationship.

[0181] Use the 2D Lanczos interpolation method to calculate the position encoding of the interpolation position in each dimension.

[0182] Merge all dimension position encodings and the class vector.

[0183] Among them, the formula for calculating the position encoding of the interpolation position in each dimension is:

[0184] ;

[0185] Among them, is the encoding result of the interpolation position at the new position, is the target pixel position of each dimension, is the value of the original two-dimensional position encoding vector at , m and n are the pixel distances for traversing the Lanczos window, and are respectively x and y the scaling factors of the two dimensions, that is, the ratio of the width of the image adjusted to the predetermined size to the width of the divided image block and the ratio of the height of the image adjusted to the predetermined size to the height of the divided image block, is the Lanczos window parameter for controlling the interpolation range, is the Lanczos function;

[0186] Among them, and The calculation formulas of are respectively

[0187] ;

[0188] ;

[0189] Among them, Wold and H old respectively represent the quantities of the original position encoding in the width and height dimensions, while W new and H new respectively represent the quantities of the position encoding required for the input image in the width and height dimensions;

[0190] The calculation formula of the Lanczos function is:

[0191] .

[0192] It should be noted that the input image size of the existing image classification model based on the Vision Transformer network structure is usually 224×224, and the size of the divided image patches is 32×32. Therefore, a total of 49 image patches can be divided, that is, a total of 49 original position encodings. The sea condition image in this embodiment is 2560×1440. If it is directly reduced to 224×224, the image ratio will be seriously damaged, affecting the accuracy of sea condition classification. Using the 2D interpolation method to expand the number of position encodings in this application can avoid the distortion of the sea wave form, better retain the wave form, and improve the accuracy of sea condition classification.

[0193] In the sea condition classification model based on ViT, the accuracy of the position encoding directly affects the model's understanding of the spatial relationship between image patches. Compared with other interpolation methods, the 2D Lanczos method has higher interpolation accuracy, ensuring that the position information of each image patch remains at a high precision, thereby improving the performance of the model.

[0194] In addition, the original position encoding is sin-cos encoding, which is a non-linear encoding. Discontinuities may be introduced during the discretization or interpolation process, resulting in the sawtooth effect. In this embodiment, the 2D Lanczos interpolation method is used in cooperation with it. On the one hand, it has less distortion when processing non-linear features, ensuring the accuracy of the position encoding. On the other hand, it can effectively reduce the sawtooth effect, ensuring the continuity and smoothness of the position encoding.

[0195] In order to further suppress the problem of sample imbalance in different category datasets, in some embodiments of the present invention, the Focal Loss loss function is used during the fine-tuning training process to accelerate the model convergence, reducing the weight for easily classified samples and increasing the weight for difficult-to-classify samples Focal Loss The expression of the loss function is:

[0196] ;

[0197] Among them, is for the kThe predicted probability of the class samples, and the sum of the predicted probabilities of all classes is 1. is the One-Hot Encoding code for the true label, that is, the corresponding to the true class, otherwise it is 0. is the weight of the class k . is the focusing coefficient, which is used to amplify the effect.

[0198] S5: Input the model input sequence into the sea state classification model for fine-tuning training of the model.

[0199] During the training process, the model input sequence is continuously passed forward in the encoder. All the encoder blocks stacked serially are used to extract the features corresponding to the class vectors for image classification. After obtaining the features, the features are further input into the multi-layer perceptron classification head to obtain the classification results of the class vectors. During the training process, the model parameters are adjusted until the model converges, and the trained sea state classification model is obtained and used for sea state classification.

[0200] S6: Use the trained sea state classification model for sea state classification.

[0201] In some embodiments of the present invention, the following steps are further included:

[0202] S7: Calculate the macro-average of the classification results of the sea state classification model; in this classification result, each classification category corresponds to a sea state level; the macro-average includes the arithmetic mean of the precision, recall, and F1-Score of each category. The macro-average calculation formulas are respectively:

[0203] ;

[0204] ;

[0205] .

[0206] Among them, is the arithmetic mean of the precision, is the arithmetic mean of the recall, is the arithmetic mean of the F1-Score.

[0207] Calculate the micro-average of the classification results of the sea state classification model; the micro-average includes the precision, recall, and F1-Score calculated based on the sum of TP, FP, and FN of each category. The micro-calculation formulas are respectively:

[0208] ;

[0209] ;

[0210] 。

[0211] Among them, is the accuracy calculated based on the sum of TP, FP, and FN for each category, is the recall calculated based on the sum of TP, FP, and FN for each category, is the F1-Score calculated based on the sum of TP, FP, and FN for each category; TP (True Positive) is the number of samples correctly predicted as positive by the model, FP (False Positive) is the number of samples incorrectly predicted as positive by the model, and FN (False Negative) is the number of samples incorrectly predicted as negative by the model; l i is the number of samples predicted correctly by the model as the i th category, m i is the number of categories predicted by the model as the i th category, n i is the number of samples actually belonging to the i-th category;

[0212] Calculate the classification accuracy Accuracy of each category in the classification results of the model:

[0213] ;

[0214] Among them, TN is the number of samples correctly predicted as negative by the model;

[0215] Evaluate the classification results of the trained sea state classification model based on the macro-average, micro-average, and the classification accuracy of each category.

[0216] Specifically, when the macro-average is much larger or much smaller than the micro-average, it indicates that there is a serious imbalance in the classification of the model. It is necessary to introduce a weighted F1-Score to measure the balance problem of the samples. The weighted F1-Score is the result of multiplying the F1-Score by the proportion of the class and then summing them, and different weights are assigned according to the proportion of each class.

[0217] In this embodiment, in the training and validation tasks of the sea state classification model based on the Vision Transformer network, the number of training set samples is 21,482, the validation set is 6,138, and the test set is 3,069. The dataset covers monocular images under different shooting angles, different time periods, and different lighting conditions. The summary of the hyperparameters of the sea state classification model is shown in Table 3.

[0218] Table 3 Hyperparameter Table for Sea State Classification Task

[0219]

[0220] Some embodiments of the present invention further provide an in-situ wave element extraction system, the hardware topology of which is as shown in the appendix Figure 8 and the system includes at least a monocular vision camera and a data processing terminal.

[0221] Among them, the monocular vision camera is installed on the platform of the offshore wind turbine and is used to capture the monocular vision video of the in-situ impact waves of the offshore structure. It is connected to a burner and an industrial control computer and transmits data through Ethernet.

[0222] The data processing terminal includes a trained sea state classification model, which communicates with the monocular vision camera to obtain the monocular video data captured by it, extracts key frames from the video data, and classifies the key frame images.

[0223] In some embodiments of the present invention, the system further includes a wave radar.

[0224] The wave radar and the monocular vision camera are installed on the same offshore wind turbine platform and are used to obtain the significant wave height and spectral peak frequency of the sea area within the range of the monocular vision camera. It is connected to the data processing terminal through RS485 and transmits the data to the industrial control computer through Ethernet.

[0225] The data processing terminal communicates with the wave radar to obtain the wave radar data.

[0226] In some embodiments of the present invention, the system further includes an industrial control computer.

[0227] The industrial control computer is connected to both the monocular vision camera and the wave radar, receives the data sent by the monocular vision camera and the wave radar, controls the monocular vision camera and the wave radar, and provides unified timing for all its connected devices to ensure the temporal consistency of the acquired data. The industrial control computer is also connected to the shore-based data center to transmit the acquired data back to the shore-based data center through the network.

[0228] In a specific embodiment of the present invention, the monocular vision camera is installed at the top of the outer platform of the column of the offshore wind turbine platform, and the height from the average water level line is 14.02m. The wave radar is installed at the bottom of the outer platform, and the height from the average water level line is 12.89m.

[0229] The monocular vision camera is from Hikvision, with the model of dome camera DS-2DE2204IW-DE3 / W / XM. It can output 30fps video with a maximum resolution of 1920×1080. The size of its CMOS is 1 / 1.28 inches. The variable range of the focal length f is 2.8 - 12mm. The horizontal field of view (FOV) ranges between 25° and 100°, the vertical FOV ranges between 14.1° and 56.3°, and the diagonal FOV ranges between 28.7° and 114.7°. Its variable focal length and large FOV can better meet the requirements of sea surface observation tasks in large scenes.

[0230] The wave radar is WaveGuide 5 Direction WG5-DR-CP produced by Radac Company in the Netherlands, with a sampling frequency of 10Hz. The measurement accuracy of the current sea surface elevation can reach ±3mm. At the same time, it can provide a wave height data per minute, and the accuracy of its wave height measurement value can reach ±1cm. The measurement accuracy of the wave propagation direction is ±2°.

[0231] Finally, it should be noted that: the embodiments in this specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts between the embodiments, reference can be made to each other.

[0232] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them; although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or perform equivalent replacements for some technical features; without departing from the spirit of the technical solutions of the present invention, they should all be covered within the scope of the technical solutions claimed by the present invention.

Claims

1. A method for classifying sea conditions in an offshore wind farm, characterized in that, It includes the following steps: S1: Obtain the original in-situ video data of marine monocular vision, and extract key frames from the video; S2: Obtain the significant wave height and wind speed of the current sea area when the key frame image is taken, and label the key frames based on the significant wave height and wind speed to obtain a sea condition dataset classified by sea condition level; S3: Construct a sea condition classification model based on the Vision Transformer network; the sea condition classification model includes an encoder and a multi-layer perceptron classification head; S4: Preprocess the images in the sea condition dataset to obtain a model input sequence that combines image patch embedding sequences, class vectors, and position encoding vectors; The method for preprocessing the pictures in the dataset in step S4 includes the following steps: S41: Adjust the images in the sea condition dataset to a predetermined size; S42: Perform image patch embedding processing on the adjusted images, divide the images into multiple fixed-size image patches, and generate an image patch embedding sequence; S43: Set a learnable embedding vector as the class vector for classification; S44: Perform two-dimensional position encoding on each image patch using sine and cosine functions to obtain the original two-dimensional position encoding vector; perform a 2D Lanczos interpolation operation on the original two-dimensional position encoding vector, calculate the position encoding of the interpolation positions in each dimension, and combine the position encodings of all dimensions to obtain a complete position encoding vector; S45: Connect the image patch embedding sequence, the class vector, and the complete position encoding vector to generate a model input sequence; S5: Input the model input sequence into the sea condition classification model for fine-tuning training, so that the model input sequence continuously performs forward propagation in the encoder to extract the features corresponding to the class vector, and input the features into the multi-layer perceptron classification head to obtain the classification result of the class vector. During the training process, adjust the model parameters until the model converges to obtain the trained sea condition classification model; S6: Use the trained sea condition classification model for sea condition classification.

2. The method for classifying sea conditions of an offshore wind farm according to claim 1, wherein [[ID=]](x,y) For any position in each image block The calculation formula for the position encoding of odd dimensions is; , the calculation formula for the position encoding of even dimensions is as follows: ; ; PE(x,2i) ; ; wherein, x and y are the abscissa and ordinate in the current image block respectively, The formula for calculating the position encoding of the interpolation positions in each dimension is: is the position x encoded in the 2nd i dimension, i and j are dimension-dependent indices, and D is the dimension of the embedding.

3. The method for classifying sea conditions of an offshore wind farm according to claim 2, wherein, The calculation formula for the Lanczos function is: ; Among them, is the encoded result of the interpolation position at the new position, is the target pixel position for each dimension, is the value of the original two-dimensional position encoding vector at ; m and n are the pixel distances for traversing the Lanczos window, and are respectively x and y the scaling factors for two dimensions, is the Lanczos window parameter for controlling the interpolation range, is the Lanczos function; Among them, and The calculation formulas are respectively as follows: ; ; Among them, W old and H old are the numbers of the original two-dimensional position encodings in the width and height dimensions respectively, W new and H new are the numbers of the position encodings required for the model input image in the width and height dimensions respectively; Focal Loss 。 4. The method for classifying sea conditions of an offshore wind farm according to claim 1, characterized in that, During the fine-tuning training process, use Focal Loss the loss function to accelerate the model convergence. The It further includes the following steps: expression of the loss function is: ; Among them, is the predicted probability of the k type of samples. The sum of the predicted probabilities of all classes is 1. is the One-Hot encoding of the true label, that is, the corresponding to the true class, otherwise it is 0. is the weight of the class k . is the focus adjustment coefficient, which is used to amplify the effect.

5. The method for classifying sea conditions of an offshore wind farm according to claim 1, wherein, Calculate the macro-average value of the classification results of the sea condition classification model; in the classification results, each classification category corresponds to a sea condition level; the macro-average value includes the arithmetic mean of the precision, recall, and F1-Score of each category. The calculation formulas for the macro-average value are respectively: Calculate the micro-average value of the classification results of the sea condition classification model; the micro-average value includes the precision, recall, and F1-Score calculated based on the sum of TP, FP, and FN of each category. The micro-calculation formulas are respectively: ; ; ; Among them, is the arithmetic mean of precision rate, is the arithmetic mean of recall rate, is the arithmetic mean of F1-Score; Calculate the classification accuracy Accuracy of each category in the classification results of the model: ; ; ; Among them, is the accuracy calculated based on the sum of TP, FP, and FN for each category. is the recall calculated based on the sum of TP, FP, and FN for each category. is the F1-Score calculated based on the sum of TP, FP, and FN for each category; TP is the number of samples correctly predicted as the positive class by the model, FP is the number of negative class samples mispredicted as the positive class by the model, and FN is the number of positive class samples mispredicted as the negative class by the model. l i is the number of samples predicted as the i th class by the model and predicted correctly. m i is the number of classes predicted as the i th class by the model. n i is the number of samples actually belonging to the i-th class. ​ ; Among them, TN is the number of samples that the model correctly predicts as negative classes; Evaluate the classification results of the trained sea state classification model based on the macro-average, the micro-average, and the classification accuracy of each category.

6. The method for classifying sea conditions of an offshore wind farm according to claim 1, wherein, The method for extracting key frames from the video specifically includes the following steps: S11: Extract the HSV histogram feature vector and the canny edge contour feature vector of each frame image of the original marine monocular vision in-situ video; S12: Concatenate the HSV histogram feature and the canny edge contour feature to obtain a fused feature vector; S13: Set a predetermined unit video duration; at every interval of the predetermined unit video duration, use the adaptive mean clustering method to classify each frame image within the unit video duration into the corresponding clustering cluster in turn, and update the clustering centers of each clustering cluster in real time; S14: Extract the image frame closest to its respective clustering center as the key frame of each clustering.

7. The method for classifying sea conditions of an offshore wind farm according to claim 6, wherein The method for extracting the HSV histogram feature vector of each frame image of the original marine monocular vision in-situ video specifically includes the following steps: Convert the image from the RGB domain to the HSV domain and then draw its color histogram; where H represents hue, S represents saturation, and V represents value; in the HSV domain, represent the color component by hue, and compress the three RGB channels into one channel for color segmentation and brightness processing; After obtaining the color histogram, perform normalization processing on the histograms of the three dimensions respectively, and connect them into an HSV histogram feature vector.

8. The method for classifying sea conditions of an offshore wind farm according to claim 6, wherein, The method for extracting the canny edge contour feature vector of each frame image of the original marine monocular vision in-situ video specifically includes the following steps: Convert the image to grayscale and apply the convolution kernel for Gaussian filtering to eliminate Gaussian noise and salt-and-pepper noise in the image; Calculate the magnitude S and direction of the gradients on the x-axis and y-axis using the Sobel operator , and the calculation formulas are respectively as follows: ; ; Among them, I is the image to be processed, S x is the horizontal Sobel operator, S y is the vertical Sobel operator; Use non-maximum suppression to retain the maximum gradient within the region and suppress other extractions. After completing non-maximum suppression, obtain a binary image, where the gray values of non-edge points are all 0, and the gray values of points that may be edges are 255; Perform double-threshold screening on the binary image, set a predetermined upper threshold and a lower threshold, identify the pixel points in the image that are greater than the upper threshold as definitely being boundaries, identify the pixel points in the image that are less than the lower threshold as definitely not being boundaries, and identify the pixel points between the upper threshold and the lower threshold as candidates. If they are connected to the boundary, retain them; otherwise, discard them.

9. The method for classifying sea conditions of an offshore wind farm according to claim 6, wherein The step S13 specifically includes the following steps: Perform an initialization operation, randomly select a feature vector as the initial clustering center, for each unselected data point, calculate the minimum distance between it and the existing clustering centers, and select the next clustering center according to the probability distribution of this distance; Allocate data points and calculate each feature vector using cosine similarity B with the cluster center A for similarity and assign it to the most similar cluster center; the calculation formula for the cosine similarity is as follows: ; wherein, A is the clustering center, B is the feature vector; Update the clustering center. For each clustering, calculate the average value of the data points within the clustering to obtain a new clustering center; Repeat the operations of assigning data points and updating the clustering center until the predetermined maximum number of iterations is reached.

Citation Information

Patent Citations

  • Video key frame extraction method based on multi-characteristic fusion shot clustering

    CN107220585A

  • Video key frame extraction method

    CN115713717A