A method for classifying nearshore wave breaking types based on video image deep learning
By using deep learning models to process video images in nearshore waters, the efficiency and accuracy problems of wave breaking type classification in traditional methods have been solved, achieving highly adaptable automatic identification and supporting real-time monitoring and early warning of coastal engineering.
Patent Information
- Application Number
- CN202510634465.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Traditional methods struggle to achieve high spatiotemporal resolution and accurate automatic classification of nearshore wave breaking types, especially in complex marine environments where the model's generalization ability is insufficient, and the amount of real-time video data is large with low processing efficiency.
Nearshore video images are acquired using fixed or temporary video acquisition equipment, preprocessed and labeled, and a deep learning model is built. Convolutional neural networks and long short-term memory networks are used for feature extraction and time-series modeling to automatically identify wave breakage types.
It achieves efficient and accurate automatic classification of wave breaking types, adapts to various sea conditions, and supports real-time monitoring and early warning of coastal engineering projects.
Smart Images

Figure CN120431405B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of ocean observation, and particularly relates to a nearshore wave breaking type classification method based on video image deep learning. BACKGROUND
[0002] Wave breaking refers to the process that wave collapses and generates foam, spray and strong turbulence under the joint action of factors such as seabed topography, wind power and wave steepness ratio when wave propagates to the nearshore shallow water area. As the core mechanism of nearshore wave energy dissipation, it plays an important role in driving the evolution of nearshore topography and geomorphology. Wave breaking type (referred to as breaking type) is one of the key characteristics of the breaking process. Different types of breaking not only reflect the pattern of wave energy release, but also directly affect the nearshore hydrodynamic environment and geomorphology. Therefore, realizing real-time and efficient breaking type identification has important significance for coastal monitoring, wave dynamics research and nearshore geomorphology evolution analysis.
[0003] Traditionally, the classification of breaking type relies on empirical formula and a large number of field observations, which is difficult to meet the high spatiotemporal resolution observation demand of coastal breaking process. A large number of observations and laboratory simulation studies show that the wave breaking type is mainly controlled by the nearshore slope (tanβ) and wave steepness (H / L). Its relationship can be represented by calculating the Iribarren number (ξ):
[0004]
[0005]
[0006] Where tanβ represents the beach slope, H is the wave height, and L is the wave length. The subscript o represents the offshore condition, and the subscript b represents the breaking point condition. The value of ξ can be used for breaking type classification and is closely related to the beach state. However, its calculation relies on the dense observation of environmental parameters such as slope and wave in the study area. Moreover, in areas with complex topography (such as sand bars and beach corners), ξ may not be an ideal predictor of breaking type, and often needs to be analyzed comprehensively in combination with other dynamic geomorphological parameters.
[0007] With the rapid development of remote sensing and computer vision technology, coastal dynamic geomorphology monitoring based on optical images has become an efficient and non-contact observation method, and has shown wide application potential in the field of wave feature extraction. Continuous multiple frames of time series video images can more completely capture the spatiotemporal evolution characteristics of the wave breaking process.
[0008] Traditional methods rely on manual observation, which has problems such as low efficiency and insufficient accuracy. How to use video image technology to realize accurate automatic classification of wave breaking types has become a technical problem to be solved. This problem involves many aspects: first, the nearshore sea environment is complex and variable, and the wave breaking shape is significantly different. How to build a representative video dataset is a basic difficulty. Second, wave breaking is a dynamic process, and single spatial feature extraction cannot accurately capture its time sequence variation rule. Third, the terrain conditions and sea state characteristics of different sea areas are different, and the generalization ability of the model is challenged. In addition, the amount of real-time video data is huge, and how to realize fast and accurate classification reasoning under limited computing resources is also a big difficulty. These sub-problems are interrelated and together constitute the technical problem of automatic identification of nearshore wave breaking types. How to develop an efficient, accurate and adaptable automatic identification method to meet the needs of real-time monitoring and early warning of coastal engineering is a major challenge in this field. SUMMARY
[0009] Therefore, the present application provides a nearshore wave breaking type classification method based on video image deep learning. By collecting time series video data, a deep learning model is trained and reasoned to automatically complete the wave breaking type identification. This method is efficient and reliable, and can adapt to various sea conditions, and has high practical value in actual application.
[0010] According to one aspect of the present application, a nearshore wave breaking type classification method based on video image deep learning is provided, the method comprising:
[0011] Based on a fixed or temporary video acquisition device, video image data of a nearshore area is collected, and the video image data covers a high position of the nearshore wave breaking zone;
[0012] The video image data is preprocessed, and overflow, roll wave and collapse wave are labeled according to the data respectively;
[0013] A training set is constructed based on the labeled data, and a neural network model is trained based on the training set;
[0014] Based on the trained neural network model, feature extraction and time sequence modeling are performed on the nearshore wave image sequence, and the wave breaking type corresponding to the nearshore wave image sequence is predicted.
[0015] In the above technical solution, the method can efficiently distinguish overflow, roll wave and collapse wave, which are three common wave breaking types, by a series of processing and modeling operations on the collected nearshore area video image. It is of great significance to many fields such as ocean research and coastal zone management, and can provide more accurate wave breaking information.
[0016] The fixed or temporary video acquisition equipment is used to obtain the nearshore video image, so as to capture the complete and representative wave breaking process and related feature information to the greatest extent. The fixed equipment can continuously monitor a specific nearshore area for a long time, obtain a large amount of stable data, and reflect the long-term law of wave breaking in the area. The temporary equipment has flexibility and can be quickly deployed to collect specific scene data at different nearshore locations according to research needs.
[0017] Based on the labeled high-quality training set, the neural network model is trained, and with the help of deep learning algorithm, the neural network can automatically learn the feature mode of different types of wave breaking from a large number of labeled data. During the training process, the model continuously adjusts the internal parameters to minimize the error between the predicted wave breaking type and the actual labeled type, so as to gradually have the ability to accurately distinguish the three types of wave breaking.
[0018] Unlike the traditional manual visual judgment method of wave breaking type, the deep learning model is used for automatic classification, which avoids subjective differences caused by human factors and can more objectively process and classify a large amount of wave breaking data. After the model is trained, it can quickly classify new nearshore wave image data, and if new data is continuously obtained to expand the training set, the model has the potential for further optimization and improvement, which can adapt to the classification needs of different regions and different time periods of nearshore wave breaking.
[0019] In some embodiments, the video image data is preprocessed, and overflow, roll-up and collapse waves are labeled according to the data, including:
[0020] The video image data is divided into multiple video slices in a manner of 1-30 seconds per segment, and the segmented time segments are cropped to focus on the wave breaking area to obtain the cropped video slices;
[0021] The video slices are labeled, and the wave breaking types are classified into overflow, roll-up and collapse waves according to visual features;
[0022] The labeled video slices are frame-extracted to generate image sequences suitable for model input.
[0023] In the above technical solution, the nearshore video image collected by the present application is divided into segments of 1-30 seconds in length, so as to finely process and analyze the wave breaking process. The 1-30 second time length setting not only covers the key process of wave breaking, but also effectively avoids the additional computational load and data storage pressure caused by too long time length, and can focus on the complete period of a specific wave breaking event, such as the whole process of wave formation and breaking, so as to better serve the subsequent wave breaking feature extraction and modeling.
[0024] Further, the segmented time segments are cropped to precisely focus on the wave breaking area, and remove background information and interference factors in the video unrelated to wave breaking, such as beaches, coastal facilities, etc. In this way, not only does it improve the labeling efficiency and accuracy, allowing the labeling personnel to more closely observe the key features of wave breaking, but the cropped video slices are more in line with the requirements of subsequent model input, reducing the model's attention to irrelevant information, and thus improving the model's learning effect on wave breaking features.
[0025] In addition, the labeled video slices are frame-extracted to generate image sequences that meet the model input requirements. Since deep learning models usually require fixed size and frame rate inputs, frame extraction can unify video slices of different lengths and frame rates into a format that meets the model requirements. At the same time, frame extraction can reduce data volume, simplify computational complexity, and improve model training and inference efficiency. Frame extraction based on key frames can further highlight the key moments of wave breaking, helping the model focus on the key features of wave breaking. In actual operation, appropriate frame extraction strategies can be selected flexibly according to model requirements and wave breaking characteristics.
[0026] In some embodiments, the video image data is preprocessed and labeled for wave overflow, wave roll, and wave collapse based on the data, including:
[0027] The video image data is segmented into multiple video slices in a manner of 1-30 seconds per segment, and the segmented time segments are multi-scale cropped to retain small-scale foam details and medium-to-large-scale wave shapes for the wave breaking area, generating multi-scale video slices;
[0028] The multi-scale video slices are feature-enhanced to generate feature-enhanced video slices;
[0029] The feature-enhanced video slices are labeled to generate labeled wave-breaking images;
[0030] The labeled video slices are frame-extracted to generate image sequences suitable for model input.
[0031] In the above technical solution, the segmented time segments are subjected to multi-scale cropping to retain small-scale foam details and medium-to-large-scale wave shapes to comprehensively capture the complete characteristics of wave breaking from different scales. The small-scale foam details can reflect the microscopic movement and energy dissipation of the water body during wave breaking, which is of great significance in distinguishing different breaking types. The medium-to-large-scale wave shapes provide overall propagation and breaking trend information of the waves, which helps to understand the macro process of wave breaking. The video slices generated by multi-scale cropping can more comprehensively reflect the characteristics of wave breaking, providing a richer data basis for subsequent feature enhancement and model training. To achieve multi-scale cropping, a suitable scale range and cropping strategy need to be determined. On the one hand, an image pyramid-based method can be used to generate image sequences of different resolutions, and then cropping is performed at each scale. On the other hand, a sliding window technique can be used to move the window at different scales to capture the key areas of wave breaking.
[0032] The purpose of feature enhancement is to improve the recognizability of wave breaking features, making different types of wave breaking more visually obvious, which facilitates subsequent labeling and model learning. Common feature enhancement methods include contrast enhancement, edge detection, and filtering. Contrast enhancement can highlight the differences between wave breaking areas and the surrounding background, making the outline of wave breaking clearer. Edge detection can emphasize the movement boundaries of the water body during wave breaking, helping to capture key features of wave breaking. Filtering techniques can remove noise interference and enhance the stability of wave breaking features.
[0033] In some embodiments, the multi-scale video slices are subjected to feature enhancement to generate feature-enhanced video slices, including:
[0034] The multi-scale video slices are subjected to feature enhancement, using a Canny edge detection operator to strengthen the wave crest outline, setting a brightness gain coefficient to highlight the contrast of the foam, extracting local binary pattern texture features to enhance dynamic textures, and generating feature-enhanced multi-scale video slices.
[0035] In the above technical solution, when the segmented time segments are subjected to multi-scale cropping, the Canny edge detection operator can accurately locate the wave crest edges in the wave breaking image. It calculates the image gradient amplitude and direction, judges the edges based on the gradient changes, and outlines the wave crest outline during wave breaking, making it stand out and facilitating subsequent identification. It can also suppress noise and reduce meaningless noise points generated by wave movement, resulting in clear and coherent wave crest outline curves, improving feature enhancement accuracy.
[0036] By setting the brightness gain coefficient, the brightness of the foam area can be increased to make it stand out against the dark background of seawater, sand, and other dark backgrounds, helping labeling personnel accurately identify the position and range of the foam and providing clear data basis for model learning of foam features.
[0037] The local binary pattern (LBP) operator can effectively capture the local texture features of the wave breaking area, such as the foam particle feeling, wave surface water marks, etc., which reflect the water movement state and energy dissipation, and are of great significance to distinguish different wave breaking types. Extracting and enhancing these texture features can provide more detailed information for the model and improve the classification performance. The LBP algorithm is simple and efficient, suitable for large data processing of multi-scale video slices, can quickly extract and enhance texture features, does not increase the computational burden of the preprocessing process, and ensures the feasibility and practicality of the method.
[0038] In some embodiments, the overflow, roll wave and collapse wave are labeled according to the data respectively, comprising:
[0039] According to the significant visual features in the slice, if it presents gentle flow and low wave crest, it is labeled as overflow, if it presents tubular curling and white foam, it is labeled as roll wave, and if it presents spattering breaking, it is labeled as collapse wave.
[0040] In the above technical solution, when labeling the video slice, if the wave presents gentle flow and low wave crest characteristics, it is labeled as overflow. In this case, the wave energy is small, and the breaking wave slowly advances along the coastline, the wave crest is low and gentle, and there is no obvious curling or collapse. For example, the wave breaking of the gentle slope coast in shallow water area, the labeling personnel need to observe the wave shape and flow characteristics carefully to determine whether it is overflow. If the wave front forms a curled tubular structure, gradually closes and produces a large amount of white foam during the advance, it is labeled as roll wave. The energy of roll wave breaking is large, which is often seen in sandy beach coast with moderate submarine slope, and the barrel-shaped wave is a typical roll wave. The labeling personnel should pay attention to the curling shape of the wave front and the distribution of foam production. If the wave breaks suddenly at the top and collapses towards the shore, producing strong spatter and water dispersion, it is labeled as collapse wave. Collapse wave usually occurs on steep coast, and the wave energy is rapidly released when it encounters steep slope. The labeling personnel need to observe the water movement situation at the moment of breaking and the shape after breaking to determine whether it is a collapse wave.
[0041] In some embodiments, when the segmented time segments are cropped, the time interval between the wave breaking slices at similar positions is more than 1 minute.
[0042] In the above technical solution, due to the complexity and variability of marine environment, the wave breaking at similar positions at different times may present different characteristics due to factors such as tide, wind and current. The interval selection can avoid sample bias caused by selecting only continuous slices in a short period, so that the model can learn the wave breaking characteristics under different conditions, thereby improving its adaptability in practical application.
[0043] If the time interval of wave breaking slices at similar positions is too short, the wave breaking characteristics of adjacent slices may be similar, resulting in limited feature patterns learned by the model. Setting a time interval of more than 1 minute can reduce the correlation between samples, so that each sample independently reflects the wave breaking characteristics at different times, which helps the model to learn the essential characteristics of different wave breaking types more accurately, thereby improving its generalization performance.
[0044] The increase of time interval makes the visual features of different wave breaking slices more distinct, and the labelers can more easily judge the wave breaking types according to the obvious visual features. For example, the characteristics of overflow, such as gentle flow and low wave crest, tubular curl and white foam of roll wave, and spatter-like breaking of breaking wave, will be more prominent and easy to distinguish. This reduces the feature ambiguity and annotation ambiguity caused by the continuity of wave breaking, reduces the difficulty and workload of annotation, and improves the accuracy and efficiency of annotation.
[0045] Because the features of different time slices are relatively independent and obvious, the labelers can make judgments based on unified annotation standards when annotating different slices at similar positions, reducing subjective judgment differences that may be caused by time continuity, and helping to improve the consistency of annotation results. At the same time, it is also easier to audit and verify subsequent annotation results, which can more accurately find and correct annotation errors and ensure the quality of annotation data.
[0046] In some embodiments, based on the trained neural network model, feature extraction and time series modeling are performed on the nearshore wave image sequence to predict the wave breaking type corresponding to the nearshore wave image sequence, comprising:
[0047] A pre-trained convolutional network is used as a feature extractor to encode the image sequence;
[0048] The encoded sequence is decoded by a long short-term memory network to generate a time series feature representation;
[0049] According to the time series feature representation, a classification label of the wave breaking type is determined.
[0050] In the above technical solution, when processing nearshore wave image sequence data, using a pre-trained convolutional network (CNN) as a feature extractor is an efficient and effective method. The rich feature hierarchy obtained by pre-training the convolutional network on large-scale image data can efficiently encode the spatial features in the wave breaking image. Through multiple convolution and pooling operations, CNN automatically learns different levels of features including edges, textures, shapes, etc., and can identify key features such as wave crest contours, foam distribution, and wave shape in the wave breaking image. After encoding, each frame of image is converted into a fixed-length feature vector. These feature vectors not only retain important information related to wave breaking types in the original image, but also reduce data dimensionality and reduce redundancy, thereby improving the efficiency of subsequent processing.
[0051] To further capture the temporal characteristics of the wave breaking process, a long short-term memory network (LSTM) is used to model the encoded feature sequence. LSTM is a neural network structure specifically designed for processing sequence data, which can effectively capture long-term dependencies in data. In the wave breaking type classification task, wave breaking is a process with time continuity, and the wave breaking features between consecutive frames are related. Using its internal gating mechanism (input gate, forget gate, and output gate), LSTM can selectively retain and update important temporal features related to wave breaking types while forgetting irrelevant or redundant information. In this way, LSTM generates a temporal feature representation that reflects the dynamic changes of the wave breaking process.
[0052] According to another aspect of the present application, a nearshore wave breaking type classification system based on video image deep learning is provided, based on the above method, comprising:
[0053] The acquisition module is used for acquiring video image data of the nearshore area based on a fixed or temporary video acquisition device, and the video image data covers a high position of the nearshore breaking wave zone in terms of perspective;
[0054] The preprocessing module is used for preprocessing the video image data and labeling overflow, roll wave and collapse wave according to the data, respectively;
[0055] The training module is used for constructing a training set based on the labeled data, and training a neural network model based on the training set;
[0056] The prediction module is used for feature extraction and temporal modeling of the nearshore wave image sequence based on the trained neural network model, and predicting the wave breaking type corresponding to the nearshore wave image sequence.
[0057] In the technical solution, in order to better use the method, the application provides a nearshore wave breaking type classification system based on video image deep learning, each module corresponds to each step of the method, and the specific principle has been described above, and will not be described here.
[0058] According to another aspect of the application, a nearshore wave breaking type classification device based on video image deep learning is provided, comprising:
[0059] at least one processor and a memory connected to the at least one processor in communication;
[0060] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.
[0061] In the technical solution, in order to better run and process the method, the method is stored in the memory, and the processor is used to execute the stored method. It should be noted that the principle and effect of each step have been described above, and will not be described here.
[0062] According to another aspect of the application, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the method described above.
[0063] In the technical solution, in order to better run and use the method, the method is stored in the computer readable storage medium, and the processor is used to implement the method. It should be noted that the principle and effect of each step have been described above, and will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0065] Figure 1 is a flowchart of an embodiment of a nearshore wave breaking type classification method based on video image deep learning of the application;
[0066] Figure 2 is a flowchart of another embodiment of a nearshore wave breaking type classification method based on video image deep learning of the application;
[0067] Figure 3is a video image collection device schematic diagram of another embodiment of a nearshore wave breaking type classification method based on video image deep learning of the present application;
[0068] Figure 4 is a data set production process schematic diagram of another embodiment of a nearshore wave breaking type classification method based on video image deep learning of the present application;
[0069] Figure 5 is a CRNN deep learning network architecture schematic diagram of another embodiment of a nearshore wave breaking type classification method based on video image deep learning of the present application;
[0070] Figure 6 is a CRNN model video image breaking wave classification task training result schematic diagram of another embodiment of a nearshore wave breaking type classification method based on video image deep learning of the present application;
[0071] Figure 7 is a structure schematic diagram of an embodiment of a nearshore wave breaking type classification system based on video image deep learning of the present application. DETAILED DESCRIPTION
[0072] The present application will be further described below in conjunction with the drawings and embodiments. It is particularly pointed out that the following embodiments are only for illustrating the present application, but do not limit the scope of the present application. Similarly, the following embodiments are only part of the embodiments of the present application, not all embodiments, and all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of the present application.
[0073] To this end, the present application proposes a nearshore wave breaking type classification method based on video image deep learning. By collecting time series video data, a deep learning model is constructed for training and reasoning, thereby automatically completing the breaking wave type identification. The method is efficient and reliable, and can adapt to various sea conditions, and has high practical value in actual application.
[0074] One of the embodiments
[0075] Please refer to Figure 1 A nearshore wave breaking type classification method based on video image deep learning, the method comprising:
[0076] S1, based on a fixed or temporary video collection device, collecting video image data of the nearshore area, the video image data visual angle covering the high position of the nearshore breaking wave belt;
[0077] In this embodiment, nearshore video image data is acquired, which is collected by fixed or mobile camera equipment deployed in the target coastal area, covering the nearshore breaking zone, for capturing dynamic information of the wave breaking process. Through the deployment of fixed or mobile video acquisition equipment in the target coastal area, video image data of the nearshore breaking zone is acquired, and high-resolution video sequences are obtained. According to the collected video image data, data preprocessing is performed using a standard of pixel resolution not less than 1920x1080 and frame rate not less than 10 frames / s, and standardized video frame sequences are obtained.
[0078] For example, through the deployment of fixed or mobile video acquisition equipment in the target coastal area, video image data of the nearshore breaking zone is acquired, using a camera with a pixel resolution of 1920x1080 and a frame rate of 30 frames / s, to ensure the high definition and continuity of the video sequence. According to the collected video image data, data preprocessing is performed using the OpenCV library, and video frame sequences that meet the input requirements of the deep learning model are obtained. If the video frame sequence contains redundant background information, the nearshore breaking zone is separated by the U-Net image segmentation algorithm, and the video segment focused on the wave breaking process is extracted.
[0079] S2, pre-process the video image data, and label the spill wave, roll wave and collapse wave according to the data respectively;
[0080] In this embodiment, S2, pre-process the video image data, and label the spill wave, roll wave and collapse wave according to the data respectively, including:
[0081] S21, divide the video image data into multiple video slices in a manner of 1-30 seconds per segment; crop the segmented time segments, focus on the wave breaking area, and obtain the cropped video slices;
[0082] Exemplarily, the fixed or temporary video acquisition device acquires nearshore wave video data of 1920x1080 resolution at a frame rate of no less than 10 frames per second to form an original video stream. A sliding window-based video segmentation technique is used to cut the video stream into multiple segments with a time window of 5 seconds, ensuring that each segment contains at least one complete wave breaking period. The inter-frame difference algorithm in the OpenCV library is used to calculate the pixel difference between adjacent frames. When the difference value of 10 consecutive frames exceeds the threshold value of 50, it is determined as the starting frame of wave breaking. When the difference value falls below the threshold value for 5 frames, it is determined as the ending frame. Based on the start and end frame markers, the FFmpeg tool is used to accurately crop the video segment, retaining the key 20-30 frames of the breaking process. If the time segment contains a wave breaking area, the YOLOv5 target detection model is used to identify the wave breaking area and determine the bounding box coordinates of the wave breaking area. Each time segment is regionally cropped based on the bounding box coordinates, and the OpenCV library is used to extract the image content focused on the wave breaking area. The resolution of the cropped image content is adjusted to 224x224 pixels, and the bilinear interpolation algorithm is used to obtain standardized wave breaking image slices.
[0083] S22, labeling the video slices, and classifying the wave breaking types into spilling, plunging and surging according to visual features;
[0084] In this embodiment, the spilling, plunging and surging are labeled according to the data, including: judging the wave breaking type according to the significant visual features in the slice. If it presents a gentle flow and a low wave crest, it is labeled as a spilling wave. If it presents a tubular curl and white foam, it is labeled as a plunging wave. If it presents a splashing breaking, it is labeled as a surging wave.
[0085] Specifically, the spilling wave (Spilling Breaker): the wave crest gradually steepens, forms white foam at the top, and the energy is slowly dissipated, commonly seen on gentle seabed slopes. The plunging wave (Plunging Breaker): the wave crest tilts to form a curled cavity, accompanied by intense energy release, often occurring on moderate slope seabed. The surging wave (Surging Breaker): the wave hardly forms obvious curling, directly rushes to the coast, and the energy is concentrated in a narrow area, often seen on steep beaches.
[0086] S23, frame extraction processing is performed on the labeled video slices to generate an image sequence suitable for model input.
[0087] As an optional embodiment, S2, the video image data is preprocessed, and the spilling, plunging and surging are labeled according to the data, including:
[0088] S21, divide the video image data into multiple video slices in a manner of 1-30 seconds per segment, perform multi-scale cropping on the segmented time segments, retain small-scale foam details and medium-large scale wave patterns for the wave breaking area, and generate multi-scale video slices;
[0089] In this embodiment, the original video stream data is segmented to generate 30-second video segments, a 224x224 pixel window is used, multi-scale cropping is combined, the cropping scale is 1 times, 0.5 times, and 0.25 times, the window sliding step length is 112 pixels, small-scale foam details and medium-large scale wave patterns are retained for the wave breaking area, and multi-scale wave breaking images are generated. The segmented 30-second video segments are cropped using a 224x224 pixel window, the cropping scale is set to 1 times, 0.5 times, and 0.25 times, and the window sliding step length is 112 pixels. The wave breaking area is cropped at multiple scales. If the cropped area contains foam details, the small-scale cropping result is retained; if the cropped area contains wave patterns, the medium-large scale cropping result is retained. Multi-scale wave breaking images are generated to ensure the complete retention of small-scale foam details and medium-large scale wave patterns.
[0090] For example, first, the original video stream is segmented in the time dimension, a video stream with a fixed frame rate of 30fps is used, and the segment command of FFmpeg is used to generate 30-second segments at an interval of 900 frames, ensuring that the time accuracy of each segment is controlled within ±0.03 seconds. In the spatial processing stage, after reading frame by frame using VideoCapture of OpenCV, Gaussian pyramid downsampling is performed to generate a multi-scale image group: three scale layers of original resolution 224x224 (1 times), 112x112 (0.5 times), and 56x56 (0.25 times). The sliding window operation uses a double loop structure, the outer loop traverses the original scale image with a step length of 112 pixels, and the inner loop simultaneously processes the corresponding regions of the three scale layers, maintaining geometric consistency through bicubic interpolation.
[0091] S22, feature enhancement is performed on the multi-scale video slices to generate feature-enhanced video slices;
[0092] In this embodiment, S22, feature enhancement is performed on the multi-scale video slices to generate feature-enhanced video slices, including:
[0093] S221, feature enhancement is performed on the multi-scale video slices, a Canny edge detection operator is used to enhance the wave peak profile, a brightness gain coefficient is set to highlight the foam contrast, a local binary pattern texture feature is extracted to enhance the dynamic texture, and feature-enhanced multi-scale video slices are generated.
[0094] S222, the feature-enhanced video slices are labeled to generate labeled wave breaking images;
[0095] S223, frame extraction processing is performed on the labeled video slice to generate an image sequence suitable for model input.
[0096] Exemplarily, the cropped image sequence is subjected to multi-scale processing to generate multi-scale breaking wave images, a Canny edge detection operator is used to extract wave peak contours, a brightness gain coefficient 1.2 is set to adjust the image brightness, the contrast between the foam region and the background is enhanced, a local binary pattern texture feature is extracted to capture dynamic texture information in the wave breaking process, the wave peak contour, the brightness-enhanced image and the dynamic texture feature are fused to generate a feature-enhanced breaking wave image.
[0097] In the image processing process, first, the multi-scale breaking wave image is subjected to Gaussian pyramid decomposition, a down-sampling operation with a scale factor of 1.2 is used to generate three images at different resolutions, the images at each layer are unified to the original size and then superimposed and fused by a bilinear interpolation algorithm to enhance the multi-scale feature representation of the image. Subsequently, a Canny edge detection operator is used, a Gaussian filter kernel size of 5x5 is set, and high and low threshold values of 120 and 60 respectively are set to extract wave peak contour information, and a morphological closing operation (3x3 rectangular structural element) is used to connect broken edges and strengthen the continuity of the wave structure. To highlight the contrast of the foam region, a linear brightness gain of 1.2 times is applied to the V channel of the HSV color space, while limiting the pixel value to not more than 255, and using histogram equalization (CLAHE algorithm, grid size 8x8, contrast limit 2.0) to enhance the local contrast. In the texture feature extraction stage, a circular neighborhood LBP operator with a radius of 2 pixels is used to divide the image into 16x16 local blocks to calculate the uniform mode feature, and after reducing the influence of light by gamma correction (γ=0.8), the texture feature map is fused with the original image by weighted fusion (weighting coefficient 0.7). Finally, non-local mean filtering (search window 15x15, similar window 5x5, decay parameter h=10) is used to eliminate the noise introduced in the enhancement process to generate a feature-enhanced breaking wave image, in which the gradient amplitude of the wave contour is increased by about 35%, the contrast standard deviation of the foam region is increased by 22%, and the discrimination of the texture feature is increased by 18%.
[0098] In this embodiment, when the segmented time segments are cropped, the time interval between the breaking wave slices at similar positions exceeds 1 minute.
[0099] S3, a training set is constructed based on the labeled data, and a neural network model is trained based on the training set;
[0100] Exemplarily, according to the training data set, a deep learning classification model is constructed, and a convolution-recurrent neural network structure is used to extract spatio-temporal features. If overfitting is detected during model training, the model parameters are adjusted through dropout and regularization techniques to obtain an optimized model structure. A preset optimizer and loss function are used to iteratively train the deep learning classification model, update the model parameters, and obtain a trained classification model. The performance of the trained classification model is evaluated by a multi-fold cross-validation method to determine the classification accuracy and generalization ability of the model. According to the evaluation results, the trained classification model is deployed to an edge computing device or a cloud server to perform real-time wave breaking type classification inference.
[0101] According to the training data set, a deep learning classification model is constructed, and a convolution-recurrent neural network structure is used to extract spatio-temporal features. If overfitting is detected during model training, the model parameters are adjusted through dropout and regularization techniques to obtain an optimized model structure. A preset optimizer and loss function are used to iteratively train the deep learning classification model, update the model parameters, and obtain a trained classification model. The performance of the trained classification model is evaluated by a multi-fold cross-validation method to determine the classification accuracy and generalization ability of the model. According to the evaluation results, the trained classification model is deployed to an edge computing device or a cloud server to perform real-time wave breaking type classification inference.
[0102] S4, based on the trained neural network model, feature extraction and time series modeling are performed on the nearshore wave image sequence to predict the wave breaking type corresponding to the nearshore wave image sequence.
[0103] In this embodiment, S4, based on the trained neural network model, feature extraction and time series modeling are performed on the nearshore wave image sequence to predict the wave breaking type corresponding to the nearshore wave image sequence, including:
[0104] S41, a pre-trained convolutional network is used as a feature extractor to encode the image sequence;
[0105] S42, the encoded sequence is decoded by a long short-term memory network to generate a time series feature representation;
[0106] S43, according to the time series feature representation, a classification label of the wave breaking type is determined.
[0107] For example, a convolutional-recurrent neural network is used to extract spatiotemporal features from the keyframe sequence. The convolutional layers use 3×3 kernels, and the recurrent layers use LSTM units to obtain dynamic feature vectors of the wave breaking process. A classification model is then used to infer the dynamic feature vectors, and the Softmax function is used to calculate the probability distribution, resulting in a probability distribution for the wave breaking type. If the highest probability value in the probability distribution is greater than 0.8, the corresponding wave breaking type label is output.
[0108] Based on the above embodiments, the present invention has the following advantages:
[0109] This method acquires nearshore wave-breaking zone video data using fixed or mobile cameras, segments, crops, and labels the videos to construct a time-series image dataset, and utilizes a deep learning model to extract spatial and temporal features for automatic classification of wave breaking types. This invention employs a deep learning model combining convolutional neural networks and recurrent neural networks to perform inference processing on real-time acquired nearshore videos, outputting wave breaking type labels, and applying the classification results to coastal engineering monitoring and early warning. Through transfer learning technology, this invention can adapt to diverse terrain and sea conditions, improving classification accuracy and generalization ability, providing efficient and reliable technical support for coastal condition monitoring and analysis.
[0110] Example 2
[0111] Based on the method described in one of the embodiments, this embodiment will be further explained and illustrated with a specific example. Please refer to [link / reference]. Figure 2 :
[0112] A1. Nearshore video image acquisition
[0113] Depending on the observation requirements, fixed or temporary video acquisition equipment will be deployed in the target coastal area. Figure 3 In the diagram, A represents a fixed network camera; B represents a mobile camera (such as a GoPro). The camera should be installed at a high position that can cover the nearshore wave-breaking zone, ensuring good visibility and stability. By continuously acquiring video image sequences, dynamic information about the wave-breaking process in the nearshore area is obtained, providing basic data support for subsequent image processing and classification. To ensure the accuracy of wave-breaking observation and the classification effect of the deep learning model, the pixel resolution of the video images should be no less than 1920×1080, and the frame rate should be no less than 10 frames per second.
[0114] A2. Creation of the Wave Classification Dataset
[0115] The original videos usually have a large field of view, and the breaking wave area occupies a small proportion in the picture. In addition, the length of a single video is usually about 10 minutes, and the frame rate is more than 10 frames per second. Directly using the video for deep learning model training may cause excessive computational burden or irrelevant information interference. In order to improve the usability of the data, we have four main processing steps for the original video: video segmentation, cropping, labeling and frame extraction Figure 2
[0116] The example data set of the application has been published on Figshare
[0117] (https: / / figshare.com / articles / dataset / VID_WBT_A_Video_Imagery_Dataset_for_Nearshore_Wave_Breaking_Type_Classification / 28814993), including two subsets: Set A and Set B.
[0118] -Set A: Video slices for nearshore wave breaking type classification, including 9000 breaking wave slices. The length of each slice is 30s, which shows a relatively complete breaking wave process. The slice is labeled as "spilling, plunging or surging" according to the most obvious breaking wave type, and the category information is directly labeled in the first word of the file name.
[0119] -Set B: Time sequence image data based on the breaking wave slices in Set A, extracted at a frequency of 2Hz. Each group of 60 frames is placed in a separate folder, and the folder name corresponds to the same breaking wave slice in Set A, which can be directly input into the deep learning model.
[0120] In this embodiment, please refer to Figure 4 , A2, breaking wave classification data set making, the specific steps are as follows:
[0121] A21, video segmentation
[0122] First, we segmented the original 10-minute video into 30-second segments. It should be particularly pointed out that in the nearshore area, the gravity wave caused by the wind has the highest energy proportion, and the period is 1-30 seconds. Therefore, by selecting 30 seconds as the segmentation unit, the complete wave breaking process can be covered while reducing data redundancy.
[0123] A22, cropping
[0124] On the basis of the segmented video clips, we identified the breaking process by visual interpretation and cropped it using a 224x224 pixel window to obtain the breaking slice. The cropping operation not only effectively focuses on the target area, but also maximizes the reduction of computing resource consumption. It should be particularly noted that the image size of 224x224 is the standard input size suitable for ResNet deep learning network. At the same time, this size has been proven in our practice to be sufficient to cover the spatiotemporal scale of the breaking process of nearshore waves. Users can also customize the cropping window size according to research needs (observation area scale and deep learning model used).
[0125] In addition, to ensure the relative independence of the data set, for multiple segmented clips generated from the same original video, we ensure that the time interval between breaking slices at similar positions is more than 1 minute when cropping. This can avoid the same single breaking process being repeatedly identified into the data set, thereby preventing data leakage and improving the generalization ability of the trained model.
[0126] A23, annotation
[0127] Selecting appropriate breaking slices, according to the most prominent visual features of the breaking process in the slice, it is labeled as "spilling, plunging and surging" one of the three types. In order to ensure the balance of the data set, the number of three types of breaking slices is strictly controlled to be 1:1:1 during the annotation process.
[0128] A24, frame extraction
[0129] Finally, in order to facilitate the deep learning model to classify the breaking type, we extracted the frame from the breaking slice. In the example of this invention, we use a frame extraction frequency of 2Hz. This frequency is common in the field of coastal video monitoring, which can provide sufficient temporal resolution to fully reflect the timing changes of different breaking types. At the same time, users can adjust the frame extraction frequency according to specific research needs to adapt to different application scenarios and analysis goals.
[0130] A3, deep learning breaking classification model design
[0131] In order to verify the feasibility of the deep learning model for the breaking type classification task of video images, we use the CNN+RNN (CRNN) model of time series deep learning to perform an exemplary classification task. CRNN combines the spatial feature extraction capability of CNN and the modeling capability of RNN for time series information, which is suitable for processing the dynamic evolution process of wave breaking. In addition, the CRNN model structure is relatively lightweight, suitable for different computing resource environments, and easy to reproduce and widely deploy. The model structure of CRNN is shown in Figure 5 .
[0132] -CNN as an encoder, by:
[0133] f CNN (x (t) )=z (t) #(3)
[0134] Each 2D image x (t) is encoded (i.e. compressed dimension) into a 1D vector z (t) .
[0135] -RNN as a decoder, receives the sequence input vector z (t) from the CNN encoder and outputs another 1D sequence h (t) . Finally, a fully connected neural network is connected to the output for wave type prediction. The CNN encoder used for validation is a ResNet-50 model pre-trained on the image dataset ILSVRC-2012-CLS, with the original fully connected layers removed and the convolutional layers retained as a spatial feature extractor; the RNN decoder uses a long short-term memory (LSTM) network. Our experiments show that this model design achieves a good balance between computational efficiency and feature expression capability.
[0136] A4、Model training
[0137] In the example model training process, the Adam optimizer (learning rate of 0.001) is used for gradient update, and the cross-entropy loss function is selected as the loss function of the model. To alleviate the problem of overfitting, the Dropout mechanism is introduced, and the dropout rate is set to 0.3. In addition, in order to comprehensively evaluate the generalization ability of the model, the 5-fold cross-validation strategy is adopted, ensuring that each sample is used for training and validation in five training cycles. This method effectively reduces the influence caused by data partition differences, thereby improving the robustness and reliability of model performance evaluation.
[0138] Figure 6 The training and validation results of 5-fold cross-validation are shown (A in the figure is the training loss result; B is the training accuracy result), highlighting the performance of the model in stability and generalization ability. As shown in Figure 6 A, the training loss remains at a low level, and although the validation loss peaks in the 4th fold, the overall trend is reasonable, indicating that the model can effectively learn features under different data partitions and does not exhibit serious overfitting. At the same time, from Figure 6It can be seen that the training accuracy is always kept above 95%, and the validation accuracy is kept above 93% in the first to third folds, although there is a short-term decline in the fourth fold, but it is restored in the fifth fold, which shows that the model has good adaptability and strong generalization ability. By adopting 5-fold cross-validation, the influence of data division randomness on the experimental results is effectively reduced, the limited data resources are maximized, and the reliability of model evaluation is improved. In addition, this scheme contains wave breaking information frame by frame, which supports users to flexibly design data enhancement strategy and frame-level sampling scheme according to their own research needs. Overall, the data set and model framework constructed by this scheme provide a solid data support and technical foundation for wave type classification tasks.
[0139] A5, model deployment and application
[0140] After completing model training and verification, the trained deep learning model is deployed in the nearshore video monitoring system to realize real-time classification and recognition of wave types. The deployment mode can select edge computing devices (such as embedded GPU modules) or cloud server architecture according to application requirements to meet different data processing and response requirements in different scenarios.
[0141] The specific application process includes:
[0142] 1) Real-time reception of nearshore video streams collected by front-end camera equipment;
[0143] 2) According to the A2 standard process, the video frames are preprocessed (segmentation, cropping, frame extraction, etc.);
[0144] 3) Use the deployed model to infer the frame sequence and output the corresponding wave type label;
[0145] 4) Real-time display or upload the classification results to the remote platform for coastal state monitoring, early warning system triggering, and topography analysis applications.
[0146] In addition, this method has good generality and scalability. In different application scenarios of different terrain, climate or hydrodynamic environment, the existing model can be fine-tuned through transfer learning, or targeted training combined with new regional data, so as to further improve the classification accuracy and adaptability. This feature makes this method flexible to meet the diversified coastal observation needs, and has a wide engineering promotion prospect.
[0147] Based on the embodiment, the present application has the following advantages:
[0148] (1) Automation processing, high recognition efficiency: Traditional wave breaking type classification relies on manual observation or rule-driven empirical methods, which not only has low efficiency and high labor intensity, but also has strong subjectivity and poor stability. The present application uses a deep learning model to automatically analyze video images, realizing real-time and unattended recognition of wave breaking types, greatly improving efficiency and automation.
[0149] (2) High classification accuracy and strong robustness: The present application builds a spatiotemporal feature extraction deep learning network (such as CRNN, 3D CNN, LSTM or Transformer) to model video frame sequences, which can effectively capture dynamic change features in the wave breaking process. In the multi-fold cross-validation experiment, the model shows high classification accuracy and stability, with good generalization ability.
[0150] (3) Support for diversified sea conditions and terrain conditions, strong adaptability: Compared with traditional methods that rely on Iribarren number and other empirical formulas, the present application does not rely on complex physical parameter calculations and shows stronger adaptability in complex terrain or measurement limited areas. Through transfer learning or regional fine-tuning, the model accuracy can be quickly deployed in new areas. Combined with high-frequency video acquisition and intelligent recognition system, the present application can be widely used in coastal engineering design, marine disaster warning, ecological environment monitoring and other fields, with significant engineering promotion and practical application value.
[0151] Example Three
[0152] Please refer to Figure 7 A nearshore wave breaking type classification system based on video image deep learning, based on the method of example one or example two, comprising:
[0153] Acquisition module: for acquiring video image data of the nearshore area based on fixed or temporary video acquisition equipment, the video image data covering the high position of the nearshore breaking wave belt;
[0154] Preprocessing module: for preprocessing video image data and labeling overflow, roll wave and collapse wave according to data;
[0155] Training module: for constructing a training set based on the labeled data, and training a neural network model based on the training set;
[0156] Prediction module: for feature extraction and time series modeling of nearshore wave image sequences based on the trained neural network model, to predict the corresponding wave breaking type of the nearshore wave image sequence.
[0157] In the above technical solution, in order to better use the method described in one of the embodiments or the second embodiment, the application proposes a nearshore wave breaking type classification system based on video image deep learning, each module corresponds to each step of the above method, and the specific principle has been described above, which will not be repeated here.
[0158] Embodiment four
[0159] A nearshore wave breaking type classification device based on video image deep learning, comprising:
[0160] At least one processor and a memory in communication connection with the at least one processor;
[0161] Wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in one of the embodiments or the second embodiment.
[0162] In the above technical solution, in order to better run and process the method described in one of the embodiments or the second embodiment, the above method is stored in the memory, and the processor is used to execute the stored method. It should be noted that the principle and effect of each step have been described above, which will not be expanded here.
[0163] Embodiment five
[0164] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the method described in one of the embodiments or the second embodiment.
[0165] In the above technical solution, in order to better run and use the method described in one of the embodiments or the second embodiment, the above method is stored in the computer readable storage medium, and the processor is used to realize the above method. It should be noted that the principle and effect of each step have been described above, which will not be expanded here.
[0166] The above is only some embodiments of the application, not to limit the protection scope of the application, any equivalent device or equivalent process transformation using the content of the application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the application.
Claims
1. A nearshore wave breakage type classification method based on deep learning of video images, characterized in that, The method includes: Based on fixed or temporary video acquisition equipment, video image data of the nearshore area is collected, and the video image data view covers the high position of the nearshore wave-breaking zone. The video image data is preprocessed, and spillover, convolution, and collapse are labeled according to the data. A training set is constructed based on the labeled data, and a neural network model is trained based on the training set. Based on the trained neural network model, feature extraction and time-series modeling are performed on nearshore wave image sequences to predict the wave breaking type corresponding to the nearshore wave image sequences. Specifically: A pre-trained convolutional network is used as a feature extractor to encode the image sequence; a long short-term memory network is used to decode the encoded sequence to generate a temporal feature representation; and a classification label for the wave breakage type is determined based on the temporal feature representation.
2. The nearshore wave breakage type classification method based on deep learning of video images as described in claim 1, characterized in that, The video image data is preprocessed, and spillover, convolution, and collapse are labeled according to the data, including: The video image data is divided into multiple video slices, each lasting 1 to 30 seconds. The segmented time segments are then cropped, focusing on the wave-breaking area, to obtain the cropped video slices. The video slices were labeled, and the wave breaking types were classified into spill waves, rolling waves, and collapse waves based on visual characteristics; The labeled video slices are processed by frame extraction to generate an image sequence suitable for model input.
3. The nearshore wave breakage type classification method based on deep learning of video images as described in claim 1, characterized in that, The video image data is preprocessed, and spillover, convolution, and collapse are labeled according to the data, including: The video image data is divided into multiple video slices in segments of 1 to 30 seconds each. The segmented time segments are then cropped at multiple scales. Small-scale foam details and medium-to-large-scale wave morphology are preserved in the wave breakage area to generate multi-scale video slices. Feature enhancement is performed on multi-scale video slices to generate feature-enhanced video slices; Annotate the feature-enhanced video slices to generate an annotated wave-breaking image; The labeled video slices are processed by frame extraction to generate an image sequence suitable for model input.
4. The nearshore wave breaking type classification method based on deep learning of video images as described in claim 3, characterized in that, Feature enhancement is performed on multi-scale video slices to generate feature-enhanced video slices, including: Feature enhancement is performed on multi-scale video slices. The Canny edge detection operator is used to strengthen the peak contour, the brightness gain coefficient is set to highlight the foam contrast, and local binary pattern texture features are extracted to enhance dynamic texture, thus generating feature-enhanced multi-scale video slices.
5. A nearshore wave breaking type classification method based on deep learning of video images as described in any one of claims 1-3, characterized in that, Based on the data, spillover, convolution, and collapse waves are labeled separately, including: The type of wave is determined by the significant visual features in the slice. If it shows a gentle flow and low wave peaks, it is labeled as an overflow wave. If it shows tubular curls and white foam, it is labeled as a curling wave. If it shows splashing and breaking, it is labeled as a collapse wave.
6. A nearshore wave breaking type classification method based on deep learning of video images as described in claim 2 or 3, characterized in that, When cropping the segmented time segments, the time interval between wave-breaking slices in close proximity exceeds 1 minute.
7. A nearshore wave breakage type classification system based on deep learning from video images, characterized in that, The method according to any one of claims 1-6 includes: Acquisition module: used to acquire video image data of the nearshore area based on fixed or temporary video acquisition equipment, wherein the video image data view covers the high position of the nearshore wave-breaking zone; Preprocessing module: used to preprocess the video image data and label the spill, convolution and collapse waves according to the data; Training module: Used to construct a training set based on labeled data and train a neural network model based on the training set; Prediction module: Based on the trained neural network model, it performs feature extraction and temporal modeling on nearshore wave image sequences to predict the wave breakage type corresponding to the nearshore wave image sequences. Specifically, it uses a pre-trained convolutional network as a feature extractor to encode the image sequences; it decodes the encoded sequences through a long short-term memory network to generate a temporal feature representation; and it determines the classification label of the wave breakage type based on the temporal feature representation.
8. A nearshore wave breakage type classification device based on deep learning of video images, characterized in that, include: At least one processor and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.