A method for detecting nearshore breaking wave propagation trajectory guided by biological physical prior in video
Patent Information
- Application Number
- CN202610932512.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-08-04
AI Technical Summary
[0009]综上所述,现有技术虽然为近岸波浪监测与视频测波分析提供了基础,但仍缺乏一种能够同时实现以下目标的技术方案:一是从Timestack影像中内生提取破波活动区域空间先验;二是利用所述先验引导深度学习模型对近岸破波时序传播轨迹进行自动、连续检测;三是基于所检测轨迹进一步计算破碎位置、传播速度、传播距离和持续历时等动力学参数
(1)本发明能够从视频时序统计特征中内生提取破波活动区域空间先验,不依赖外部环境数据。本发明不是依赖外部水深、潮位、地形或数值模型结果来限定破波发生区域,而是直接基于Timestack影像自身的时序统计特征,内生构建破波活动区域空间先验。该方法降低了对外部多源数据的依赖,减少了数据获取、维护及时空匹配带来的复杂性,具有更好的可移植性和工程应用便利性。
Smart Images

Figure CN122510801A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of marine observation technology and artificial intelligence technology, and in particular to a method for detecting nearshore wave-breaking propagation trajectories guided by video-based endogenous physics priors. Background Technology
[0002] Nearshore wave breaking is a key process in beach dynamic geomorphological evolution, sediment transport, wave energy dissipation, and shoreline safety assessment. As waves propagate from the open sea towards the shore, they gradually become shallower and steeper under the combined effects of decreasing water depth and topographic changes, eventually breaking up. Information on the location of breakup, propagation path, duration, and propagation speed is crucial for revealing nearshore hydrodynamic processes and beach response mechanisms. Especially in natural beach environments, wave breaking typically exhibits significant spatial migration, process nonlinearity, and typological differences. Therefore, how to continuously, automatically, and quantitatively monitor wave breaking processes in natural beach environments has always been a significant technical challenge in marine observation and nearshore dynamic research.
[0003] For a long time, nearshore wave breaking observation has mainly relied on pressure sensors, current meters, buoys, and other in-situ measurement methods. These methods can acquire local wave elements or hydrodynamic parameters and have high point measurement accuracy, but they generally suffer from limited spatial coverage, insufficient continuity, and high long-term maintenance costs, making it difficult to meet the needs of reconstructing long-term, event-level wave breaking processes under natural beach conditions. In contrast, shore-based video monitoring has advantages such as non-contact operation, high frequency, strong continuity, and wide coverage. It can continuously record wave propagation and breaking processes over a large spatial area and has gradually become an important technical approach for nearshore wave breaking research. In particular, video-derived data formats such as timestack images can compress and characterize the wave propagation and breaking process in the transshore direction into a two-dimensional spatiotemporal structure, presenting the wave propagation trajectory as brightness stripes with a certain tilt angle, providing an effective data carrier for the continuous analysis of single wave breaking events.
[0004] Existing video measurement and analysis technologies have made some progress in wave parameter extraction, nearshore topography inversion, and wave-breaking feature recognition. For example, some existing technologies use sea surface video images to generate time-stack images and combine image transformation methods to extract overall wave parameters such as wave direction, wave speed, wavelength, and period. Other technologies extract feature parameters such as wave-breaking point location, wave height, and wave depth through pixel sampling, temporal moving average, white pixel region identification, and threshold segmentation of nearshore video images. These technologies demonstrate the clear feasibility of nearshore wave information extraction based on video images and lay the foundation for the development of subsequent methods. Represented by CN101034004A and CN120147920B, these methods embody two typical technical routes: video wave parameter measurement and rule-based nearshore wave-breaking feature extraction, respectively.
[0005] However, considering the monitoring needs of nearshore wave propagation processes in complex natural beach scenarios, existing technologies still have significant limitations. First, most existing methods focus on extracting parameters or feature points such as wave direction, wave velocity, wavelength, period, and wave-breaking point location, lacking dedicated automatic detection technology for the thin, continuously extending wave-breaking time-series propagation trajectories in Timestack imagery. Second, existing technologies rely heavily on traditional digital image processing methods such as Radon transform, temporal moving average, white pixel recognition, threshold segmentation, or geometric relationship calculation. When the observation angle, lighting conditions, sea surface reflection, foam residue, and sea state background change, the stability of these methods tends to decrease, making them unsuitable for the complex and ever-changing real-world scenarios of natural beaches. Third, existing methods often focus more on local brightness anomalies or feature point recognition, paying insufficient attention to the continuous temporal and spatial structure of wave-breaking propagation trajectories. Therefore, under conditions of overlapping adjacent events, complex wave group separation, and strong textured background interference, problems such as false detections, missed detections, or trajectory breaks are prone to occur.
[0006] In recent years, deep learning methods have been gradually applied to nearshore video image analysis, demonstrating significant automation potential in tasks such as wave break location detection, foam identification, wave-by-wave scale tracking, and related video interpretation. Compared to traditional methods, deep learning can automatically learn multi-scale representation features of wave break trajectories from a large number of samples, eliminating reliance on fixed thresholds, single image transformations, or manually set rules. Therefore, it is more suitable for handling complex wave break morphologies, strong background disturbances, and overlapping adjacent events in natural beach scenes. For thin, continuous targets in Timestack imagery, it is necessary not only to identify local brightness changes but also to maintain the integrity and continuity of the trajectory structure in both temporal and spatial dimensions. Learning-based detection methods show significant potential in this regard.
[0007] Further analysis reveals that existing deep learning methods largely rely on image pixel statistical features as the primary driver. These models tend to easily learn brightness distribution, foam texture, and background structure under specific site conditions, but lack explicit constraints on the spatial distribution patterns of wave breaking. Under different site conditions, tide levels, topography, and lighting backgrounds, these purely visual-driven methods are prone to cross-scene performance degradation, manifesting as increased false positives, increased false negatives, and decreased trajectory continuity. The reason for this is that wave breaking on natural beaches does not occur randomly across the shore, but is typically concentrated in a limited area of wave-breaking activity controlled by tide level, water depth, and topography. Compared to simple visual texture, the spatial concentration of wave-breaking activity areas is itself a more stable and physically meaningful piece of information. However, current technology lacks a method to intrinsically extract this spatial prior from the temporal statistical features of the video itself and use it to guide the detection process.
[0008] Furthermore, some existing technologies rely on external water depth, tide level, topography, or empirical relationships for auxiliary calculations during the wave-breaking parameter inversion process. While these methods have certain applicability under specific conditions, in multi-site, long-term, and complex environment applications, the acquisition, maintenance, and spatiotemporal matching of external data all increase deployment complexity and application costs, and also limit the portability and scalability of the methods to some extent. In contrast, if the spatial distribution patterns of wave-breaking activity areas can be extracted directly from the temporal statistical characteristics of Timestack imagery itself, and this can be introduced into the detection model as prior information, it is expected that the model's ability to focus on real wave-breaking activity areas can be improved without relying on external environmental data, reducing false detections caused by non-wave-breaking bright fringes, reflections, and foam residues, and enhancing the model's stability in complex backgrounds and cross-scene conditions.
[0009] In summary, while existing technologies provide a foundation for nearshore wave monitoring and video wave measurement analysis, a solution is still lacking that can simultaneously achieve the following objectives: first, to intrinsically extract spatial priors of wave-breaking activity areas from Timestack imagery; second, to use these priors to guide a deep learning model for automatic and continuous detection of nearshore wave-breaking time-series propagation trajectories; and third, to further calculate dynamic parameters such as breakage location, propagation velocity, propagation distance, and duration based on the detected trajectories. Therefore, it is necessary to propose an intelligent detection method for nearshore wave-breaking time-series propagation trajectories based on intrinsic physical priors. This method aims to improve the accuracy, continuity, and cross-scenario adaptability of nearshore wave-breaking trajectory detection, and to provide technical support for long-term continuous monitoring and event-level analysis of wave breaking processes on natural beaches. Summary of the Invention
[0010] In view of this, the purpose of this invention is to propose a nearshore wave breaking propagation trajectory detection method guided by video-derived intrinsic physical priors. This method uses spatiotemporal stacked images derived from shore-based video as input. Through analysis of the temporal statistical features of the images, it intrinsically constructs spatial priors for wave breaking activity areas. These priors are then used to guide a deep learning model to automatically detect and continuously represent the temporal propagation trajectory of nearshore wave breaking. Simultaneously, based on the detected trajectory, it further automatically calculates dynamic parameters such as breaking location, propagation velocity, propagation distance, and duration. This invention can be applied to scenarios such as long-term continuous monitoring of wave breaking processes on natural beaches, identification of wave breaking activity areas, extraction of wave propagation features, and analysis of nearshore dynamic geomorphological processes, providing core algorithmic support for shore-based intelligent video monitoring systems.
[0011] According to one aspect of the present invention, a method for detecting nearshore wave-breaking propagation trajectories guided by video-based endogenous physics priors is provided, the method comprising: Acquire shore-based video images, set up a cross-shore sampling section within the video field of view, extract pixel brightness values frame by frame along the cross-shore sampling section and stack them in time order to generate a spatiotemporal stacked image. Temporal statistical feature analysis is performed on the spatiotemporal stacked images to generate endogenous physical priors. The endogenous physical priors include a characterization of enhanced wave-breaking activity and a two-dimensional spatial prior. The characterization of enhanced wave-breaking activity is obtained by weighting the normalized images with the intensity of cross-shore wave-breaking activity. The two-dimensional spatial prior is copied from the one-dimensional spatial prior in the cross-shore direction along the time direction. The one-dimensional spatial prior is determined based on the cross-shore distribution characteristics of wave-breaking activity intensity. A wave-breaking temporal propagation trajectory detection network is constructed. The wave-breaking activity enhancement representation is embedded into the feature extraction process of the model, and the two-dimensional spatial prior is introduced as a spatial weight into the loss function of the training end of the detection network. The wave-breaking activity enhancement representation and the two-dimensional spatial prior are used to guide the detection network to perform pixel-level detection of the wave-breaking temporal propagation trajectory in the spatiotemporal stacked image, and output the continuous geometric shape representation result of the wave-breaking propagation trajectory. Based on the continuous geometric morphology characterization results of the wave breaking trajectory, the vectorized propagation trajectory of a single wave breaking event is extracted, and the dynamic parameters of the wave breaking process are automatically calculated based on the vectorized propagation trajectory.
[0012] In the aforementioned technical solution, a detection method based on endogenous physical priors is proposed to address the problem of automatic detection of nearshore wave breaking propagation trajectories in shore-based video monitoring. Its core lies in endogenously constructing enhanced wave-breaking activity representations and two-dimensional spatial priors from the temporal statistical features of the spatiotemporal stacked images themselves. These priors then guide a deep learning model to achieve pixel-level continuous detection of wave-breaking trajectories, thereby automatically calculating key dynamic parameters.
[0013] Unlike existing detection methods that rely on external water depth, tide data, or purely visual-driven approaches, this invention first extracts the intensity distribution of wave-breaking activity through variance analysis along the time dimension of spatiotemporal stacked images. This generates two types of prior information with clear physical meaning: first, a wave-breaking activity enhancement characterization, obtained by weighting standardized images with cross-shore wave-breaking activity intensity, which highlights the wave-breaking activity area and suppresses background noise at the input; second, a two-dimensional spatial prior, replicated from the cross-shore one-dimensional prior along the time direction, quantitatively characterizing the spatial concentration range of the wave-breaking zone. Existing technologies either employ traditional image processing methods such as fixed thresholds and Radon transforms, which lack stability under complex lighting and wave group overlap conditions; or they use deep semantic segmentation networks that only output wave-breaking foam areas, making it difficult to obtain the continuous propagation trajectory of a single wave. This invention integrates the aforementioned intrinsic physical priors into the detection network in two complementary ways: firstly, it embeds the enhanced representation of wave-breaking activity into the network's feature extraction process, enabling the model to prioritize physically plausible wave-breaking activity regions during training and inference; secondly, it introduces two-dimensional spatial priors as spatial weights into the loss function, allowing the model to impose a higher penalty on prediction errors of pixels within the wave-breaking band and suppress errors of out-of-band interfering pixels. This dual-layer guidance mechanism of "feature-level embedding + loss-level weighting" significantly improves the continuity and anti-interference capability of trajectory detection compared to baseline models that rely solely on data-driven approaches. It maintains a more complete and stable trajectory structure even in scenarios such as complex wave group separation, interference from non-wave-breaking targets, and strong texture backgrounds. Furthermore, this invention extracts the vectorized propagation trajectory of a single wave-breaking event based on the output continuous geometric morphology representation results and automatically calculates dynamic parameters such as breaking location, wave velocity, propagation distance, duration, and breaking period. This achieves closed-loop analysis from trajectory detection to dynamic process quantification, overcoming the limitations of existing technologies that can only extract discrete wave-breaking points or overall wave parameters.
[0014] This invention significantly improves the continuity and cross-scene robustness of wave propagation trajectory detection by endogenously constructing wave-breaking activity enhancement representations and two-dimensional spatial priors from temporal stacked images, and guiding the detection network in a two-layer manner of feature embedding and loss weighting, without relying on external environmental data.
[0015] In some embodiments, generating the enhanced characterization of the wave-breaking activity specifically includes: standardizing the spatiotemporal stacked image to obtain a standardized image; calculating the gray-level variance of each cross-shore location along the time dimension to obtain the intensity of the wave-breaking activity; and weighting the standardized image with the intensity of the wave-breaking activity to obtain the enhanced characterization of the wave-breaking activity.
[0016] In the above technical solution, the intensity of wave activity at each cross-shore location is quantified by calculating the gray-scale variance of the standardized image along the time dimension. Then, the intensity is used as a spatial weight and fused with the standardized image to achieve signal enhancement in the wave-breaking activity area and suppression of the non-wave-breaking background.
[0017] Existing technologies for extracting wave-breaking information from spatiotemporal stacked images often rely on raw pixel brightness or simple threshold segmentation, lacking utilization of the spatiotemporal statistical characteristics of wave activity. While some deep learning methods can learn features, the input remains the raw image or a single standardized result, without explicitly injecting the spatial distribution patterns of wave-breaking activity. The core of this feature lies in: firstly, standardizing the spatiotemporal stacked image to eliminate systematic biases caused by illumination, exposure, and sea surface reflection at different observation times; then, calculating the gray-level variance of each shore-crossing location along the time dimension—this variance directly reflects the intensity of pixel brightness fluctuations over time at that location. In areas where wave-breaking occurs frequently, the brightness abrupt changes caused by wave breaking will result in high variance values, while calm water surfaces or stable foam areas will have lower variances; finally, the obtained wave-breaking activity intensity is used as a spatial weight and multiplied pixel-by-pixel with the standardized image to generate an enhanced wave-breaking activity representation. The larger the variance at a shore-crossing location, the stronger the wave-breaking activity, and the corresponding pixels in the weighted area are amplified; locations with near-zero variance are suppressed. Compared to the original image or simple normalized results, this enhanced representation highlights the temporal dynamics within the wave-breaking zone at the input end, while suppressing static or weakly dynamic backgrounds such as sand and reflections. Using this enhanced representation as an additional channel input to the detection network allows the model to directly obtain prior knowledge of "which regions are more temporally active" during the first convolutional layer, thereby guiding the network to focus on the actual wave-breaking propagation trajectory and reducing the learning of interference from non-wave-breaking bright fringes, foam residue, etc.
[0018] In some embodiments, generating the two-dimensional spatial prior specifically includes: calculating the variance of the wave-breaking activity enhancement characterization along the time dimension to obtain the cross-shore wave-breaking activity intensity distribution; determining at least one wave-breaking activity area boundary based on the cross-shore wave-breaking activity intensity distribution; defining a one-dimensional spatial prior in the cross-shore direction for each wave-breaking activity area boundary; and copying the one-dimensional spatial prior along the time direction to obtain the corresponding two-dimensional spatial prior.
[0019] In the above technical solution, based on the enhanced characterization of wave-breaking activities, the variance is calculated twice along the time dimension to obtain the intensity distribution of cross-shore wave-breaking activities, and then a two-dimensional prior mask covering the wave-breaking activity area is constructed through boundary detection and spatial replication.
[0020] Existing technologies, when defining the wave-breaking area, either rely on external water depth, tide level, and topographic data, or use fixed spatial windows, making it difficult to adapt to the dynamic migration of the wave-breaking zone under different station, tide level, and dynamic conditions. Some methods directly predict foam areas through semantic segmentation, but lack explicit quantification of the distribution pattern of wave-breaking activity intensity along the trans-shore direction. The essential innovation of this feature lies in: firstly, calculating the variance of the generated enhanced wave-breaking activity representation along the time dimension again. The first variance (used when generating the enhanced representation) is used to extract the wave-breaking activity intensity at each trans-shore location, while the second variance, based on the spatial weights already applied to the enhanced representation, further characterizes the fluctuation energy distribution of wave-breaking activity intensity in the time dimension, thus obtaining a more robust trans-shore wave-breaking activity intensity distribution curve. This curve typically exhibits a single-peak or multi-peak structure, with the peak corresponding to the trans-shore location with the most concentrated wave-breaking activity, and the peak width reflecting the spatial extension range of the wave-breaking zone. Secondly, based on this intensity distribution, methods such as peak detection, threshold segmentation, or region growing are used to determine the boundaries of one or more wave-breaking activity areas, thereby adaptively extracting the spatial intervals of the wave-breaking zone and handling multi-peak scenarios where the outer main wave-breaking zone and the inner secondary wave-breaking zone coexist. Thirdly, a one-dimensional spatial prior is defined for each wave-breaking activity area in the trans-shore direction (e.g., a binary prior with a value of 1 within the interval and 0 outside the interval, or a soft prior that continuously changes with distance from the center). Then, a simple copy is performed along the time direction to generate a two-dimensional spatial prior with the same size as the original spatiotemporal stacked image. This copying operation is based on the physical fact that the spatial position of the wave-breaking zone is relatively stable in the short term, and is therefore reasonable. Compared to schemes that rely on external data or fixed windows, the two-dimensional spatial prior generated by this feature is entirely derived from the temporal statistical features of the same video itself, requiring no auxiliary information, and can automatically adjust with changes in tide level and wave conditions. By introducing this prior as a spatial weight into the loss function of the detection network, the model imposes a higher penalty on the prediction error of pixels within the wave-breaking zone and significantly reduces the weight on the error of interfering pixels outside the zone. This effectively suppresses false detections of non-wave-breaking targets such as beach reflections and oblique bright streaks on the sea surface, while enhancing the continuity of trajectory detection within the wave-breaking zone.
[0021] In some embodiments, the training loss function is a loss function that incorporates the two-dimensional spatial prior as spatial weights; the pixel-level loss function adopts any one or a combination of weighted cross-entropy, Focal Loss, Dice Loss, Tversky Loss, IoU Loss, boundary preservation loss, connectivity constraint loss, and curve smoothing constraint loss.
[0022] In the above technical solution, the spatial distribution of the loss function is modulated by the spatial prior constructed by endogenous construction, so that the model focuses on the wave-breaking activity region differently during the training process.
[0023] Existing deep learning-based image segmentation or trajectory detection methods typically employ pixel-level cross-entropy, Dice, or Focal Loss as loss functions, assigning equal weight to errors at all spatial locations. These methods have inherent limitations when handling complex natural beach scenes: wave-breaking activity is concentrated within a limited trans-shore zone, while pixels in non-wave-breaking areas (such as sand, calm sea surfaces, or foam remnants) constitute the majority of the spatiotemporal image stack. If the loss function uniformly weights the entire image, the model is prone to overfitting the dominant background pixel class while failing to adequately learn the sparse wave-breaking trajectory pixels, leading to trajectory breaks or numerous false detections in the detection results. This paper introduces the intrinsically extracted two-dimensional spatial prior P(x,t) from step S2 as a spatial weighting factor into the loss function. This prior has a clear physical meaning in the cross-shore direction. Specifically, this feature multiplies or weights the prior weights with pixel-level loss functions (such as weighted cross-entropy, Focal Loss, Dice Loss, Tversky Loss, IoULoss, and even optional boundary preservation loss, connectivity constraint loss, or curve smoothing constraint loss) pixel by pixel. The resulting technical effects are multifaceted: First, the loss contribution of pixels within the break zone is preserved or even enhanced, forcing the model to pay more attention to the prediction accuracy of these key regions, thereby improving the recall and continuity of trajectory detection; Second, the loss contribution of non-break zone regions outside the break zone is significantly suppressed, and the model no longer spends capacity learning background interference such as beach reflections and oblique bright streaks on the sea surface, significantly reducing false positives; Third, this spatial weighting mechanism can be seamlessly superimposed on any basic loss function, forming joint supervision with auxiliary losses such as boundary preservation and connectivity constraints, further strengthening the topological integrity and geometric smoothness of the trajectory. Compared to traditional uniformly weighted loss functions, this feature introduces data-native physical space priors, enabling spatial attention guidance during the training phase. This allows the model to quickly focus on the actual wave-breaking activity area even with limited samples, improving training sample efficiency and generalization ability. Furthermore, the feature's compatibility with multiple loss functions provides excellent flexibility—the most suitable loss combination can be selected for different beach scenarios or monitoring needs without altering the core architecture guided by the priors.
[0024] In some embodiments, the training loss function is a physically guided hybrid loss function:
[0025] in, For two-dimensional space prior, These are the prior weighting coefficients. To modulate Focal Loss, For Tversky Loss, and This is the balance coefficient.
[0026] In the above technical solution, the loss function is composed of a combination of modulated Focal Loss and Tversky Loss with two-dimensional spatial prior weighting, which aims to solve the two technical problems of extreme imbalance between positive and negative samples and preservation of topology in wave trajectory detection at the same time.
[0027] Existing pixel-level classification-based detection methods face two inherent difficulties when processing spatiotemporally stacked images: First, the proportion of broken wave trajectory pixels (positive samples) in the entire image is extremely low, while background pixels (negative samples) account for the vast majority, causing the model under standard cross-entropy loss to be biased towards predicting the background and prone to trajectory breaks; Second, as a thin, continuous target, the topological connectivity of the trajectory is more important than the accuracy of simple pixel classification, and traditional loss functions lack explicit constraints on structural integrity. This paper proposes a specific form of combined loss function, the advantages of which are reflected in three aspects.
[0028] First, a multiplicative weighting factor of two-dimensional spatial prior is incorporated into the modulated Focal Loss term. Focal Loss itself uses a modulation factor to focus the model on hard-to-classify samples, and this feature further... As a location-related weighting coefficient: within the wave-breaking activity area Larger values amplify the loss weight. The weights are approximately 1 times the normal value outside the region; outside the region, the weights are close to 1. This design differs from conventional spatially weighted loss in that the prior weights and the focusing mechanism of Focal Loss complement each other. The former guides the model to focus on physically reasonable wavebands from a spatial location perspective, while the latter guides the model to focus on pixels with blurred classification boundaries from a sample difficulty perspective. Combined, the model neither wastes capacity learning out-of-band background nor ignores difficult-to-distinguish in-band samples.
[0029] Second, Tversky Loss is introduced as a structural constraint term. Compared to Dice Loss, which uses a weighted approach to balance false positives and false negatives, Tversky Loss uses parameters... and Independently control the penalty weights for false positives and false negatives; for targets with elongated trajectories, appropriately increase the penalty weights. (i.e., the penalty for false negatives) can effectively suppress trajectory breakage. Combining this with the aforementioned weighted modulated Focal Loss forms a dual supervision of "pixel-by-pixel accuracy and region topology": the modulated Focal Loss is responsible for improving the recognition accuracy of each trajectory pixel, while the Tversky Loss is responsible for maintaining the overall connectivity structure of the trajectory.
[0030] Third, adopt the balance coefficient and The two losses are weighted and combined. The practical value of this design lies in the fact that the sparsity and continuity requirements of trajectories differ across different beach scenarios; by adjusting... and This allows for flexible control over the model's emphasis on pixel accuracy and structural integrity without modifying the underlying loss function. Compared to using cross-entropy, Dice Loss, or Focal Loss alone, the hybrid loss function proposed in this feature unifies intrinsic spatial priors, hard sample focusing, and topology preservation into a single optimization objective for the first time. Experiments show that, guided by this loss function, the model significantly outperforms the pure visual baseline model in terms of trajectory continuity and false detection resistance in complex wavegroup separation and strongly textured background interference scenarios (see specification). Figure 6 Furthermore, this hybrid loss does not require additional post-processing repair steps; instead, structural constraints are directly implemented at the training stage, reducing computational overhead during inference.
[0031] In some embodiments, the continuous geometric morphology representation result of the wave propagation trajectory includes any of the following: pixel-level trajectory probability map, binary trajectory mask, center line, broken line sequence, spline curve, curve parametric equation.
[0032] The above technical solution provides a flexible representation interface to adapt to different downstream analysis needs. Different output formats serve different application scenarios, and multiple optional formats make the method highly adaptable to various tasks. Specifically: pixel-level trajectory probability maps retain the model's uncertainty estimates for all pixels, making them suitable for scenarios requiring subsequent fine-grained post-processing or probability fusion; binary trajectory masks can be directly used for visualization or region statistics; centerlines and polyline sequences describe the topological skeleton of the trajectory in a lightweight manner, facilitating storage and transmission; spline curves and curve parametric equations provide a smooth geometric representation, which is beneficial for subsequent calculation of differential geometric quantities such as trajectory curvature and tangent direction.
[0033] In some embodiments, the dynamic parameters include breaking wave velocity, breaking location, propagation distance, duration, and breaking period; the breaking wave velocity is obtained by the local slope at the beginning of the trajectory, the breaking location is obtained by the cross-shore coordinates of the starting point of the trajectory, the propagation distance is obtained by the cross-shore displacement or arc length between the starting and ending points of the trajectory, the duration is obtained by the time difference between the ending and starting points of the trajectory, and the breaking period is obtained by the difference between the starting times of two adjacent wave breaking events.
[0034] The aforementioned technical solution covers parameters such as breaking wave velocity, breaking location, propagation distance, duration, and breaking period, and specifies the detailed calculation method for each parameter extracted from the trajectory's geometric features. Existing video-based wave parameter extraction methods often employ Radon transform, spatiotemporal stack brightness peak detection, or global wavefront tracking. The output is often a statistically significant average value for wave direction, wave velocity, and period, or the location of a single breaking point. These methods struggle to automatically correlate the complete process parameters of the same wave breaking event from its inception to its dissipation, and are even less capable of obtaining event-level propagation distance, duration, and breaking periods of adjacent events. The essence of this feature lies in: using a vectorized propagation trajectory as input, directly analyzing the dynamic parameters with clear physical meaning using the trajectory's geometric shape in the spatiotemporal coordinate system.
[0035] According to another aspect of the present invention, a nearshore wave-breaking propagation trajectory detection device guided by video endogenous physics priors is provided, the device comprising, based on the above-described method: The video acquisition module is used to acquire shore-based video images, set up a cross-shore sampling section within the video field of view, extract pixel brightness values frame by frame along the cross-shore sampling section and stack them in time order to generate a spatiotemporal stacked image. An endogenous physics prior generation module is used to perform temporal statistical feature analysis on the spatiotemporal stack image to generate endogenous physics priors. The endogenous physics priors include a wave-breaking activity enhancement characterization and a two-dimensional spatial prior. The wave-breaking activity enhancement characterization is obtained by weighting the normalized image with the trans-shore wave-breaking activity intensity. The two-dimensional spatial prior is copied from the one-dimensional spatial prior in the trans-shore direction along the time direction. The one-dimensional spatial prior is determined based on the trans-shore distribution characteristics of the wave-breaking activity intensity. The trajectory detection module has a built-in wave-breaking temporal propagation trajectory detection network, which is used to embed the wave-breaking activity enhancement representation into the feature extraction process of the model, and introduce the two-dimensional spatial prior as spatial weight into the training loss function of the detection network. The wave-breaking activity enhancement representation and the two-dimensional spatial prior are used to guide the detection network to perform pixel-level detection of the wave-breaking temporal propagation trajectory in the spatiotemporal stacked image, and output the continuous geometric shape representation result of the wave-breaking propagation trajectory. The parameter calculation module is used to extract the vectorized propagation trajectory of a single wave breaking event based on the continuous geometric morphology characterization results of the wave breaking trajectory, and automatically calculate the dynamic parameters of the wave breaking process based on the vectorized propagation trajectory.
[0036] In order to better utilize the above method, this application proposes a nearshore wave-breaking propagation trajectory detection device guided by video endogenous physics priors. Each module corresponds to a step of the above method, and its specific principle has been described above and will not be repeated here.
[0037] According to another aspect of the present invention, a near-shore wave-breaking propagation trajectory detection device guided by video-endogenous physics priors is provided, comprising: At least one processor and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described above.
[0038] In the above technical solution, to better operate and process the method, the method is stored in memory, and the processor executes the stored method. It should be noted that the principle and effect of each step have been described above and will not be elaborated upon here.
[0039] According to another aspect of the present invention, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the above-described method.
[0040] In the above technical solution, to better operate and use the method, the method is stored in a computer-readable storage medium and implemented using a processor. It should be noted that the principle and effect of each step have been described above and will not be elaborated upon here. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a flowchart illustrating an embodiment of a video-driven, physics-guided method for detecting nearshore wave-breaking propagation trajectories according to the present invention. Figure 2 This is a flowchart of the Timestack image generation process of an embodiment of the video-based endogenous physics prior-guided nearshore wave propagation trajectory detection method of the present invention: (a) the cross-shore sampling section in the original video; (b) the standardized sampling section after reprojection; (c) the Timestack spatiotemporal stack image. Figure 3 (a) A Timestack image of an embodiment of a video-based endogenous physics-guided nearshore wave-breaking propagation trajectory detection method of the present invention; (b) Enhanced wave-breaking activity characterization (c) Cross-shore wave-breaking activity profile; (d) Prior information on dynamic wave-breaking activity zone. ; Figure 4 This is an experimental result of the sensitivity of the prior construction parameters of the dynamic wave-breaking activity area in an embodiment of the nearshore wave-breaking propagation trajectory detection method guided by video endogenous physics priors of the present invention: (a) Mean IoU under different parameter combinations; (b) Mean Recall under different parameter combinations; Figure 5 This is an embodiment of the BreakTrajNet-PG overall architecture diagram of a video-based endogenous physics prior-guided nearshore wave propagation trajectory detection method of the present invention: LiteDexi backbone network and input-end enhancement, training-end physics prior weighting mechanism; Figure 6 This is a comparison of the detection results of a pure visual baseline model and BreakTrajNet-PG in a complex scene according to an embodiment of the video-based endogenous physics prior-guided nearshore wave-breaking propagation trajectory detection method of the present invention: (a–c) complex wave group separation scene; (d–f) non-wave-breaking target interference scene; (g–i) strong texture background interference scene; Figure 7 This is a schematic diagram of an embodiment of a nearshore wave-breaking time-series propagation trajectory detection device based on endogenous physical priors according to the present invention. Detailed Implementation
[0043] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] To address the shortcomings of existing technologies in monitoring and identifying nearshore wave breaking processes on natural beaches, such as insufficient utilization of the spatial patterns of wave breaking activity areas, limited automatic detection capabilities of wave breaking time-series propagation trajectories, and difficulty in automatically extracting event-level dynamic parameters, this invention aims to provide an intelligent detection method for nearshore wave breaking time-series propagation trajectories based on endogenous physical priors. This method enables automatic, continuous, and quantitative characterization of the propagation trajectory of a single wave breaking event from spatiotemporal stacked images derived from shore-based video, and further facilitates the automatic calculation of key dynamic parameters. Specifically, it includes the following three aspects: (1) In view of the fact that existing technologies generally lack explicit utilization of the spatial distribution pattern of wave-breaking activity areas and that some methods need to rely on external water depth, tide level or topographic data, this invention proposes a method to endogenously extract the spatial prior of wave-breaking activity areas from the temporal statistical features of the Timestack image itself. Without the need to introduce additional external environmental data, the method can automatically construct the spatial concentration pattern of wave-breaking activity, thereby improving the portability and cross-scenario applicability of the method.
[0045] (2) In view of the lack of a dedicated detection mechanism for thin, continuous wave propagation trajectories in Timestack images in existing methods, and the problems of false detection, missed detection and trajectory breakage that are prone to occur under complex background, illumination changes, sea surface reflection and wave group overlap, this invention introduces the aforementioned endogenous physical prior into the model learning process to establish an intelligent detection method for wave propagation trajectories in time series, so as to improve the accuracy, continuity and cross-scene adaptability of trajectory detection.
[0046] (3) In view of the problem that existing technologies are difficult to automatically obtain key dynamic parameters of a single wave breaking event from continuous video data, this invention further automatically calculates parameters such as breaking location, propagation speed, propagation distance and duration based on the detected wave breaking propagation trajectory, thereby realizing integrated analysis from propagation trajectory detection to dynamic process quantification, providing technical support for long-term continuous monitoring and event-level research of wave breaking processes on natural beaches.
[0047] Example 1 Please see Figure 1 A method for detecting nearshore wave-breaking propagation trajectories guided by video-based intrinsic physics priors, the method comprising: S1. Acquire shore-based video images, set up a cross-shore sampling section within the video field of view, extract pixel brightness values frame by frame along the cross-shore sampling section and stack them in time order to generate a spatiotemporal stacked image. For example, S1, shore-based video acquisition, parameter calibration, and Timestack image generation; Fixed shore-based video acquisition equipment is deployed in the target coastal area to continuously acquire video images of nearshore wave activity. The video acquisition equipment includes fixed high-definition network cameras, local storage units, and edge computing devices. Internal and external parameter calibrations are performed on the cameras to establish a mapping relationship between image pixel coordinates and real-world coordinates.
[0048] It is important to note that fixed high-definition network cameras are not the only solution for front-end video acquisition. Besides network cameras, industrial cameras, PTZ cameras, binocular video equipment, multi-view synchronous camera systems, or other optical acquisition devices capable of continuously acquiring near-shore video images can also be used. Acquisition equipment can be deployed on building rooftops, observation decks, lighthouses, or other high and stable locations, or installed on shore-based supports, towers, or temporary observation platforms. The front-end system can use a single camera to cover a localized area of wave-breaking activity, or multiple cameras can be used to jointly cover a longer shoreline, forming a zoned monitoring or full-shore monitoring mode. The platform is not limited to a single implementation; the shore-based intelligent video monitoring platform can adopt an overall architecture of "front-end acquisition—edge computing—remote management," so the specific deployment methods between on-site acquisition, edge processing, and remote analysis can be adjusted according to application conditions. Regarding data transmission and processing, all preprocessing and detection tasks can be completed by on-site edge devices, or only video caching and basic screening can be performed on-site, with subsequent Timestack generation, prior extraction, and trajectory detection tasks transferred to a remote server; a hybrid processing mode combining edge and cloud collaboration can also be adopted. Therefore, the present invention is not limited to a fixed hardware deployment form.
[0049] Please see Figure 2 After parameter calibration, a fixed cross-shore sampling section is set within the video field of view. Frames of the original video are extracted at a preset frequency, and the corresponding pixel brightness values are extracted frame by frame along the fixed cross-shore sampling section. These values are then stacked in chronological order to generate a spatiotemporal stack image, Timestack. For any sampling point in the world coordinate system... Its pixel coordinates in the image plane Represented as:
[0050] in, The camera intrinsic parameter matrix, and These are the rotation matrix and translation vector in the extrinsic parameters, respectively. This represents the projection operator from three-dimensional world coordinates to two-dimensional image coordinates. Furthermore, for time... Image The pixel values corresponding to each spatial sampling point are extracted along the cross-section, which can be represented as:
[0051] in, Indicates the first section on the cross-section One cross-border sampling location, This indicates the position at time [time]. The pixel brightness values. By stitching together the sampling results from consecutive moments, a two-dimensional spatiotemporal stacked image is formed with the distance across the shore as one axis and time as the other axis.
[0052] The sampling interval, video frame rate, and sample duration in the cross-shore direction are determined based on actual observation conditions, which include at least camera resolution, observation range, spatial scale of the wave-breaking activity area, wave propagation time scale, and model input requirements. Specifically, the cross-shore sampling interval ensures the spatial continuity of the wave-breaking propagation trajectory, the video frame rate ensures the temporal resolution of the wave propagation and breaking process, and the sample duration ensures that a single sample contains a complete or nearly complete wave-breaking propagation process. Preferably, in one embodiment of the invention, the cross-shore sampling interval is 0.1 m, the video frame rate is 4 Hz, and the sample duration is 60 s, corresponding to 1300 sampling points on the spatial axis and 240 rows on the time axis. The resulting Timestack image can better preserve the spatiotemporal continuity structure of a single wave-breaking propagation trajectory, making it appear as bright continuous stripes with a certain tilt angle.
[0053] It is important to note that for parameter calibration, methods such as the checkerboard method, ground control point method, or other methods that can establish a mapping relationship between image pixel coordinates and real-world coordinates can be used. If the application scenario has low requirements for absolute spatial quantification, only relative geometric correction or simplified projection transformation can be performed, without having to be limited to the exact same calibration process. Regarding the setting of cross-shore sampling sections, a single fixed section, multiple parallel sections, multiple coastal distributed sections, or adaptive section selection schemes can be used. Sections can be set strictly along the cross-shore direction, or they can be tilted or curved according to the actual field of view, shoreline orientation, and wave-breaking zone geometry. For long shorelines or complex coastal topography, multiple Timestack samples can be generated in segments as needed, followed by partitioned detection or result stitching. Timestack samples are essentially two-dimensional reconstructions of fixed section pixels changing over time; therefore, the number, location, and length of sections can be flexibly adjusted according to the monitoring objectives. Regarding the Timestack generation method, grayscale luminance value stacks, RGB three-channel stacks, single-channel luminance component stacks, luminance component stacks after color space transformation, or preprocessing results such as local contrast, texture response, and edge response can be used to generate the Timestack. The video frame sampling frequency, spatial sampling interval, and sample duration can also be set according to the camera resolution, observation range, width of the wave-breaking activity area, wave propagation timescale, and model input requirements, and are not limited to a fixed value. The 0.1 m sampling interval, 4 Hz frame sampling frequency, and 60 s sample length in the preferred embodiment are only a preferred combination and not the only implementation.
[0054] S2. Perform temporal statistical feature analysis on the spatiotemporal stacked image to generate endogenous physical priors; the endogenous physical priors include a wave-breaking activity enhancement characterization and a two-dimensional spatial prior; wherein, the wave-breaking activity enhancement characterization is obtained by weighting the normalized image with the trans-shore wave-breaking activity intensity, and the two-dimensional spatial prior is copied from the one-dimensional spatial prior in the trans-shore direction along the time direction, and the one-dimensional spatial prior is determined based on the trans-shore distribution characteristics of the wave-breaking activity intensity; In this embodiment, generating the enhanced characterization of the wave-breaking activity specifically includes: standardizing the spatiotemporal stacked image to obtain a standardized image; calculating the gray-level variance of each cross-shore location along the time dimension to obtain the intensity of the wave-breaking activity; and weighting the standardized image with the intensity of the wave-breaking activity to obtain the enhanced characterization of the wave-breaking activity.
[0055] In this embodiment, generating the two-dimensional spatial prior specifically includes: calculating the variance of the wave-breaking activity enhancement characterization along the time dimension to obtain the cross-shore wave-breaking activity intensity distribution; determining at least one wave-breaking activity area boundary based on the cross-shore wave-breaking activity intensity distribution; defining a one-dimensional spatial prior in the cross-shore direction for each wave-breaking activity area boundary; and copying the one-dimensional spatial prior along the time direction to obtain the corresponding two-dimensional spatial prior.
[0056] For example, S2, the construction of enhanced characterization of wave-breaking activity and the endogenous extraction of spatial priors of wave-breaking activity regions; combined with Figure 3 The specific steps for example S2 are as follows: The Timestack images generated in step S1 are subjected to sample-by-sample grayscale normalization to reduce systematic brightness differences caused by variations in light intensity, camera exposure settings, and sea surface reflectivity at different sites and observation times. The normalization expression is:
[0057] in, This indicates the location of the Timestack image across the shore. With time grayscale value at that location and These represent the average gray value and standard deviation of the entire time series, respectively. It is a numerically stable term.
[0058] Based on standardization, the gray-scale variance of each cross-shore location is calculated along the time dimension to quantitatively characterize the wave activity intensity at different cross-shore locations:
[0059] Then, the above activity intensity is applied as a spatial weight to the standardized grayscale image to construct a wave-breaking activity enhancement representation:
[0060] Furthermore, the variance of the enhanced characterization is calculated again along the time dimension to obtain the wave-breaking activity profile in the trans-shore direction:
[0061] Determine the location of the main peak of the active profile:
[0062] And a certain percentage of the main peak value (such as 0.01 to 0.2 times the main peak value) is used as the stopping threshold:
[0063] Searching outwards from the main peak location, the initial interval boundaries are determined when the activity intensity drops below the threshold. and Based on this, the initial interval is appropriately expanded:
[0064] Therefore, a one-dimensional prior function for the wave-breaking activity region in the trans-shore direction is defined:
[0065] Then, this one-dimensional prior is copied along the time direction to generate a two-dimensional dynamic wave-breaking activity region spatial prior with the same size as the original Timestack image:
[0066] in, This is the threshold scaling factor. This is the interval expansion factor (e.g., 1 to 2 times). The spatial prior of the two-dimensional dynamic wave-breaking activity area is entirely extracted endogenously from the temporal statistical features of the Timestack image itself, without relying on external water depth, tide level, or topographic data.
[0067] Furthermore, the threshold scaling factor and interval expansion coefficient It can be configured according to the actual application scenario. Preferably, please refer to [link / reference needed]. Figure 4 Based on pre-labeled wave-breaking activity zone samples, the parameters can be optimized through parameter sensitivity experiments. In the parameter sensitivity experiments, the intersection-union ratio (IUU) and recall rate between the prior interval and the actual labeled wave-breaking activity zone are used as evaluation indicators to measure the spatial overlap between the prior interval and the actual wave-breaking activity zone and the completeness of coverage of the actual wave-breaking zone, respectively. Preferably, in one embodiment of the present invention, the parameters are selected as follows: , .
[0068] It should be noted that, in terms of grayscale standardization, in addition to mean-standard deviation standardization, other methods such as maximum-minimum normalization, robust standardization, quantile normalization, local brightness normalization, or adaptive contrast enhancement can also be used to reduce systematic deviations caused by changes in lighting, exposure differences, and sea surface reflection at different sites and time periods.
[0069] In constructing enhanced activity representations, besides using time variance to characterize wave activity intensity, time standard deviation, absolute deviation, local energy, temporal gradient, inter-frame differential cumulative value, local frequency domain energy, and time-frequency response obtained from short-time Fourier transform or wavelet transform can also be used as activity intensity indicators. Correspondingly, enhanced activity representations can employ multiplicative weighting, additive enhancement, exponential weighting, or other equivalent combinations that can highlight the wave-breaking activity area and suppress background disturbances. The essence of the current scheme lies in utilizing the temporal statistical characteristics of Timestack itself to extract the spatial concentration of wave-breaking activity in the cross-shore direction. Therefore, any temporal statistical quantity that can reflect this spatial concentration pattern can, in principle, be used as an alternative.
[0070] In terms of constructing one-dimensional priors, in addition to determining the prior interval based on the main peak position of the activity profile, the threshold ratio coefficient, and the interval expansion coefficient, other methods such as the double threshold method, the local peak detection method, the significant peak screening method, the percentile interval method, the region growth method, the cluster segmentation method, the probability density estimation method, or the fitted distribution method can also be used to determine the boundary of the wave-breaking activity area. The representation of the prior does not have to be limited to a 0 / 1 binary prior; continuous value priors, piecewise weighted priors, Gaussian priors, triangular priors, or other soft prior forms that continuously change with spatial location can also be used.
[0071] Furthermore, the current method is well applicable to scenarios with relatively clear main wave-breaking bands. However, under conditions of strong dynamic environment or multi-peak wave-breaking structure, a multi-peak prior construction strategy can be further developed. The single-peak prior can also be extended to a bi-peak or multi-peak prior. That is, by searching for local peaks and constraining peak significance, multiple main peaks can be transformed into multiple wave-breaking activity zone priors to adapt to the situation where the outer main wave-breaking band and the inner secondary wave-breaking band coexist.
[0072] Regarding the parameter determination method, parameters such as the threshold ratio coefficient and interval expansion coefficient in the prior construction can be determined either through parameter sensitivity experiments using pre-labeled samples, or through empirical setting, adaptive estimation, online updating, or unsupervised optimization. Therefore, it is not required to use manual annotation for implementation.
[0073] S3. Construct a wave-breaking temporal propagation trajectory detection network, embed the wave-breaking activity enhancement representation into the feature extraction process of the model, and introduce the two-dimensional spatial prior as spatial weight into the training loss function of the detection network. Use the wave-breaking activity enhancement representation and the two-dimensional spatial prior to guide the detection network to perform pixel-level detection of the wave-breaking temporal propagation trajectory in the spatiotemporal stacked image, and output the continuous geometric shape representation result of the wave-breaking propagation trajectory. In this embodiment, the training loss function is a loss function that incorporates the two-dimensional spatial prior as spatial weights; the pixel-level loss function adopts any one or a combination of weighted cross-entropy, Focal Loss, Dice Loss, Tversky Loss, IoULoss, boundary preservation loss, connectivity constraint loss, and curve smoothing constraint loss.
[0074] In this embodiment, as a preferred example, the training loss function is a physically guided hybrid loss function:
[0075] in, For two-dimensional space prior, These are the prior weighting coefficients. To modulate Focal Loss, For Tversky Loss, and This is the balance coefficient.
[0076] In this embodiment, the continuous geometric morphology representation result of the wave propagation trajectory includes any of the following: pixel-level trajectory probability map, binary trajectory mask, center line, broken line sequence, spline curve, curve parametric equation.
[0077] For example, S3, intelligent detection of wave propagation trajectory based on endogenous physics prior guidance. Please see Figure 5 A BreakTrajNet-PG network for detecting wave-breaking time-series propagation trajectories of thin, continuous linear targets in Timestack imagery is constructed. The detection network includes a LiteDexi lightweight backbone network and a two-layer physical information guidance mechanism. In the backbone network, let the first... The output feature map of the convolutional sub-layer is The calculation process is as follows:
[0078] in, For convolution kernel parameters, For bias terms, This represents the convolution operation. The ReLU activation function is used. Preferably, LiteDexi employs four serial encoder blocks to extract multi-scale spatiotemporal structural features step by step.
[0079] Perform the following steps on the feature maps output by each encoder block: Convolution mapping, unified number of channels:
[0080] Subsequently, bilinear interpolation upsampling is performed on the feature maps of each layer, and they are then concatenated along the channel dimension to obtain the fused features:
[0081] Then apply fusion features Convolution, outputting a single-channel trajectory probability map:
[0082] in, It is the Sigmoid activation function. This represents the probability distribution of the output trajectory.
[0083] In the physical enhancement section at the input end, the wave-breaking activity enhancement characterization constructed in step S2 is performed. As an additional input channel, it is stitched together with the original image at the input end:
[0084] This allows the network to simultaneously receive visual information and wave-breaking activity enhancement information during the first layer feature extraction. In the physical prior weighting part of the training process, a modulation loss function is first constructed based on the binary cross-entropy loss:
[0085]
[0086] in, This represents the trajectory probability predicted by the model. Indicates the true label, These are the modulation weighting coefficients. To focus on the index.
[0087] Furthermore, Tversky Loss is introduced as the structural constraint loss:
[0088] in, , and These represent the number of pixels representing true positives, false positives, and false negatives, respectively. For numerically stable terms, and These are the weight parameters.
[0089] Then, the two-dimensional dynamic wave-breaking activity region spatial prior constructed in step S2 is used... By introducing a loss function and weighting the pixel-level loss contributions at different spatial locations, the final physically guided hybrid loss function is obtained:
[0090] in, These are the prior weighting coefficients. and These are the balance coefficients for the pixel-level loss term and the structural constraint loss term, respectively. Preferably, , , , , , and The settings can be adjusted according to actual sample conditions and training requirements; in one embodiment of the present invention, the following can be selected: , , , , , , .
[0091] Furthermore, the model training process is as follows: Based on the Timestack image samples generated in step S1, the BreakTrajNet-PG model described in step S3 is trained. During training, manually labeled pixel coordinate sequences of the broken wave propagation trajectory are used as supervision signals. The labeling start point is defined as the position where the broken white foam is first clearly identifiable in the time series, and the labeling end point is defined as the position where the foam dissipates to the point where it can no longer be effectively distinguished from the background noise.
[0092] Preferably, the model training employs the AdamW optimizer, and the initial learning rate, batch size, maximum number of training epochs, and learning rate decay strategy can be set according to actual computing power and sample size. More preferably, in one embodiment of the present invention, the initial learning rate is set to... The batch size is set to 4, the maximum number of training epochs is set to 100, and the learning rate is scheduled using a cosine annealing strategy to smoothly decay the learning rate from its initial value to... It also features an early stop mechanism, which terminates training prematurely if the validation performance does not improve for 20 consecutive epochs. The model has approximately 2.3 M parameters and is capable of being deployed at the edge.
[0093] It is important to note that, regarding the organization of training data, in addition to using fully supervised, manually labeled trajectory samples for training, semi-supervised training, weakly supervised training, self-training, teacher-student distillation, transfer learning, incremental learning, domain-adaptive learning, or cross-site fine-tuning can also be used to train or optimize the model. Regarding the inference output method, besides directly outputting pixel-level trajectory probability maps, binary trajectory masks, centerlines, polyline sequences, spline curves, curve parametric equations, or other result formats that can characterize the continuous geometric shape of the wave propagation trajectory can also be output.
[0094] Please see Figure 6 In complex scenarios such as wavegroup separation, interference from non-breaking targets, and strong textured background interference, the BreakTrajNet-PG proposed in this invention exhibits significant advantages over pure visual baseline models (pure visual baseline models refer to models that detect wave propagation trajectories based solely on visual features such as pixel brightness, texture, and edges of the original Timestack image, without introducing enhanced wave activity representations, spatial priors of dynamic wave activity regions, and their corresponding guidance mechanisms). In complex wavegroup separation scenarios, while baseline models can identify the general shape of the broken trajectory, they are prone to trajectory breakage, whereas BreakTrajNet-PG... jNet-PG maintains a clearer and more continuous trajectory structure and achieves more stable separation of adjacent wave trains. In non-wave-breaking target interference scenarios, the baseline model is prone to local false responses to irrelevant small-scale signals on the beach side, while BreakTrajNet-PG can effectively suppress such non-wave-breaking isolated target interference, making the model response more concentrated in the real wave-breaking zone region. In strong textured background interference scenarios, facing a large area of oblique bright textures on the sea surface, the baseline model is prone to a large number of scattered false detections, while BreakTrajNet-PG can maintain continuous detection of real wave-breaking propagation trajectories while suppressing false responses. This shows that the intrinsic physics prior guidance mechanism, which combines input-side enhancement and training-side weighting, can effectively improve the model's ability to maintain trajectory continuity, determine breakage locations, and suppress background false detections in complex backgrounds, thereby enhancing the stability and cross-scene adaptability of wave-breaking propagation trajectory detection.
[0095] It is important to note that while LiteDexi is a preferred implementation for the backbone network structure, it is not the only one. Any other convolutional neural network, encoder-decoder network, feature pyramid network, dilated convolutional network, edge detection network, lightweight visual Transformer network, hybrid convolutional-transformer network, or other spatiotemporal structure detection network can be used as alternatives, as long as the detection of thin, continuous targets in Timestack images is achievable. This invention has compared different backbone networks such as UNet, FPN, ResDilated, and LiteDexi, verifying that the backbone form itself has alternative possibilities. The key is how to couple it with the physical prior guidance mechanism. Regarding the input-side prior guidance method, in addition to directly stitching the enhanced wave activity representation as an additional channel with the original image, methods such as intermediate layer feature fusion, attention gating, feature modulation, layer-by-layer fusion of prior and feature maps, cross-layer guidance, channel recalibration, or spatial mask enhancement can be used to embed the intrinsic physical prior into the model's feature extraction process. Regarding training-side constraints, in addition to the current prior weighted hybrid loss function, weighted cross-entropy, Focal Loss, Dice Loss, Tversky Loss, IoU Loss, boundary preservation loss, connectivity constraint loss, curve smoothing constraint loss, or combinations of multiple losses can also be used. Any method that can increase the model's focus on the wave-breaking activity region and enhance trajectory continuity during training is an equivalent alternative to this invention. The current embodiment has demonstrated that input-side enhancement and training-side weighting are two loosely coupled and complementary injection methods. Therefore, applying physical priors to the model's forward feature learning, backward optimization process, or both simultaneously, can all be included within the scope of this invention.
[0096] S4. Based on the continuous geometric morphology characterization results of the wave breaking trajectory, extract the vectorized propagation trajectory of a single wave breaking event, and automatically calculate the dynamic parameters of the wave breaking process based on the vectorized propagation trajectory.
[0097] In this embodiment, the dynamic parameters include breaking wave velocity, breaking position, propagation distance, duration, and breaking period; the breaking wave velocity is obtained by the local slope at the beginning of the trajectory, the breaking position is obtained by the cross-shore coordinates of the starting point of the trajectory, the propagation distance is obtained by the cross-shore displacement or arc length between the starting and ending points of the trajectory, the duration is obtained by the time difference between the ending and starting points of the trajectory, and the breaking period is obtained by the difference between the starting times of two adjacent wave breaking events.
[0098] For example, S4, automatic calculation and output of dynamic parameters based on propagation trajectory. In step S3, the Timestack image to be tested is input into the trained model, and the corresponding pixel-level wave breaking trajectory probability map is output. This step further processes the wave breaking trajectory probability map through thresholding, connected component analysis, and vectorization to extract the vectorized propagation trajectory; key dynamic parameters of a single wave breaking event are automatically calculated. It is important to note that in the post-processing stage, in addition to thresholding and connected component analysis, methods such as thinning algorithms, skeleton extraction, shortest path tracking, active contour models, curve fitting, graph search, or topology repair can also be used to complete trajectory vectorization and continuity optimization.
[0099] The parameters include at least the breaking wave velocity. Location of breakage Transmission distance and duration And can further extend the calculation of crushing cycle The parameter definitions are shown in Table 1.
[0100] Table 1 Key dynamic parameters and definitions of wave breaking process
[0101] in, This represents the initiation moment of a single wave breaking event. This is the time when the event ends. and This represents the start time of two consecutive wave breaking events. Through the calculation of these parameters, an automatic conversion is achieved from wave breaking propagation trajectory detection to the quantification of event-level dynamic processes.
[0102] Finally, the detection results and parameter information are output, stored, or uploaded to a remote management platform. The output results include, but are not limited to: trajectory probability maps, vectorized propagation trajectories, single-event parameter tables, and statistical analysis results, which are used for long-term continuous monitoring of wave breaking processes on natural beaches, identification of wave breaking activity areas, extraction of propagation features, and analysis of nearshore dynamic geomorphological processes.
[0103] It is important to note that, in terms of dynamic parameter extraction, the current preferred implementation method automatically calculates parameters such as breaking wave velocity, breaking location, propagation distance, duration, and breaking cycle based on the trajectory start and end points and local slope. This parameter system naturally matches the native spatiotemporal dimension of Timestack imagery and can be directly extracted from trajectory geometric features. In addition to the above parameters, the calculation can be further extended to include trajectory curvature, average propagation velocity, maximum propagation velocity, acceleration, breaking wave width, event interval, active area coverage, categorical statistical parameters, or other derived indicators based on trajectory geometry, depending on the application requirements. In terms of specific calculation methods, the propagation velocity can be obtained by local linear fitting at the initial stage of the trajectory, or by using sliding window fitting, piecewise fitting, spline curve differentiation, overall average slope, or other numerical differentiation methods; the propagation distance can be calculated using the cross-shore displacement between the start and end points, or by calculating along the trajectory arc length or piecewise cumulative distance; the duration can be directly calculated from the start and end times, or it can be defined by combining confidence thresholds, adaptive termination conditions, or multi-stage event recognition results. Therefore, parameter calculation is not limited to a single formula, but covers all equivalent schemes for dynamic quantization based on propagation trajectories.
[0104] Furthermore, in terms of output, the system can output trajectory probability maps, binary detection maps, vectorized trajectories, parameter tables, time series statistical results, early warning information, process reconstruction results, or graphical display results. The output format can be local files, database records, interface messages, visualization panels, or results from cloud management platforms. Regarding system integration, this invention can run offline as a standalone software method, be embedded in the edge computing module of a shore-based intelligent video monitoring platform, or be deployed on a remote control center server. It can be used for single-site automatic detection or multi-site collaborative analysis. The platform can adopt an overall architecture of front-end acquisition, edge computing, and remote management, indicating that the method of this invention can serve both scientific research scenarios and long-term operational monitoring scenarios. Further, the propagation trajectory detection results described in this invention can be used in conjunction with other video recognition results, such as fusion with wave breaking type recognition, beach geomorphological zoning recognition, nearshore hydrodynamic parameter inversion results, or spatiotemporal matching strategies, thereby achieving automatic reconstruction of event-level wave breaking processes and higher-level dynamic geomorphological analysis. This system can support event-level reconstruction and subsequent mechanism analysis through spatiotemporal matching strategies; therefore, joint analysis is also a scalable alternative to this invention.
[0105] Compared with the prior art, the present invention has the following advantages: (1) This invention can intrinsically extract the spatial prior of the wave-breaking activity area from the temporal statistical features of video, without relying on external environmental data. This invention does not rely on external water depth, tide level, topography or numerical model results to limit the wave-breaking area, but directly constructs the spatial prior of the wave-breaking activity area based on the temporal statistical features of the Timestack image itself. This method reduces the dependence on external multi-source data, reduces the complexity of data acquisition, maintenance and spatiotemporal matching, and has better portability and engineering application convenience.
[0106] (2) This invention realizes intelligent detection of wave-breaking time-series propagation trajectories for thin, continuous targets, and can more completely characterize the propagation process of a single wave-breaking event. Existing technologies mostly focus on overall wave parameter measurement or extraction of local features such as wave-breaking points, while this invention is aimed at thin, continuously extended wave-breaking propagation trajectories in Timestack images, and establishes a specialized intelligent detection method that can automatically and continuously characterize the propagation process of a single wave-breaking event from occurrence to dissipation, which is more suitable for the event-level analysis needs under natural beach conditions.
[0107] (3) This invention improves detection stability and cross-scene generalization ability in complex scenarios through a two-layer physical guidance mechanism of "input-end enhancement + training-end weighting". This invention applies endogenous physical priors to both the model input representation and the loss function optimization process, so that the model is constrained by the spatial laws of wave group activity areas while learning visual features. Compared with a pure visual baseline model, the method of this invention reduces the generalization performance decay (F1 Drop) by about 21% under the current cross-site testing conditions, and shows better robustness in scenarios such as complex wave group separation and strong sea surface reflection.
[0108] (4) This invention balances detection performance with lightweight design and has the potential for edge deployment. The invention adopts a lightweight detection network structure, which, while ensuring the accuracy and continuity of wave propagation trajectory detection, helps to reduce the scale of model parameters and computational overhead. It is suitable for integration with the edge computing module of the shore-based intelligent video monitoring system, thereby meeting the needs of on-site automated processing and long-term continuous operational use.
[0109] (5) This invention can extend from propagation trajectory detection to automatic calculation of dynamic parameters, realizing integrated analysis from "identification" to "quantification". This invention can not only detect the propagation trajectory of broken waves, but also automatically extract key dynamic parameters such as breakage location, propagation speed, propagation distance and duration based on the detected trajectory, thereby realizing the automatic conversion from time-series propagation trajectory identification to event-level dynamic process quantification, improving the application depth and analytical value of the method.
[0110] Example 2 Please see Figure 7A nearshore wave-breaking propagation trajectory detection device guided by video-based endogenous physics priors, based on the method described in one embodiment, the device comprising: The video acquisition module is used to acquire shore-based video images, set up a cross-shore sampling section within the video field of view, extract pixel brightness values frame by frame along the cross-shore sampling section and stack them in time order to generate a spatiotemporal stacked image. An endogenous physics prior generation module is used to perform temporal statistical feature analysis on the spatiotemporal stack image to generate endogenous physics priors. The endogenous physics priors include a wave-breaking activity enhancement characterization and a two-dimensional spatial prior. The wave-breaking activity enhancement characterization is obtained by weighting the normalized image with the trans-shore wave-breaking activity intensity. The two-dimensional spatial prior is copied from the one-dimensional spatial prior in the trans-shore direction along the time direction. The one-dimensional spatial prior is determined based on the trans-shore distribution characteristics of the wave-breaking activity intensity. The trajectory detection module has a built-in wave-breaking temporal propagation trajectory detection network, which is used to embed the wave-breaking activity enhancement representation into the feature extraction process of the model, and introduce the two-dimensional spatial prior as spatial weight into the training loss function of the detection network. The wave-breaking activity enhancement representation and the two-dimensional spatial prior are used to guide the detection network to perform pixel-level detection of the wave-breaking temporal propagation trajectory in the spatiotemporal stacked image, and output the continuous geometric shape representation result of the wave-breaking propagation trajectory. The parameter calculation module is used to extract the vectorized propagation trajectory of a single wave breaking event based on the continuous geometric morphology characterization results of the wave breaking trajectory, and automatically calculate the dynamic parameters of the wave breaking process based on the vectorized propagation trajectory.
[0111] In this embodiment, in order to better utilize the method described in one of the embodiments, this application proposes a nearshore wave-breaking propagation trajectory detection device guided by video endogenous physics priors. Each module corresponds to each step of the above method, and its specific principle has been described above and will not be repeated here.
[0112] Example 3 A video-driven, endogenous physics-guided nearshore wave propagation trajectory detection device includes: At least one processor and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described in one of the embodiments.
[0113] In this embodiment, to better run and process the method described in one of the embodiments, the above method is stored in a memory, and the stored method is executed using a processor. It should be noted that the principle and effect of each step have been described above and will not be elaborated further here.
[0114] Example 4 A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in one of the embodiments.
[0115] In this embodiment, to better operate and use the method described in one of the embodiments, the above method is stored in a computer-readable storage medium, and the above method is implemented using a processor. It should be noted that the principle and effect of each step have been described above and will not be elaborated further here.
[0116] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for detecting near-shore wave-breaking propagation trajectories guided by video-based endogenous physics priors, characterized in that, The method includes: Acquire shore-based video images, set up a cross-shore sampling section within the video field of view, extract pixel brightness values frame by frame along the cross-shore sampling section and stack them in time order to generate a spatiotemporal stacked image. Temporal statistical feature analysis is performed on the spatiotemporal stacked images to generate endogenous physical priors. The endogenous physical priors include a characterization of enhanced wave-breaking activity and a two-dimensional spatial prior. The characterization of enhanced wave-breaking activity is obtained by weighting the normalized images with the intensity of cross-shore wave-breaking activity. The two-dimensional spatial prior is copied from the one-dimensional spatial prior in the cross-shore direction along the time direction. The one-dimensional spatial prior is determined based on the cross-shore distribution characteristics of wave-breaking activity intensity. A wave-breaking temporal propagation trajectory detection network is constructed. The wave-breaking activity enhancement representation is embedded into the feature extraction process of the model, and the two-dimensional spatial prior is introduced as a spatial weight into the loss function of the training end of the detection network. The wave-breaking activity enhancement representation and the two-dimensional spatial prior are used to guide the detection network to perform pixel-level detection of the wave-breaking temporal propagation trajectory in the spatiotemporal stacked image, and output the continuous geometric shape representation result of the wave-breaking propagation trajectory. Based on the continuous geometric morphology characterization results of the wave breaking trajectory, the vectorized propagation trajectory of a single wave breaking event is extracted, and the dynamic parameters of the wave breaking process are automatically calculated based on the vectorized propagation trajectory.
2. The method for detecting near-shore wave-breaking propagation trajectories guided by video-based endogenous physics priors as described in claim 1, characterized in that, The generation of the enhanced characterization of the wave-breaking activity specifically includes: standardizing the spatiotemporal stacked image to obtain a standardized image; calculating the gray-level variance of each cross-shore location along the time dimension to obtain the intensity of the wave-breaking activity; and weighting the standardized image with the intensity of the wave-breaking activity to obtain the enhanced characterization of the wave-breaking activity.
3. The method for detecting near-shore wave-breaking propagation trajectories guided by video-based endogenous physics priors as described in claim 1, characterized in that... The generation of the two-dimensional spatial prior specifically includes: calculating the variance of the wave-breaking activity enhancement characterization along the time dimension to obtain the cross-shore wave-breaking activity intensity distribution; determining at least one wave-breaking activity area boundary based on the cross-shore wave-breaking activity intensity distribution; defining a one-dimensional spatial prior in the cross-shore direction for each wave-breaking activity area boundary; and copying the one-dimensional spatial prior along the time direction to obtain the corresponding two-dimensional spatial prior.
4. The method for detecting near-shore wave-breaking propagation trajectories guided by video-based endogenous physics priors as described in claim 1, characterized in that, The training loss function is a loss function that incorporates the two-dimensional spatial prior as spatial weights; the pixel-level loss function adopts any one or a combination of weighted cross-entropy, Focal Loss, Dice Loss, Tversky Loss, IoU Loss, boundary preservation loss, connectivity constraint loss, and curve smoothing constraint loss.
5. The method for detecting near-shore wave-breaking propagation trajectories guided by video-based endogenous physics priors as described in claim 1, characterized in that, The training loss function is a physically guided hybrid loss function. in, For two-dimensional space prior, These are the prior weighting coefficients. To modulate Focal Loss, For Tversky Loss, and This is the balance coefficient.
6. The method for detecting near-shore wave-breaking propagation trajectories guided by video-based endogenous physics priors as described in claim 1, characterized in that, The continuous geometric morphology representation of the wave propagation trajectory includes any of the following: pixel-level trajectory probability map, binary trajectory mask, center line, broken line sequence, spline curve, and curve parametric equation.
7. The method for detecting near-shore wave-breaking propagation trajectories guided by video-based endogenous physics priors as described in claim 1, characterized in that, The dynamic parameters include breaking wave velocity, breaking position, propagation distance, duration, and breaking period; the breaking wave velocity is obtained by the local slope at the beginning of the trajectory, the breaking position is obtained by the cross-shore coordinates of the starting point of the trajectory, the propagation distance is obtained by the cross-shore displacement or arc length between the starting and ending points of the trajectory, the duration is obtained by the time difference between the ending and starting points of the trajectory, and the breaking period is obtained by the difference between the starting times of two adjacent wave breaking events.
8. A near-shore wave-breaking propagation trajectory detection device guided by video-based endogenous physics priors, characterized in that, Based on the method according to any one of claims 1-7, the apparatus comprises: The video acquisition module is used to acquire shore-based video images, set up a cross-shore sampling section within the video field of view, extract pixel brightness values frame by frame along the cross-shore sampling section and stack them in time order to generate a spatiotemporal stacked image. An endogenous physics prior generation module is used to perform temporal statistical feature analysis on the spatiotemporal stack image to generate endogenous physics priors. The endogenous physics priors include a wave-breaking activity enhancement characterization and a two-dimensional spatial prior. The wave-breaking activity enhancement characterization is obtained by weighting the normalized image with the trans-shore wave-breaking activity intensity. The two-dimensional spatial prior is copied from the one-dimensional spatial prior in the trans-shore direction along the time direction. The one-dimensional spatial prior is determined based on the trans-shore distribution characteristics of the wave-breaking activity intensity. The trajectory detection module has a built-in wave-breaking temporal propagation trajectory detection network, which is used to embed the wave-breaking activity enhancement representation into the feature extraction process of the model, and introduce the two-dimensional spatial prior as spatial weight into the training loss function of the detection network. The wave-breaking activity enhancement representation and the two-dimensional spatial prior are used to guide the detection network to perform pixel-level detection of the wave-breaking temporal propagation trajectory in the spatiotemporal stacked image, and output the continuous geometric shape representation result of the wave-breaking propagation trajectory. The parameter calculation module is used to extract the vectorized propagation trajectory of a single wave breaking event based on the continuous geometric morphology characterization results of the wave breaking trajectory, and automatically calculate the dynamic parameters of the wave breaking process based on the vectorized propagation trajectory.
9. A near-shore wave-breaking propagation trajectory detection device guided by video-endogenous physics priors, characterized in that, include: At least one processor and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.