Extraction method of intertidal zone terrain profile based on time-average video image

By using a method for extracting intertidal topographic profiles based on time-averaged video images, and employing neural networks to identify breakpoints along the waterline and combining them with a local slope model for elevation calculation, this method solves the problem of high-frequency, long-term beach profile monitoring in traditional methods, and achieves high-precision and stable acquisition of beach topographic information.

CN121353984APending Publication Date: 2026-01-16THIRD INSTITUTE OF OCEANOGRAPHY STATE OCEANI C ADMINISTRATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511540922.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing technologies are insufficient for high-frequency, long-term beach profile monitoring, and traditional methods suffer from high costs, low accuracy, and complex data quality control.

Method used

By using a method for extracting intertidal topographic profiles based on time-averaged video imagery, a neural network model is used to identify the location of waterline breakpoints. An elevation calculation is performed by combining a local slope model. By integrating video monitoring and physical models and optimizing the data stream processing link, high-frequency and high-precision profile extraction is achieved.

Benefits of technology

It enables high-frequency, long-term beach profile monitoring, improves data continuity and accuracy, adapts to complex tidal environments, provides high-confidence beach topographic information, and supports refined coastal zone management and extreme event research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353984A_ABST
    Figure CN121353984A_ABST
Patent Text Reader

Abstract

The invention discloses an intertidal zone terrain profile extraction method based on time-average video images. The method comprises the following steps: acquiring an intertidal zone continuous video image, and correcting the intertidal zone continuous video image to a plane geographic coordinate system through a coordinate conversion model; establishing single-valued mapping between ground coordinates and image plane coordinates, and back-projecting pixels of the video image to a unified geographic grid; constructing an affine matrix of a single-pixel image column based on single-valued mapping, and embedding a geographic direction vector of a target section into spatial transformation of the image column; intercepting an image column sample of the image at each moment in the video image according to a preset pixel width along the target section direction; identifying the position of a waterline breakpoint in the image column by using a neural network model; and carrying out waterline elevation calculation and the like by referring to a local slope model based on a balanced section. According to the method, the high-frequency characteristics of video monitoring and the priori knowledge of the physical model are fused, so that the efficiency and precision of profile extraction are effectively improved while the data continuity is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of shore-based video monitoring and the technical field of photogrammetry, and in particular to a method for extracting an intertidal zone topographic profile based on time-averaged video images, which extracts a characteristic topography from a target profile single-pixel image column directly intercepted from a time-averaged image. BACKGROUND

[0002] A beach is a geomorphologic unit commonly developed on sandy coasts, and is the forefront for resisting the action of marine power and defending the rear highlands. As one of the main topographic features of a beach, a beach profile not only directly reflects the morphological structure thereof, but also is a key indicator reflecting the evolution of a nearshore environment. In recent years, an increase in extreme storm events and refined management have put forward higher requirements on the spatiotemporal resolution of monitoring: not only should the rapid evolution of a topography be continuously tracked on an event scale, but also long-term morphological changes should be stably depicted on a seasonal or even interannual scale. Therefore, developing a high-frequency and continuous beach profile observation technology and algorithm has become a key scientific problem to be solved in the current field of coastal mapping.

[0003] Traditional profile monitoring means such as RTK-GPS, airborne or ground-based LiDAR, and unmanned aerial vehicle aerial photography, although have high precision, are difficult to form a high-frequency and long-time series of profiles due to the cost, operation window, and coverage frequency. Remote sensing images have an advantage in macroscopic grasp, but lack the ability to identify on a microscale. In comparison, a shore-based video monitoring system has the advantages of low deployment cost, high spatiotemporal resolution, and long-term operation, and provides a new data source for high-frequency beach profile monitoring.

[0004] The existing path for extracting a profile from a video image is mostly to first reconstruct an intertidal zone digital elevation model based on a waterline method, and then cut the target profile. Meanwhile, a spectrum method based on wave inversion and a spatiotemporal filtering fusion algorithm can be used to inversely calculate the intertidal zone topography from the color dispersion relationship of video time series. However, the link from the overall topography to the profile is long, is often affected by joint fusion, interpolation smoothing, and error propagation, and has complex quality control and high update cost. SUMMARY

[0005] In view of this, the purpose of the present application is to provide a method for extracting an intertidal zone topographic profile based on time-averaged video images, which can at least solve one of the technical problems mentioned in the background.

[0006] According to one aspect of the present application, a method for extracting an intertidal zone topographic profile based on time-averaged video images is provided, and the method comprises:

[0007] obtaining continuous video images of an intertidal zone, and correcting them to a plane geographic coordinate system through a coordinate conversion model;

[0008] A single-value mapping between ground coordinates and image plane coordinates is established, and the pixels of the video image are back-projected onto a unified geographic grid. Based on the single-value mapping, an affine matrix of single-pixel image columns is constructed, and the geographic direction vector of the target profile is embedded in the spatial transformation of the image columns. Image column samples of the video image at each time point are extracted along the target profile direction according to a preset pixel width. The location of the waterline breakpoint in the image columns is identified using a neural network model.

[0009] Referring to the local slope model based on the equilibrium profile, the elevation of the waterline is calculated and the calculation results are assigned to the corresponding waterline breakpoints.

[0010] A topographic map of the beach profile was drawn based on the distance from the shore and elevation data of the breakpoints at each time point.

[0011] In the above technical solution, this method effectively improves the efficiency and accuracy of profile extraction while ensuring data continuity by integrating the high-frequency characteristics of video monitoring with the prior knowledge of physical models.

[0012] Specifically, this scheme establishes a single-valued mapping between ground coordinates and image plane coordinates, embedding the geographic direction vector of the target profile into the spatial transformation of the image pillars. This optimizes the path from whole-domain terrain reconstruction to direct extraction of profile directions, avoiding the error propagation problems introduced by seam fusion and interpolation smoothing in traditional waterline methods. By using a neural network model to identify waterline breakpoints and combining this with a priori beach profile model to obtain local slope in real time, the scheme not only enhances the robustness of waterline detection in complex wave environments but also reduces the uncertainty of relying solely on image inversion through physical constraints. Furthermore, based on the elevation assignment mechanism of multi-time-lapse breakpoint locations, the scheme incorporates the dynamic influence of tide levels and wave elements, enabling the profile data to capture rapid topographic responses at the event scale while maintaining consistency in long-term monitoring, thus effectively balancing the dual requirements of high-frequency acquisition and data stability.

[0013] In summary, this technical solution, by optimizing the data stream processing chain, integrating machine learning and physical models, and enhancing adaptability to dynamic environments, forms an efficient, stable, and sustainable method for beach profile monitoring, providing strong technical support for refined coastal zone management and extreme event research.

[0014] In some embodiments, if continuous high-frequency video images from different perspectives are acquired through several camera positions, then after back-projecting the pixels onto a unified geographic grid, the method further includes:

[0015] By projecting images acquired from several camera positions onto the sea surface height that changes over time using time-varying tide levels, and then using a weighted stitching strategy based on distance transformation to complete the image fusion of the several camera positions.

[0016] In the aforementioned technical solution, the extraction of profiles from a single perspective is transformed into a collaborative and robust regional monitoring system. Specifically, the fusion of multi-camera data significantly expands the spatiotemporal coverage of monitoring, effectively overcoming blind spots and occlusion issues at individual cameras, thus ensuring the continuity and integrity of data acquisition of the target profile in both the horizontal and vertical directions. Its core advantage lies in the physical rationality and algorithmic sophistication of the fusion strategy: by projecting images from each camera position onto dynamically changing sea levels using time-varying tides, this process adheres to the physical reality of the marine dynamic environment, providing a precise registration benchmark for images from different perspectives in the vertical direction, fundamentally avoiding stitching misalignments that may result from projections based on static planes. Furthermore, a weighted stitching strategy based on distance transformation achieves pixel-level fusion in overlapping areas. It can smoothly balance the observation quality and reliability of different cameras at specific spatial locations, for example, prioritizing image information closer to the orthophoto perspective, thereby generating a seamless, consistent, and geometrically accurate time-averaged image.

[0017] In summary, this solution not only overcomes the inherent limitations of single-station monitoring, but also enhances the geometric consistency and physical authenticity of data sources, thus constructing a high-frequency beach monitoring solution that can adapt to complex tidal environments and provide high-confidence topographic information.

[0018] In some embodiments, based on single-value mapping, an affine matrix of single-pixel image pillars is constructed, and the geographic direction vector of the target profile is embedded in the spatial transformation of the image pillars; along the target profile direction, image pillar samples of the video image at each time point are extracted according to a preset pixel width, including:

[0019] Based on single-valued mapping, an affine matrix for single-pixel image pillars is constructed, allowing the image pillars to advance along a given target profile direction; let the profile geographic increment vector be:

[0020]

[0021] In the formula, (x1,y1) and (x2,y2) represent the start and end points of the target profile image coordinates, respectively; M represents the pixel row index of the image column;

[0022] The affine matrix of a single-pixel image column is:

[0023]

[0024] The affine matrix embeds the geographic direction vector of the target profile into the spatial transformation of the image bars. Along the target profile direction, image bar samples are extracted from the video image at each time point according to a preset pixel width. Then, the world coordinates of the pixel in the i-th row are (x1+i*dX, y1+i*dY), and its azimuth angle θ is calculated using the following formula:

[0025]

[0026] In the above technical solution, the profile extraction process is transformed from a raster-based empirical operation to a vector-based deterministic computation.

[0027] Specifically, an affine matrix is ​​used to achieve a linear transformation between geographic space and image space. By embedding the geographic orientation vector and starting coordinates of the target profile into the affine matrix, this design constructs a mapping channel from image column index to world coordinates. This method solves the lengthy link of traditional methods that first reconstruct the entire digital elevation model and then extract the profile through resampling, avoiding information loss and error introduction caused by interpolation smoothing and resampling processes, and achieving "direct geocoding" of profile data. Furthermore, the scheme provides a formula for calculating the azimuth angle, considering all possible profile orientations to ensure accurate azimuth information can be obtained in any geographic direction.

[0028] In some embodiments, using a neural network model to identify the location of waterline breakpoints in an image column includes:

[0029] A pre-trained stacked Bi-LSTM network model processes single-pixel image pillars column-by-column; this network model includes:

[0030] The input subnetwork is used to input the width and column height of the image bars as time steps and the pixel channel X as features into the stacked Bi-LSTM network for time-by-time encoding.

[0031] The feature extraction subnetwork includes two bidirectional Bi-LSTM layers, each of which contains a forward LSTM layer and a backward LSTM layer. This feature extraction subnetwork is used to extract temporal features from the forward and backward directions of the sequence, respectively, to capture contextual information.

[0032] The classification subnetwork is used to perform binary classification of each pixel of the image column as "land" or "ocean" based on the features extracted by the feature extraction subnetwork;

[0033] The output and feedback subnetworks are used to output the final classification results, identify the positions of different classifications, and determine the last "land" pixel in the image column that is closer to the "ocean" side as the breakpoint of the waterline.

[0034] In the above technical solution, a stacked bidirectional long short-term memory network (Bi-LSTM) is used to process spatiotemporal image columns. This design transforms the problem of waterline recognition into a segmentation task of temporal pixel sequences. Compared with traditional image processing methods, it achieves a significant improvement in recognition accuracy and robustness in complex hydrodynamic environments.

[0035] Specifically, its core advantage lies in the neural network architecture's ability to capture and integrate temporal and spatial contextual information from videos. By using the width and column height of a single-pixel image column as the time step and the pixel channel as the feature input, the model can learn the essential differences between "land" and "ocean" from the temporal spectral changes at each geographically fixed point, rather than relying solely on the color or texture features of a single frame. This effectively overcomes interference from lighting changes, reflections from wet sand, and edge blurring caused by thin water cover. The key lies in the two-layer bidirectional Bi-LSTM structure in the feature extraction sub-network: the forward LSTM layer captures the dependencies from the past to the present, while the backward LSTM layer captures the context from the future to the present. This bidirectional mechanism allows the model to simultaneously refer to the "historical" and "future" states of any pixel in the entire time series when determining its category, thus providing a strong ability to suppress dynamic noise such as wave ebb and water turbidity.

[0036] In summary, this stacked Bi-LSTM model, with its powerful temporal modeling and bidirectional context-aware capabilities, provides a high-precision and high-stability solution for waterline breakpoint detection. It effectively transforms the high-frequency advantages of video surveillance into reliable ground feature classification information, laying a solid data foundation for subsequent elevation calculations and profile mapping.

[0037] In some embodiments, the bidirectional Bi-LSTM layer includes:

[0038] The first bidirectional Bi-LSTM layer encodes the local visual modal information of the input sequence through a gating mechanism, extracts the spatiotemporal distribution features of fine-grained chromaticity gradient and texture primitives, and forms a discriminative representation of the underlying visual features.

[0039] The secondary bidirectional Bi-LSTM layer takes the output of the primary bidirectional Bi-LSTM layer as input, models longer-range contextual dependencies through hidden state interactions across time steps, incorporates the prior structural constraint of "land-ocean single jump" in the feature abstraction process, and achieves the integration of cross-regional semantic associations through hidden layer dynamic evolution, ultimately generating feature vectors that have both high-order semantic discriminativeness and structural consistency.

[0040] In the above technical solution, the dual-layer bidirectional Bi-LSTM architecture achieves accurate and robust mapping from raw pixels to high-order semantics through a hierarchical and progressive information processing mechanism with clear division of labor, which greatly improves the cognitive intelligence level of waterline recognition.

[0041] Specifically, its advantages lie in the efficient collaboration and functional specialization of the two layers. The first-layer bidirectional Bi-LSTM acts as a "local feature perceptron," its core value lying in using gating mechanisms (such as input gate, forget gate, and output gate) to finely encode the input sequence, selectively memorizing and forgetting information, thereby effectively capturing and integrating the spatiotemporal distribution of fine-grained visual features such as chromaticity gradients and texture primitives from complex visual flows. This process lays a solid perceptual foundation for the entire recognition task, ensuring the model's high sensitivity to pixel-level appearance changes (such as wet sand and water reflection). Building on this, the second-layer bidirectional Bi-LSTM plays the role of a "global semantic integrator." It takes the information-rich feature sequence output from the previous layer as input and models longer-range contextual dependencies through hidden state interactions across time steps. Its key innovation lies in incorporating the crucial prior physical knowledge of "single land-to-ocean transition" as a structural constraint into the feature abstraction process. This enables the network not only to identify local features, but also to understand the semantic roles and sequential logic of these features in the overall sequence. Through the dynamic evolution of its internal hidden layers, it can effectively distinguish between real, structurally consistent land-water boundaries and local disturbances caused by instantaneous waves, sprays, or shadows, thereby integrating cross-regional semantic associations and outputting more discriminative and overall consistent feature vectors.

[0042] In summary, this two-tiered design, from local to global and from perception to cognition, ensures that the determination of the final waterline break point is based on visual evidence and conforms to the constraints of the physical world, thus achieving high-precision recognition in complex and ever-changing natural environments.

[0043] In some embodiments, the establishment of the local slope model based on the equilibrium profile includes:

[0044] Collect several RTK-GPS measured beach profile data and fit a profile model based on the beach profile data;

[0045] Based on the profile fitting model, a beach alluvial plain slope model is established. Using this beach alluvial plain slope model and the breakpoint locations at each time point, the local slope corresponding to the breakpoint location is obtained in real time.

[0046] In the above technical solution, based on the breakpoint locations obtained at each moment, the corresponding local slope is obtained in real time through the prior beach profile model, which reflects the strategic advantage of deep integration of physical model and observation data.

[0047] In summary, by introducing a parameterized prior beach profile model, this approach combines discrete, sparse shoreline observation information with the overall morphological regularity of the beach topography. This solves the problem of missing slope information encountered when relying solely on shoreline elevation assignment, significantly enhancing the physical rationality and spatial continuity of the topography reconstruction. Its core advantage lies in using a model with clear geophysical significance to characterize the typical morphology of the beach profile. The model's strength lies in its concise form, parameterization, and clear physical meaning, which can be robustly determined using several collected RTK-GPS measured profile data, thus solidifying the specific topographic characteristics of the beach within the model. Based on this profile model, an analytical and continuous beach alluvial slope model is obtained. This design enables the system to calculate the corresponding local slope in real-time and efficiently using this analytical model based on the distance from the shoreline discontinuity identified at any given time, without relying on additional time-consuming field measurements or complex numerical fitting processes. This not only greatly improves the automation and computational efficiency of the entire profile extraction process, but more importantly, it assigns a spatially self-consistent slope value to each independent waterline point that conforms to the evolution of the overall profile shape. This effectively constrains the solution space for waterline elevation conversion and suppresses random errors caused by factors such as wave fluctuations and instantaneous water surface fluctuations, thereby improving the accuracy and reliability of the final generated beach profile topographic map.

[0048] In summary, this method for real-time acquisition of local slope based on a priori profile model transforms discrete waterline location information into continuous terrain features in a computationally efficient manner by embedding physical terrain constraints into the data processing flow, ensuring that it conforms to both instantaneous observation and macroscopic geomorphological patterns.

[0049] In some embodiments, the waterline elevation is calculated using the following formula:

[0050]

[0051] In the formula, Z sl Z represents the elevation of the waterline. o K represents the nearshore water level elevation, a combination of astronomical tides and meteorological data. osc Representing the empirical coefficient of impulsive flow, β s Let H0 be the average slope of the sluice zone, H0 be the significant wave height in nearshore waters, and L0 be the wavelength in deep offshore waters. L0 = gT 2 / 2π, where T is the effective wave period and g is the gravitational acceleration.

[0052] In the above technical solution, the formula for calculating the waterline elevation achieves dynamic and accurate assignment of instantaneous waterline elevation by precisely coupling tidal level, wave dynamics and beach morphology parameters.

[0053] Specifically, the significant advantage of this formula lies in its high degree of physical completeness and adaptability to dynamic environments. It surpasses traditional methods that merely equate the waterline elevation with the static tidal level (Z). o The simplified assumptions explicitly introduce wave elements (significant wave height H0, significant wave period T) and local slope (β). s () as the core variable. Complex structure terms in the formula.

[0054] Essentially, this is a quantitative modeling of the physical process of wave run-up, based on wave theory and hydrodynamics of the run-up zone. This quantitative relationship allows the model to accurately characterize the vertical height of wave run-up towards the land on a beach with a specific slope, thus correcting the actual position of the waterline to a combined result of "tide level + wave run-up". This mechanism enables the technology to sensitively respond to instantaneous fluctuations in waterline elevation caused by different wave conditions (such as large waves during storm surges and small waves on ordinary days), thereby capturing the rapid response of the terrain under extreme events or daily fluctuations. Furthermore, the empirical coefficient of run-up (K) introduced in the formula... osc It provides an interface for fine-tuning and calibration of the model based on specific site data, further enhancing its applicability and calculation accuracy in different geographical environments.

[0055] In summary, this formula for calculating the waterline elevation is a dynamic corrector deeply embedded in physical mechanisms. It ensures that the final generated beach profile not only reflects the slow changes of tides but also records the decisive role of wave dynamics in shaping the instantaneous beach topography in greater detail. This allows the topographic profile based on video image inversion to more realistically and dynamically reproduce the actual state of the intertidal zone topography, meeting the stringent requirements for data quality in high-frequency and high-precision monitoring.

[0056] According to another aspect of the present invention, an apparatus for extracting intertidal topographic profiles based on time-averaged video imagery is provided, wherein the system, based on the above-described method, comprises:

[0057] The acquisition module is used to acquire continuous video images of the intertidal zone and correct them to a planar geographic coordinate system through a coordinate transformation model.

[0058] The recognition module is used to establish a single-value mapping between ground coordinates and image plane coordinates, and back-project the pixels of the video image onto a unified geographic grid. Based on the single-value mapping, it constructs an affine matrix of single-pixel image columns, and embeds the geographic direction vector of the target profile into the spatial transformation of the image columns. Along the direction of the target profile, it extracts image column samples of the video image at each time point according to a preset pixel width. It uses a neural network model to identify the position of the waterline breakpoint in the image columns.

[0059] The slope calculation module is used to calculate the elevation of the waterline by referring to the local slope model based on the equilibrium profile and assign the calculation results to the corresponding waterline breakpoints.

[0060] The elevation assignment module is used for

[0061] The terrain module is used to draw topographic maps of the beach profile based on the distance from the shore and elevation data of the breakpoints at each time point.

[0062] In order to better utilize the above method, this application proposes an extraction device for intertidal topographic profile based on time-averaged video images. Each module corresponds to a step of the above method, and its specific principle has been described above and will not be repeated here.

[0063] According to another aspect of the present invention, an apparatus for extracting intertidal topographic profiles based on time-averaged video imagery is provided, comprising:

[0064] At least one processor and a memory communicatively connected to said at least one processor;

[0065] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described above.

[0066] In the above technical solution, to better operate and process the method, the method is stored in memory, and the processor executes the stored method. It should be noted that the principle and effect of each step have been described above and will not be elaborated upon here.

[0067] According to another aspect of the present invention, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the above-described method.

[0068] In the above technical solution, to better operate and use the method, the method is stored in a computer-readable storage medium and implemented using a processor. It should be noted that the principle and effect of each step have been described above and will not be elaborated upon here. Attached Figure Description

[0069] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0070] Figure 1This is a flowchart illustrating an embodiment of a method for extracting intertidal topographic profiles based on time-averaged video images according to the present invention.

[0071] Figure 2 This is a schematic diagram of the video image to planar geographic coordinate image transformation according to an embodiment of the method for extracting intertidal topographic profiles based on time-averaged video images of the present invention.

[0072] Figure 3 This is a schematic diagram of target profile image bar extraction according to an embodiment of the method for extracting intertidal topographic profiles based on time-averaged video images of the present invention.

[0073] Figure 4 This is a schematic diagram of the neural network structure for detecting waterline breakpoints, an embodiment of a method for extracting intertidal topographic profiles based on time-averaged video images according to the present invention.

[0074] Figure 5 This is a schematic diagram of a slope model of an embodiment of a method for extracting intertidal topographic profiles based on time-averaged video images according to the present invention.

[0075] Figure 6 This is a schematic diagram of the profile generation process of an embodiment of a method for extracting intertidal topographic profiles based on time-averaged video images according to the present invention.

[0076] Figure 7 This is a schematic diagram of an embodiment of a device for extracting intertidal topographic profiles based on time-averaged video images according to the present invention. Detailed Implementation

[0077] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0078] This embodiment utilizes time-averaged images generated from video footage to perform refined extraction of intertidal topographic profiles under complex backgrounds. A stacked bidirectional long short-term memory network model (Stacked Bi-LSTM) with strong noise robustness is introduced to identify the location of waterline breakpoints in single-pixel images, and a local profile slope model is established to calculate wave run-up, thereby achieving refined extraction of the target profile. This effectively compensates for the shortcomings of traditional monitoring and analysis methods and provides strong technical support for nearshore beach dynamic geomorphology research and refined management.

[0079] Example 1

[0080] Please see Figure 1A method for extracting intertidal topographic profiles based on time-averaged video imagery, the method comprising:

[0081] S1. Acquire continuous video images of the intertidal zone and correct them to a planar geographic coordinate system through a coordinate transformation model; establish a single-value mapping between ground coordinates and image plane coordinates, and back-project the pixels of the video images onto a unified geographic grid.

[0082] In this embodiment, a nearshore area was selected for investigation, and a shore-based video monitoring system was constructed. Continuous time-series video image data was acquired, with a time resolution of 2Hz or higher. The time-series video images were then converted and corrected to a planar geographic coordinate system to form a time-series planar geographic coordinate image. Please refer to [link to relevant documentation]. Figure 2 The coordinate transformation model is as follows:

[0083]

[0084] In the formula, Let be a point in a planar coordinate system, and let its image distortion coordinates be (c, r); (u, v) represent the undistorted coordinates to be solved, and d 2 =u 2 +v 2 ; Indicates the distortion coefficient; s represents the pixel size; (o c ,o r () represents the coordinates of the principal point; This represents the position of the camera in a planar world coordinate system. It is a rotation matrix defined by Euler angles (φ,σ,τ).

[0085] In this embodiment, if continuous high-frequency video images from different perspectives are acquired by several camera positions, after back-projecting the pixels onto a unified geographic grid, the method further includes: projecting the images acquired by several camera positions onto the sea surface height that changes over time using time-varying tide levels, and completing the image fusion of several camera positions based on a weighted stitching strategy of distance transformation.

[0086] For example, with four camera positions, please refer to [link / reference]. Figure 3 First, the camera's intrinsic and extrinsic parameters are input into the coordinate transformation model to establish a single-value mapping between ground coordinates (X,Y,Z) and image plane coordinates (u,v). Then, the pixels are back-projected onto a unified geographic grid, and a weighted stitching strategy based on distance transformation is used to complete the fusion of images from four camera positions. It is worth noting that since each photo was taken at a different time, using a single plane would lead to distortion of the coastline. Therefore, this application uses time-varying tide levels to project all four camera positions onto the sea level that changes over time before stitching them into an orthophoto, thereby reducing near-shore reprojection errors.

[0087] S2. Based on single-value mapping, construct the affine matrix of single-pixel image pillars and embed the geographic direction vector of the target profile into the spatial transformation of the image pillars; along the target profile direction, extract image pillar samples of the video image at each time point according to the preset pixel width; use a neural network model to identify the position of the waterline breakpoint in the image pillars.

[0088] In this embodiment, please refer to Figure 3 To ensure that image bars inherit the coordinate reference and true orientation of the source image in geospatial space, a 4×4 ModelTransformationTag was constructed, allowing the pixel row index (from 1 to M) of the image bars to advance along the given target profile direction. Let the profile geographic increment vector be:

[0089]

[0090] In the formula, (x1,y1), (x s y1 and y2 represent the start and end points of the target profile image coordinates, respectively; M represents the pixel row index of the image column;

[0091] The affine matrix m of a single-pixel image column is constructed as follows:

[0092]

[0093] The affine matrix embeds the profile direction vector into the spatial transformation of the image cylinder. The "row" axis of the image coordinates advances along the target profile direction. Therefore, the world coordinates of the i-th row pixel are (x1+i*dX, y1+i*dY), and its azimuth angle θ is calculated using the following formula:

[0094]

[0095] In this embodiment, the identification of the waterline breakpoint location in the image column using a neural network model includes:

[0096] A pre-trained stacked Bi-LSTM network model processes single-pixel image columns sequentially.

[0097] Please see Figure 4 The network model includes:

[0098] The input subnetwork is used to input the width and column height of the image bars as time steps and the pixel channel X as features into the stacked Bi-LSTM network for time-by-time encoding.

[0099] The feature extraction subnetwork includes two bidirectional Bi-LSTM layers, each of which contains a forward LSTM layer and a backward LSTM layer. This feature extraction subnetwork is used to extract temporal features from the forward and backward directions of the sequence, respectively, to capture contextual information.

[0100] The classification subnetwork is used to perform binary classification of each pixel of the image column as "land" or "ocean" based on the features extracted by the feature extraction subnetwork;

[0101] The output and feedback subnetworks are used to output the final classification results, identify the positions of different classifications, and determine the last land pixel on the ocean side of the image column as the breakpoint of the waterline.

[0102] In this embodiment, traditional RNNs are prone to gradient vanishing or exploding problems when processing long sequence data, making it difficult to learn long-distance dependencies. To address this challenge, Hochreiter and Schmidhuber proposed the LSTM network model. The internal structure of LSTM adds gates to the classic RNN to selectively add and remove past temporal information; that is, it adds input gates, output gates, and forget gates to control the input and output of the current unit and the addition or subtraction of information from the previous unit, respectively. To overcome the limitation of unidirectional LSTMs, which can only utilize past context and not future context, Schuster and Paliwal proposed the Bidirectional Recurrent Neural Network (BRNN), which merges two hidden LSTM layers with opposite time directions into a single output. This allows the output layer to utilize relevant information from both the past and future. In this embodiment, one LSTM layer processes data in the forward direction of the input sequence (land → sea), and the other processes data in the reverse direction (sea → land). Bidirectional gating memory simultaneously gathers context from both the land and sea sides, improving the overall ability to distinguish between false boundaries and real boundaries such as white waves and seepage layers. The first layer of Bi-LSTM focuses on local chromaticity and fine-grained texture, while the second layer integrates longer context to form a structural prior for "single land-to-sea transition" and stronger discriminative features. Each layer is normalized to stabilize training and suppress overfitting. Finally, row-by-row features are fed into fully connected layers and Softmax layers to perform binary classification ("land" or "water") on each pixel in each column of the image, and the last "land" pixel is determined as the breakpoint of the waterline.

[0103] The bidirectional Bi-LSTM layer includes:

[0104] The first bidirectional Bi-LSTM layer encodes the local visual modal information of the input sequence through a gating mechanism, extracts the spatiotemporal distribution features of fine-grained chromaticity gradient and texture primitives, and forms a discriminative representation of the underlying visual features.

[0105] The secondary bidirectional Bi-LSTM layer takes the output of the primary bidirectional Bi-LSTM layer as input, models longer-range contextual dependencies through hidden state interactions across time steps, incorporates the prior structural constraint of "land-ocean single jump" in the feature abstraction process, and achieves the integration of cross-regional semantic associations through hidden layer dynamic evolution, ultimately generating feature vectors that have both high-order semantic discriminativeness and structural consistency.

[0106] In this embodiment, a stacked Bi-LSTM is defined to process single-pixel image pillars column by column sequentially, thereby realizing an artificial neural network for waterline breakpoint extraction. The network starts with a sequential input layer, with an image pillar of 1 pixel width as the input. The column height H of the image pillar is used as the time step, and the pixel channel X is used as the feature to be encoded by the stacked Bi-LSTM at each time step. Bidirectional gating memory simultaneously gathers the context of both the land side and the sea side, thereby improving the overall discrimination power between false boundaries and real boundaries such as white waves and seepage layers. The first layer of Bi-LSTM focuses on local color and fine-grained texture, while the second layer integrates a longer context to form a structural prior of "single land → sea transition" and stronger discriminative features. Each layer is normalized to stabilize training and suppress overfitting. Finally, the row-by-row features are fed into a fully connected layer and a Softmax layer to perform binary classification ("land" or "water") on each pixel of each column of the image, and the last "land" pixel is determined as the waterline breakpoint position.

[0107] In this embodiment, bidirectional gated memory simultaneously aggregates context from both the landside and seaside, improving the overall ability to distinguish between false boundaries and true boundaries such as white waves and seepage layers. The first-layer Bi-LSTM focuses on local chromaticity and fine-grained texture, while the second layer integrates a longer context, forming a structural prior of "single land-to-sea transition" and stronger discriminative features. The main purpose of setting two layers is that intertidal video images are subject to various interferences in the judgment of waterline breakpoints due to surface seepage, white wave breaking, etc. Therefore, this application can further improve classification or regression performance by stacking Bi-LSTM layers in the neural network. In addition, deep hierarchical models are more efficient than shallow models when representing certain functions, so a two-layer bidirectional Bi-LSTM layer was designed. In actual operation, it was found that this design is suitable for waterline breakpoint judgment and can significantly improve the judgment accuracy.

[0108] S3. Referencing the local slope model based on the equilibrium profile, calculate the elevation of the waterline and assign the calculation results to the corresponding waterline breakpoints.

[0109] In this embodiment, several equilibrium profile models have been proposed by the academic community, such as the Brunn-Dean model and the exponential equilibrium profile. The specific model used depends on the physical parameters calculated on-site. Those skilled in the art can select a suitable profile model based on the actual shoreline conditions. In this embodiment, Xisha Bay is used as the calculation scenario. Since the Xisha Bay shoreline approximates the line shape of the exponential model, the exponential model is used as an example for illustration. The establishment of the local slope model based on the equilibrium profile includes:

[0110] Collect several RTK-GPS measured beach profile data; establish a fitted profile model:

[0111] h(x)=B(1-e -kx )

[0112] The profile model parameters B and k are fitted based on the beach profile data;

[0113] A beach alluvial slope model was established based on this profile fitting model.

[0114] h′(x)=tanβ f (x)=Bke -kx

[0115] In the formula, x is the horizontal distance from the coastline, h is the corresponding water depth, B is the asymptotic water depth scale, k is the lateral transition rate constant, and β f Let be the slope at any point offshore.

[0116] Using the beach alluvial slope model and the breakpoint locations at various times, the local slope corresponding to the breakpoint locations can be obtained in real time.

[0117] In this embodiment, please refer to Figure 5 Using 10 RTK-GPS measured beach profile data, a fitted profile model was established: h(x)=B(1-e -kx ),like Figure 5 As shown in Figure a: the fitting result range of parameter B in the fitted profile model is 4.6824–4.9954, with an expected value of 4.8491; the fitting result range of parameter k is 0.01665–0.01895, with an expected value of 0.01739. The model fitting evaluation parameters were calculated, and the coefficient of determination (R-square) of the fitted profile model is 0.931, indicating a high goodness of fit. Based on this profile fitting model, a slope model for the alluvial plain of Xisha Bay is established, as follows... Figure 5 As shown in b, the slope β ranges from 1.09° to 10.78°, with an average slope of 2.18°. The corresponding average tanβ value is 0.038, which is close to the calculated average slope of 0.04 from the measured beach foreshore. Therefore, the model has high reliability.

[0118] In this embodiment, the waterline elevation is calculated using the following formula:

[0119]

[0120] In the formula, Z sl Z represents the elevation of the waterline. o K represents the nearshore water level elevation, a combination of astronomical tides and meteorological data. osc Representing the empirical coefficient of impulsive flow, β s Let H0 be the average slope of the sluice zone, H0 be the significant wave height in nearshore waters, and L0 be the wavelength in deep offshore waters. L0 = gT 2 / 2π, where T is the effective wave period and g is the gravitational acceleration.

[0121] The standardized method for calculating the elevation of the waterline is expressed as follows:

[0122]

[0123] In the formula, Z sl Z represents the elevation of the waterline. o η represents the nearshore water level elevation, which is a combination of astronomical tides and meteorological data. sl This indicates the elevation change caused by wave-induced water level rise. K represents the maximum elevation reached by the wave-driven motion. osc This is the empirical coefficient for the impulsive flow, which needs to be calibrated through on-site experiments.

[0124] The empirical relation proposed by Stockdon et al. (2006) is directly η. sl η osc Compared with the commonly used extreme climb index R 2% A parameterized expression is provided, allowing the two to correspond one-to-one and integrate equation (4) into:

[0125]

[0126] In the formula, β s Let H0 be the average slope of the embankment, H0 be the significant wave height in nearshore waters, and L0 be the wavelength in deep offshore waters. This can be determined through the dispersion relation: L0 = gT 2 The value is calculated as / 2π (T is the effective wave period, and g is the gravitational acceleration).

[0127] Existing studies often use the average slope to uniformly calculate wave run-up (Atkinson et al., 2017). While this method is simple to operate, it neglects the spatiotemporal drift of the tidal zone caused by tidal fluctuations, leading to significant variations in local beach slopes at different tide levels. This, in turn, introduces a systematic bias into the estimation accuracy of extreme run-up values. To improve the accuracy of wave run-up estimation, this paper abandons the use of the average slope of the tidal zone. Given the sensitivity of video analysis and empirical formulas to slope parameters, a model fitting based on measured profile data from the study area is performed using an exponential equilibrium profile model. The formula is as follows:

[0128] h(x)=B(1-e -kx In equation (6), x is the horizontal distance from the coastline, h is the corresponding water depth, B is the asymptotic water depth scale, and k is the lateral transition rate constant. Further, this is extended by establishing a functional relationship between the tide level and the distance from the shore as independent variables, and obtaining the slope β at any offshore location by taking the first derivative of this function. f ,Right now:

[0129] h′(x)=tanβ f (x)=Bke -kx (7)

[0130] Please see Figure 6 Equation (6) allows the slope under any combination of tide level and distance from shore to be directly called, thereby controlling the wave rise η sl With the oscillation of the flow η osc The introduction of high-resolution, time-varying slope input in the piecewise solution provides strict terrain constraints for refined calculation of climbing elevation extrema.

[0131] S4. Draw a topographic map of the beach profile based on the distance from the shore and elevation data of the breakpoints at each time point.

[0132] In this embodiment, the distance from the shore of the breakpoint at each moment is input into the slope model to obtain the local slope of the breakpoint. The local slope corresponding to each breakpoint, the nearshore tide level measured by the tide gauge station, and the wave elements such as the effective wave height and period obtained by the wave buoy at that moment are substituted into equation (5) to calculate the elevation of the breakpoint. The elevation of each breakpoint is obtained. Based on the elevation and distance from the shore of each breakpoint, a topographic map of the corresponding beach profile is drawn.

[0133] Traditional methods (such as on-site instrument measurement, satellite remote sensing, and aerial photogrammetry) cannot simultaneously meet the demands for high frequency, continuous operation, high resolution, low cost, and real-time processing. This embodiment utilizes a Stacked Bi-LSTM network for deep learning to achieve precise extraction of waterline breakpoints in target profiles under single-pixel image pillars. The waterline elevation, corrected for wave run-up using a local slope model, is assigned to the waterline breakpoints at the corresponding time and geographical location, completing high spatiotemporal resolution beach profile reconstruction. This provides strong technical support for nearshore beach dynamic geomorphology research and refined management, aiming to solve the current bottleneck problems faced by nearshore dynamic geomorphology monitoring research.

[0134] Example 2

[0135] Please see Figure 7 An apparatus for extracting intertidal topographic profiles based on time-averaged video imagery, based on the method described in one embodiment, wherein the system comprises:

[0136] The acquisition module is used to acquire continuous video images of the intertidal zone and correct them to a planar geographic coordinate system through a coordinate transformation model.

[0137] The recognition module is used to establish a single-value mapping between ground coordinates and image plane coordinates, and back-project the pixels of the video image onto a unified geographic grid. Based on the single-value mapping, it constructs an affine matrix of single-pixel image columns, and embeds the geographic direction vector of the target profile into the spatial transformation of the image columns. Along the direction of the target profile, it extracts image column samples of the video image at each time point according to a preset pixel width. It uses a neural network model to identify the position of the waterline breakpoint in the image columns.

[0138] The elevation assignment module is used to calculate the elevation of the waterline by referring to the local slope model based on the equilibrium profile and assign the calculation results to the corresponding waterline breakpoints.

[0139] The terrain module is used to draw topographic maps of the beach profile based on the distance from the shore and elevation data of the breakpoints at each time point.

[0140] In the above technical solution, in order to better use the method described in one of the embodiments, this application proposes an extraction device for intertidal topographic profile based on time-averaged video images. Each module corresponds to each step of the above method, and its specific principle has been described above and will not be repeated here.

[0141] Example 3

[0142] An apparatus for extracting intertidal topographic profiles based on time-averaged video imagery, comprising:

[0143] At least one processor and a memory communicatively connected to said at least one processor;

[0144] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method as described in one of the embodiments.

[0145] In the above technical solution, in order to better operate and process the method described in one of the embodiments, the method is stored in a memory, and the stored method is executed by a processor. It should be noted that the principle and effect of each step have been described above and will not be elaborated further here.

[0146] Example 4

[0147] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in one of the embodiments.

[0148] In the above technical solutions, to better operate and use the method described in one of the embodiments, the method is stored in a computer-readable storage medium, and a processor is used to implement the method described in one of the embodiments. It should be noted that the principle and effect of each step have been described above and will not be elaborated further here.

[0149] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for extracting intertidal topographic profiles based on time-averaged video imagery, characterized in that, The method includes: Acquire continuous video images of the intertidal zone and correct them to a planar geographic coordinate system using a coordinate transformation model; establish a single-value mapping between ground coordinates and image plane coordinates, and back-project the pixels of the video images onto a unified geographic grid. Based on single-value mapping, an affine matrix of single-pixel image columns is constructed, and the geographic direction vector of the target profile is embedded in the spatial transformation of the image columns. Image column samples of each moment in the video image are extracted along the target profile direction according to a preset pixel width. The location of the waterline breakpoint in the image column is identified using a neural network model. Referring to the local slope model based on the equilibrium profile, the elevation of the waterline is calculated and the calculation results are assigned to the corresponding waterline breakpoints. A topographic map of the beach profile was drawn based on the distance from the shore and elevation data of the breakpoints at each time point.

2. The method for extracting intertidal topographic profiles based on time-averaged video imagery as described in claim 1, characterized in that, If continuous high-frequency video images from different perspectives are acquired through several camera positions, then after back-projecting the pixels onto a unified geographic grid, the process also includes: By projecting images acquired from several camera positions onto the sea surface height that changes over time using time-varying tide levels, and then using a weighted stitching strategy based on distance transformation to complete the image fusion of the several camera positions.

3. The method for extracting intertidal topographic profiles based on time-averaged video imagery as described in claim 1, characterized in that, Based on single-value mapping, an affine matrix of single-pixel image columns is constructed, and the geographic direction vector of the target profile is embedded in the spatial transformation of the image columns. Image pillar samples are extracted from the video image at various time points along the target profile direction, with a preset pixel width, including: Based on single-valued mapping, an affine matrix for single-pixel image pillars is constructed, allowing the image pillars to advance along a given target profile direction; let the profile geographic increment vector be: In the formula, (x1,y1), (x s y1 and y2 represent the start and end points of the target profile image coordinates, respectively; M represents the pixel row index of the image column; The affine matrix of a single-pixel image column is: The affine matrix embeds the geographic direction vector of the target profile into the spatial transformation of the image bars. Along the target profile direction, image bar samples are extracted from the video image at each time point according to a preset pixel width. Then, the world coordinates of the pixel in the i-th row are (x1+i*dX, y1+i*dY), and its azimuth angle θ is calculated using the following formula:

4. The method for extracting intertidal topographic profiles based on time-averaged video imagery as described in claim 1, characterized in that, Using a neural network model to identify the locations of waterline breakpoints in image columns, including: A pre-trained stacked Bi-LSTM network model processes single-pixel image pillars column-by-column; this network model includes: The input subnetwork is used to input the width and column height of the image bars as time steps and the pixel channel X as features into the stacked Bi-LSTM network for time-by-time encoding. The feature extraction subnetwork includes two bidirectional Bi-LSTM layers, each of which contains a forward LSTM layer and a backward LSTM layer. This feature extraction subnetwork is used to extract temporal features from the forward and backward directions of the sequence, respectively, to capture contextual information. The classification subnetwork is used to perform binary classification of each pixel of the image column as "land" or "ocean" based on the features extracted by the feature extraction subnetwork; The output and feedback subnetworks are used to output the final classification results, identify the positions of different classifications, and determine the last "land" pixel in the image column closest to the "ocean" side as the breakpoint of the waterline.

5. The method for extracting intertidal topographic profiles based on time-averaged video imagery as described in claim 4, characterized in that, The dual-layer bidirectional Bi-LSTM layer includes: The first bidirectional Bi-LSTM layer encodes the local visual modal information of the input sequence through a gating mechanism, extracts the spatiotemporal distribution features of fine-grained chromaticity gradient and texture primitives, and forms a discriminative representation of the underlying visual features. The secondary bidirectional Bi-LSTM layer takes the output of the primary bidirectional Bi-LSTM layer as input, models longer-range contextual dependencies through hidden state interactions across time steps, incorporates the prior structural constraint of "land-ocean single jump" in the feature abstraction process, and achieves the integration of cross-regional semantic associations through hidden layer dynamic evolution, ultimately generating feature vectors that have both high-order semantic discriminativeness and structural consistency.

6. The method for extracting intertidal topographic profiles based on time-averaged video imagery as described in claim 1, characterized in that, The local slope model based on the equilibrium profile is established by including: Collect several RTK-GPS measured beach profile data and fit a profile model based on the beach profile data; Based on the profile fitting model, a beach alluvial plain slope model is established. Using this beach alluvial plain slope model and the breakpoint locations at each time point, the local slope corresponding to the breakpoint location is obtained in real time.

7. The method for extracting intertidal topographic profiles based on time-averaged video imagery as described in claim 1, characterized in that, The formula for calculating the waterline elevation is as follows: In the formula, Z sl Z represents the elevation of the waterline. o K represents the nearshore water level elevation, a combination of astronomical tides and meteorological data. osc Representing the empirical coefficient of impulsive flow, β s Let H0 be the average slope of the sluice zone, H0 be the significant wave height in nearshore waters, and L0 be the wavelength in deep offshore waters. L0 = gT 2 / 2π, where T is the effective wave period and g is the gravitational acceleration.

8. A device for extracting intertidal topographic profiles based on time-averaged video imagery, characterized in that, Based on the method according to any one of claims 1-7, the system comprises: The acquisition module is used to acquire continuous video images of the intertidal zone and correct them to a planar geographic coordinate system through a coordinate transformation model; it establishes a single-value mapping between ground coordinates and image plane coordinates and back-projects the pixels of the video images onto a unified geographic grid. The recognition module is used to construct the affine matrix of single-pixel image pillars based on single-value mapping, embedding the geographic direction vector of the target profile into the spatial transformation of the image pillars; along the target profile direction, it extracts image pillar samples of the video image at each time point according to a preset pixel width; and uses a neural network model to identify the position of the waterline breakpoint in the image pillars. The elevation assignment module is used to calculate the elevation of the waterline by referring to the local slope model based on the equilibrium profile and assign the calculation results to the corresponding waterline breakpoints. The terrain module is used to draw topographic maps of the beach profile based on the distance from the shore and elevation data of the breakpoints at each time point.

9. A device for extracting intertidal topographic profiles based on time-averaged video imagery, characterized in that, include: At least one processor and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.