Three-dimensional wave field time sequence visual intelligent reconstruction and dynamic parameter evaluation system

The three-dimensional wave field time-series visual intelligent reconstruction system, which integrates multi-frame time-series fusion and adaptive extrinsic parameter analysis, solves the problems of insufficient real-time performance, strong dependence on extrinsic parameters, and weak environmental adaptability in existing technologies. It achieves high-precision wave field monitoring and dynamic analysis and is suitable for multi-platform environments.

CN121746293APending Publication Date: 2026-03-27HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing wave field observation and 3D reconstruction technologies suffer from insufficient real-time performance, strong dependence on external parameters, poor dynamic consistency, and weak environmental adaptability, making it difficult to meet the dual requirements of high-precision 3D reconstruction and time-series dynamic analysis.

Method used

Employing a multi-frame temporal fusion mechanism and an adaptive extrinsic parameter analysis model, this system couples image acquisition, stereo vision gated perception, temporal fusion, camera extrinsic parameter analysis, and hydrodynamic parameter evaluation modules to achieve continuous reconstruction of wave surface geometry and accurate inversion of dynamic hydrodynamic characteristics. It is suitable for non-contact wave monitoring on multiple platforms, including shore-based, shipborne, and UAV-based systems.

Benefits of technology

It enables stable and continuous non-contact wave monitoring and hydrodynamic analysis under complex sea conditions and dynamic platform conditions. It features high precision, strong robustness and good versatility, and is suitable for nearshore observation, shipborne monitoring, ocean energy assessment and intelligent navigation sensing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746293A_ABST
    Figure CN121746293A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional wave field time sequence vision intelligent reconstruction and dynamic parameter evaluation system, and belongs to the crossing field of computer vision and ocean engineering. The system comprises six coupling modules including an image acquisition module, a stereoscopic vision gating sensing module, a time sequence fusion module, a camera external parameter analysis module, a three-dimensional reconstruction module and a hydrodynamic parameter evaluation module, and end-to-end collaboration is realized by unifying spatial-temporal characteristic flow. The image acquisition module acquires multi-platform synchronous images, high-precision space-time consistent depth information is extracted through the stereoscopic vision gating sensing and time sequence fusion module, the camera external parameter analysis module realizes uncalibrated attitude self-correction, and the three-dimensional reconstruction module generates low-noise wave surface data. The hydrodynamic parameter evaluation module inverts parameters such as wave height, wave velocity, energy spectrum and the like and predicts an evolution trend. The system has high precision, strong robustness and multi-platform adaptability, and is suitable for non-contact wave monitoring scenes such as near-shore observation, shipborne monitoring and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of computer vision and marine engineering, and specifically relates to a three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system. Background Technology

[0002] With the deepening of intelligent ocean observation and hydrodynamic process research, how to achieve high-precision, non-contact three-dimensional reconstruction and dynamic parameter inversion of wave fields under complex and variable sea conditions has become a key scientific and technological challenge in marine engineering, ship navigation, and marine energy development. The spatiotemporal evolution of waves is directly related to the force response and energy transfer laws of marine structures, but the current refined observation and dynamic characterization of wave fields still face significant bottlenecks.

[0003] Existing contact wave measurement methods (such as buoys, wave stakes, and pressure sensors) can provide local wave height or period information, but their spatial resolution is limited, they are significantly affected by environmental interference, and their equipment deployment is complex and maintenance costs are high, making it difficult to operate stably for a long time in harsh sea conditions. Non-contact measurement technologies (such as radar wave measurement, laser scanning, and satellite remote sensing) have wide coverage capabilities, but their spatial and temporal resolution is insufficient, and they are significantly limited by meteorological conditions and imaging angles. They are unable to continuously and in real-time capture details such as local nonlinear fluctuations and short-period breaking waves, thus failing to meet the requirements for accurate reconstruction and time-series analysis of highly dynamic wave fields.

[0004] The rise of computer vision technology has provided a new research path for the 3D reconstruction of wave fields. Binocular stereo vision-based wave surface reconstruction methods offer advantages such as non-contact operation, high resolution, and flexible deployment, enabling direct recovery of wave geometry from optical images. However, these methods typically rely on explicit calibration and static feature matching, making them highly sensitive to changes in illumination, sea surface reflection, and rapid wave surface movement. In real-world marine environments, they are prone to feature drift and matching errors, leading to unstable depth estimation and error accumulation. Furthermore, traditional stereo matching algorithms are primarily based on local window or cost volume optimization, resulting in high computational cost, poor real-time performance, and limited adaptability to complex dynamic scenes.

[0005] In recent years, the introduction of deep learning has improved the matching accuracy and robustness of stereo vision to some extent. However, models based on convolutional neural networks or 3D convolution still have significant shortcomings: First, the deep network structure leads to high computational complexity and low inference efficiency, making it difficult to meet the real-time processing requirements of dynamic wave fields. Second, the models have insufficient generalization performance under limited samples or new scene conditions, requiring a large amount of labeled data. Third, existing networks are mostly optimized for static scenes and lack effective modeling of temporal consistency and dynamic features, making it difficult to accurately capture the continuous changes in the wave surface. In addition, traditional stereo vision methods rely on fixed camera calibration parameters. On dynamic platforms such as shipboard or UAVs, extrinsic parameter drift will seriously affect the accuracy and coordinate consistency of 3D reconstruction, further limiting their applicability in real ocean observation.

[0006] In summary, existing wave field observation and 3D reconstruction technologies generally suffer from insufficient real-time performance, strong dependence on external parameters, poor dynamic consistency, and weak environmental adaptability. They are unable to meet the dual requirements of high-precision 3D reconstruction and time-series dynamic analysis, which restricts the development of marine environmental monitoring and non-contact hydrodynamic assessment technologies. Summary of the Invention

[0007] This invention aims to overcome the problems of insufficient real-time performance, strong dependence on external parameters, poor dynamic consistency, and weak environmental adaptability in existing three-dimensional wave field reconstruction technologies. It proposes a three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system. By introducing a multi-frame temporal fusion mechanism and an adaptive external parameter analysis model, the system realizes continuous reconstruction of wave surface geometry and accurate inversion of dynamic hydrodynamic characteristics. It is suitable for non-contact wave monitoring environments on multiple platforms such as shore-based, shipborne, and UAV.

[0008] The technical solution adopted by this invention to solve the technical problem is as follows:

[0009] A three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system includes an image acquisition module, a stereo vision gated perception module, a temporal fusion module, a camera extrinsic parameter analysis module, a three-dimensional reconstruction module, and a hydrodynamic parameter evaluation module coupled in sequence. Each module achieves end-to-end collaboration through a unified spatiotemporal feature flow, forming an integrated process of dynamic perception and physical parameter inversion.

[0010] The image acquisition module is used to simultaneously acquire multiple wave image sequences on shore-based, shipborne, or UAV platforms. It employs a high frame rate camera array with a fixed baseline and uses hardware triggering to achieve synchronous exposure of the left and right cameras and consecutive frames. The acquired images are jointly calibrated using a pinhole model and a radial distortion model, and then subjected to epipolar correction and adaptive illumination compensation to maintain spatiotemporal consistency.

[0011] The stereo vision gated perception module extracts high-precision depth information based on depth feature matching and gated control mechanism. It utilizes a multi-layer feature pyramid and variable receptive field mechanism to achieve adaptive perception of waves at different scales, effectively distinguishing between reflection areas and real wave crests and troughs. By updating the correlation field of the left and right views through iterative perception units, a continuous parallax optimization chain is established, ensuring that the wavefront reconstruction results maintain consistency between high-frequency changes and low-frequency fluctuations.

[0012] The temporal fusion module achieves spatiotemporal dynamic consistency reconstruction of wave images through multi-frame feature fusion and cyclic feedback mechanism, automatically learns the spatiotemporal dependence of wave propagation, and suppresses matching drift caused by rapid motion or camera shake through cross-frame residual compensation, thereby improving the continuity and stability of reconstruction.

[0013] The camera extrinsic parameter analysis module, based on adaptive geometric constraints and feature space consistency estimation, calculates camera attitude changes under dynamic platform conditions, updates intrinsic and extrinsic parameters in real time, and unifies the coordinate system, maintaining temporal consistency of the 3D reconstructed coordinate system without calibration plates or external sensors. This module achieves self-updating of intrinsic and extrinsic parameters and coordinate system unification by estimating the geometric offset of the feature field between consecutive frames, without requiring external calibration plates or GPS-assisted equipment. This mechanism can maintain temporal consistency of the 3D reconstructed coordinate system under conditions of ship swaying, platform pitching, or UAV attitude changes.

[0014] The 3D reconstruction module generates a dense point cloud based on the stereo matching and temporal fusion results, and restores wavefront details through multi-scale geometric optimization and feature-guided interpolation to obtain high-resolution, low-noise 3D wavefront temporal data.

[0015] The hydrodynamic parameter evaluation module extracts wave height, wave velocity, energy spectrum, and hydrodynamic parameters of the main propagation direction based on the reconstructed three-dimensional wavefront time series data, and constructs a frequency-direction joint spectrum model. It combines the elevation changes extracted from the visual domain with the fluid dynamics model to realize the mapping from image space to physical quantity space, and can predict and analyze the short-term wave evolution trend based on energy distribution and spectral peak migration law.

[0016] Furthermore, the stereo vision gated perception module is used to extract high-precision depth information based on binocular image sequences, and constructs a deep network model through a feature encoder, a correlation pyramid, and a multi-layer GRU update unit; its depth optimization process satisfies the following definition:

[0017]

[0018] In the formula, These are pixel coordinates; These are the feature maps of the left and right images in the feature coding layer, respectively. Candidate values ​​for depth; Let be the feature similarity function; the network uses an iterative recursive mechanism to update the depth estimate, as expressed below:

[0019]

[0020] In the formula, s represents the GRU layer number; This represents the hidden state at layer s; It is a convolutional gate unit; For contextual features; For depth increment; through multi-scale feature pyramids, joint modeling of peaks and troughs at different scales is achieved;

[0021] The network employs a multi-layer stacked GRU structure and achieves cross-frame deep consistency learning through hidden state propagation; the expression satisfied by the hidden state update is as follows:

[0022]

[0023] In the formula, To update the gate control quantity; subscript For time indexing; This is the candidate hidden state; This indicates the Hadamard product operation.

[0024] Furthermore, the stereo vision gated perception module employs a multi-stage depth supervision loss function, the expression of which is as follows:

[0025]

[0026] In the formula, It is the attenuation factor; For the mask matrix, For smoothing loss; It is a constant; this loss function causes the depth estimation to converge gradually during the iterative phase.

[0027] The expressions for spatial smoothing and edge preservation loss terms are also included, as follows:

[0028]

[0029] in,

[0030]

[0031] To reduce the constraints of textured areas and enhance the smoothness of flat areas;

[0032] The expression for the timing consistency loss function is defined as follows:

[0033]

[0034] In the formula, The current frame depth; Displacement is estimated for optical flow to constrain the depth consistency of the same physical point at consecutive time steps;

[0035] The expression for the KL divergence regularization based on wave height distribution in hydrodynamic statistical constraints is as follows:

[0036]

[0037] In the formula, and These are the probability density functions for the actual measured and predicted wave heights, respectively, used to improve the model's fit to the statistical characteristics of waves.

[0038] Taking all constraints into account, the overall loss can be expressed as follows:

[0039]

[0040] In the formula, The weight coefficients for each loss term are adjusted using the validation set to obtain the optimal combination, ensuring that the reconstruction results achieve a balance between spatial accuracy and temporal consistency.

[0041] Furthermore, the network employs the AdamW optimizer, with an initial learning rate set to... The single-cycle learning rate strategy is used to achieve rapid convergence in the early stages of training and smooth fine-tuning in the later stages, with training samples... The random cropped block input is set to Furthermore, random scaling, saturation perturbation, and right-view vertical offset were added to the binocular image pairs to simulate common alignment errors and illumination changes in real observations, thereby improving the model's domain generalization ability. The training lasted for 200 epochs, with each epoch containing 7200 image pair samples.

[0042] Furthermore, the temporal fusion module is used to establish spatiotemporal dynamic consistency among multiple frames of images. It updates the depth features of consecutive frames through multi-frame feature fusion and a recursive feedback mechanism, with the update relationship being:

[0043]

[0044] In the formula, As a weighting factor; These are the features of the current frame and the previous frame, respectively; To predict residuals, a cross-frame residual compensation mechanism is used to suppress matching drift caused by camera shake and high-frequency wavefront motion.

[0045] Furthermore, the camera extrinsic parameter analysis module is used to dynamically estimate camera attitude changes and unify the three-dimensional coordinate system; the module estimates the geometric offset between point clouds in consecutive frames, using the following expression for the geometric consistency constraint function:

[0046]

[0047] In the formula, The wavefront normal vector; For planar offset terms; Point cloud coordinates; by minimizing The wavefront normal estimation and attitude calculation are performed, resulting in the following expression for the non-homogeneous rigid body transformation matrix from the camera to the wave coordinate system:

[0048]

[0049] In the formula, It is a rotation matrix; It is a translation vector;

[0050] For any camera system 3D point Its position in the wave coordinate system The expression for the coordinates below is as follows:

[0051] .

[0052] Furthermore, the three-dimensional reconstruction module expresses the wavefront elevation field as follows:

[0053]

[0054] The expression for Gaussian filtering is as follows:

[0055]

[0056] Remove discrete noise to obtain a continuous and smooth three-dimensional wavefront.

[0057] Furthermore, the hydrodynamic parameter evaluation module is used to calculate wave height, energy spectrum, and velocity field parameters based on time-series three-dimensional wavefront data. The expression for its vertical velocity field is as follows:

[0058]

[0059] The expression for horizontal velocity is as follows:

[0060]

[0061] In the formula, Wave number; It is the acceleration due to gravity; Phase velocity;

[0062] The hydrodynamic parameter evaluation module obtains the wave energy spectrum expression through three-dimensional fast Fourier transform as follows:

[0063]

[0064] In the formula, Represents wavenumber components in different directions; The expression representing frequency, and converting it to a direction-frequency spectrum, is as follows:

[0065]

[0066] In the formula, This indicates the wave direction angle, used to extract the main propagation direction and energy distribution of the wave.

[0067] Furthermore, the three-dimensional wave field time-series visual intelligent reconstruction and dynamic parameter evaluation system described above is characterized by further including a visualization module, which is used to generate a three-dimensional pseudo-color elevation map and dynamic wave surface animation based on the reconstructed wave surface, and supports energy spectrum superposition display, main wave direction tracking and velocity vector field visualization output.

[0068] Furthermore, as described above, the three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system achieves end-to-end network training through joint loss constraints. After attitude correction and wave surface fitting, the output depth field forms a continuous and high-fidelity three-dimensional wave model, which can realize the automatic extraction and predictive analysis of wave height, energy and velocity fields.

[0069] The overall technical solution of this invention establishes an end-to-end mapping mechanism from raw images to dynamic wave physical parameters by fusing deep visual perception with adaptive geometric modeling. The system design incorporates temporal recursive feature updates and a multi-scale correlation pyramid structure, resulting in stronger spatiotemporal consistency and environmental adaptability in wavefront morphology reconstruction. Its main technical highlights are reflected in the following three aspects:

[0070] (1) Adaptive pose analysis mechanism without external calibration: The camera pose self-estimation and spatial alignment under the dynamic platform are realized through geometric consistency learning, eliminating the dependence of traditional stereo vision on fixed calibration plates.

[0071] (2) Multi-frame recursive temporal fusion mechanism: Through temporal feature feedback and residual compensation, continuous reconstruction of wave surface and short-term prediction are realized, which improves the continuity and stability of wave field dynamic identification.

[0072] (3) Visual-physical dual-domain fusion hydrodynamic parameter evaluation framework: establish a unified mapping relationship from visual reconstruction to hydrodynamic parameter inversion, and realize the interpretable transformation of visual signals into physical quantities.

[0073] The beneficial effects and advantages of this invention are as follows: This invention enables stable and continuous non-contact wave monitoring and hydrodynamic analysis under complex sea conditions and dynamic platform conditions. The system possesses high precision, strong robustness, and good versatility, and can be widely applied in nearshore observation, shipborne monitoring, ocean energy assessment, and intelligent navigation sensing, providing reliable technical support for future intelligent ocean dynamics research. Attached Figure Description

[0074] Figure 1 This is a flowchart of the three-dimensional wave field temporal visual intelligent reconstruction method provided in the embodiments of the present invention;

[0075] Figure 2 This is a schematic diagram of the depth label provided in an embodiment of the present invention;

[0076] Figure 3 This is a schematic diagram of the gated recursive temporal stereo vision network structure provided in an embodiment of the present invention;

[0077] Figure 4 This is a schematic diagram comparing the network output depth map with the ground truth provided in an embodiment of the present invention;

[0078] Figure 5 This is a network output error diagram provided in an embodiment of the present invention;

[0079] Figure 6 This is a schematic diagram illustrating the unified transformation from camera coordinates to wave coordinates provided in an embodiment of the present invention;

[0080] Figure 7 This is a schematic diagram of the three-dimensional mapping result of the wave field provided in an embodiment of the present invention;

[0081] Figure 8 This is a schematic diagram of the wave field point cloudification result provided in an embodiment of the present invention;

[0082] Figure 9 This is a graph showing the comparison results of wave height time series evaluation and probability density function provided by the embodiments of the present invention. Detailed Implementation

[0083] To make the technical solution and beneficial effects of the present invention clearer, the following detailed description of a three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system of the present invention is provided in conjunction with the accompanying drawings and specific embodiments. It should be understood that the described embodiments are only for illustrating the present invention and are not intended to limit the scope of protection of the present invention; without departing from the overall concept of the present invention, those skilled in the art can make various modifications and substitutions to the processing flow, network structure, parameter settings, and implementation details in the following embodiments, and all such equivalent modifications or substitutions should fall within the scope of protection of the present invention.

[0084] Example 1

[0085] like Figure 1 As shown, the present invention provides a three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system, including an image acquisition module, a stereo vision gated perception module, a temporal fusion module, a camera extrinsic parameter analysis module, a three-dimensional reconstruction module and a hydrodynamic parameter evaluation module coupled in sequence. Each module achieves end-to-end collaboration through a unified spatiotemporal feature flow, forming an integrated process of dynamic perception and physical parameter inversion.

[0086] The image acquisition module is used to simultaneously acquire multiple wave image sequences on shore-based, shipborne, or UAV platforms. It employs a high frame rate camera array with a fixed baseline and uses hardware triggering to achieve synchronous exposure of the left and right cameras and consecutive frames. The acquired images are jointly calibrated using a pinhole model and a radial distortion model, and then subjected to epipolar correction and adaptive illumination compensation to maintain spatiotemporal consistency.

[0087] The stereo vision gated perception module extracts high-precision depth information based on depth feature matching and gated control mechanism. It utilizes a multi-layer feature pyramid and variable receptive field mechanism to achieve adaptive perception of waves at different scales, effectively distinguishing between reflection areas and real wave crests and troughs. By updating the correlation field of the left and right views through iterative perception units, a continuous parallax optimization chain is established, ensuring that the wavefront reconstruction results maintain consistency between high-frequency changes and low-frequency fluctuations.

[0088] The temporal fusion module achieves spatiotemporal dynamic consistency reconstruction of wave images through multi-frame feature fusion and cyclic feedback mechanism, automatically learns the spatiotemporal dependence of wave propagation, and suppresses matching drift caused by rapid motion or camera shake through cross-frame residual compensation, thereby improving the continuity and stability of reconstruction.

[0089] The camera extrinsic parameter analysis module calculates the camera attitude change under the dynamic platform based on adaptive geometric constraints and feature space consistency estimation, updates intrinsic and extrinsic parameters in real time and unifies the coordinate system, and maintains the temporal consistency of the 3D reconstructed coordinate system under the condition of no calibration plate or external sensor.

[0090] The 3D reconstruction module generates a dense point cloud based on the stereo matching and temporal fusion results, and restores wavefront details through multi-scale geometric optimization and feature-guided interpolation to obtain high-resolution, low-noise 3D wavefront temporal data.

[0091] The hydrodynamic parameter evaluation module extracts wave height, wave velocity, energy spectrum, and hydrodynamic parameters of the main propagation direction based on the reconstructed three-dimensional wavefront time series data, and constructs a frequency-direction joint spectrum model. It combines the elevation changes extracted from the visual domain with the fluid dynamics model to realize the mapping from image space to physical quantity space, and can predict and analyze the short-term wave evolution trend based on energy distribution and spectral peak migration law.

[0092] Its implementation process mainly includes:

[0093] S1: Acquire binocular wave image sequences from multiple platforms and time points, and perform geometric correction and brightness normalization.

[0094] S2, high-precision depth maps of consecutive frames are obtained based on a gated recursive stereo network;

[0095] S3 performs camera extrinsic self-calibration and coordinate unification on all depth maps to restore the three-dimensional wave surface in a unified wave coordinate system.

[0096] S4 converts the sequence depth field into a three-dimensional point cloud and elevation field to achieve three-dimensional mapping of the wave field;

[0097] S5 extracts hydrodynamic parameters such as wave height, wave velocity, and energy spectrum based on the time-series elevation field and performs statistical evaluation. The specific network training and data processing steps in this process are as follows:

[0098] This embodiment uses an industrial-grade binocular camera, with a single camera resolution of 2058×2456 pixels and a frame rate of 12fps, where the focal length is... The principal point is the coordinate of the center point. The binocular baseline length of the shore-based platform is... This is used to balance near-field wavefront accuracy and depth range. The above parameters are not unique limitations; those skilled in the art can adjust the baseline length within the range of 0.1–10 m according to the observation distance and target wavelength, but the two cameras should be arranged in parallel as much as possible to reduce errors caused by deflection. Fixed-baseline binocular cameras are deployed on shore-based platforms, shipborne platforms, and UAV platforms. Hardware triggering is used to achieve synchronous exposure of the left and right cameras and consecutive frames, obtaining original image pairs:

[0099]

[0100] in , Corresponding pixel coordinates, subscript Indicates the left and right cameras.

[0101] The original image is jointly calibrated using the pinhole model and the radial distortion model to obtain the intrinsic parameter matrix. And distortion parameters. Then, epipolar correction is used to transform the left and right views to a unified image plane, so that corresponding pixels satisfy approximately the horizontal epipolar constraint, i.e.

[0102]

[0103] To reduce contrast differences caused by variations in lighting, image normalization and gamma correction were performed:

[0104]

[0105] in , The mean and standard deviation of the current frame's brightness. The adjustable gamma coefficient is preferably set within the range of [0.6, 1.2]. In this embodiment... To enhance the contrast of water surface textures. To prevent tiny constants with a denominator of zero.

[0106] To further ensure the quality of the input images, this embodiment adds an abnormal frame detection and removal step after normalization and gamma correction. For detected abnormal frames, linear interpolation is used to fill in the gaps between adjacent normal frames to ensure the consistency of the time series.

[0107] In this embodiment, the training dataset consists of measured wave sequences. The measured data comes from a near-shore observation platform, collecting six typical sea state conditions and accumulating approximately 12,000 pairs of original binocular images. Online random cropping, scaling, and color perturbation were applied. Figure 2 As shown, the training phase requires providing deep labels to the network.

[0108] This embodiment uses the SGBM stereo matching algorithm and related image post-processing to fill in offset values ​​to construct wave dataset labels. Specifically, semi-global matching is used to perform stereo matching on the corrected left and right images to obtain the depth field. Then, based on the stereo geometry, the depth is converted to depth in the camera coordinate system:

[0109]

[0110] in The equivalent focal length in the horizontal direction. Baseline length To prevent tiny constants with a denominator of zero.

[0111] Because traditional algorithms have large errors in areas with strong reflections and occlusions, an error mask is constructed for depth labels to avoid introducing incorrect labels into training.

[0112]

[0113] in, Depth estimates obtained by independent algorithms This is the threshold. Only retain [the specified value]. The pixels are used as supervision regions for subsequent network training.

[0114] like Figure 3As shown, the stereo matching network employs a multi-layer GRU recursive field transform structure, specifically comprising three parts: a feature encoder, a correlation pyramid, and a multi-layer GRU update unit. The feature encoder is applied to the left and right image sequences respectively to obtain dense features:

[0115]

[0116] To characterize the similarity between pixels, a one-dimensional correlation volume is constructed in the feature domain:

[0117]

[0118] in, As a deep candidate, It is the vector inner product.

[0119] The relevant volumes are average pooled to obtain multi-level relevance pyramids for matching searches at different scales. During the update phase, the network maintains multi-scale hidden states. (where s represents the scale), gated convolutional units are used to recursively correct the current depth estimate:

[0120]

[0121] in For convolutional GRU units, For contextual features.

[0122] The network employs a multi-layer stacked GRU structure and achieves cross-frame deep consistency learning through hidden state propagation; the expression satisfied by the hidden state update is as follows:

[0123]

[0124] In the formula, To update the gate control quantity; subscript For time indexing; This is the candidate hidden state; This indicates the Hadamard product operation.

[0125] Highest resolution scale output depth increment A new depth estimate is obtained:

[0126]

[0127] To utilize the temporal correlation of the wave field, the network not only receives the current frame image pair but also additionally inputs the left-view features from the previous moment as a temporal prior. For example... Figure 3 As shown, the timing encoder... Perform compression encoding to generate time-series features This is injected additively into each GRU update to guide the network in learning the wavefront evolution trend. The update formula is:

[0128]

[0129] in, As a weighting factor, These are the features of the current frame and the previous frame, respectively. To predict residuals, a cross-frame residual compensation mechanism is used to suppress matching drift caused by camera shake and high-frequency wavefront motion.

[0130] The detailed loss function design and training strategy are as follows:

[0131] 1. Losses from multi-stage in-depth supervision

[0132] A network can generate multiple iterations during a single forward propagation. To encourage gradual convergence, a multi-stage monitoring system with progressively weighted adjustments is introduced:

[0133]

[0134] in, As the attenuation factor, For smoothing loss, It is a small constant.

[0135] 2. Spatial smoothing and edge preservation loss

[0136] To suppress depth noise while preserving high-frequency structures such as peaks and troughs, a spatial regularization based on image gradients is introduced:

[0137]

[0138] in:

[0139]

[0140] That is, reduce the smoothing constraint in areas with rich texture and increase the smoothing constraint in areas with flat sea surface.

[0141] 3. Temporal consistency loss

[0142] By leveraging the continuity of depth prediction between adjacent frames, constraints are imposed on the depth variation of the same physical point at consecutive time points. Considering the optical flow approximation of displacement, a definition is defined...

[0143]

[0144] in Depth is derived from the camera model. This loss encourages smooth changes in the depth field over time, reducing flicker artifacts.

[0145] 4. Wave height statistical constraints

[0146] The statistical characteristics of the wave height time series should be consistent with the reference measurement. Wave height sequences should be extracted from both network predictions and baseline measurements. , The probability density function is obtained through kernel density estimation. , And calculate the KL divergence between distributions as statistical regularity:

[0147]

[0148] 5. Comprehensive Objective Function

[0149] Taking all constraints into account, the overall loss is

[0150]

[0151] in The weighting coefficients are determined through adjustment using the validation set. In this embodiment, a preferred set of parameters is obtained by performing a grid search on the validation set for different weight combinations. The network reconstruction accuracy and stability remain at a relatively high level. Those skilled in the art can fine-tune the above weights according to specific applications without exceeding the scope of protection of this invention.

[0152] 6. Training Strategies

[0153] The network uses the AdamW optimizer, with an initial learning rate set to... This approach utilizes a single-cycle learning rate strategy for rapid convergence in the early stages of training and smooth fine-tuning in the later stages. The training samples are... The random pruning block input is typically set to... Furthermore, random scaling, saturation perturbation, and slight vertical offset of the right view were added to the binocular image pairs to simulate common alignment errors and illumination variations in real-world observations, thereby improving the model's domain generalization ability. Training lasted 200 epochs, with each epoch containing 7200 image pairs. Figure 4 , Figure 5As shown, the depth output of the network in this invention is basically consistent with the actual spatial variation of wave depth, and the maximum error is concentrated in the central region and the four corners of the image, with the maximum depth error not exceeding 0.8%. Depth is converted into parallax for quantitative evaluation, and the average End-Point Error (EPE) is approximately 1.95 pixels, with a Bad1px index (the proportion of pixels with a parallax error greater than 1 pixel) below 16%, indicating that the depth field has good spatial consistency and detail fidelity. On an NVIDIA RTX 4090 GPU, the network's single-frame inference time is less than 1 second, which meets the requirements of online monitoring in engineering projects.

[0154] Subsequent post-processing of the network output results first requires camera extrinsic parameter self-calibration, such as... Figure 6 As shown, the camera coordinate system is denoted as The wave coordinate system is denoted as ,in The axis points perpendicularly to the water surface normal. A planar optimization constraint is established:

[0155] Suppose we reconstruct the point cloud Approximately located on the wavefront plane:

[0156]

[0157] Define the optimization objective:

[0158]

[0159] The plane normal vector is obtained through a random sampling consensus algorithm. With offset .

[0160] Solving for the rotation matrix and translation vector using the alignment constraints of point cloud planes in adjacent frames:

[0161]

[0162] The final transformation matrix is ​​obtained as follows:

[0163]

[0164] Achieve dynamic attitude compensation.

[0165] Based on the estimation results from the extrinsic parameter analysis module, the non-homogeneous rigid body transformation matrix from camera coordinates to wave coordinates can be obtained:

[0166]

[0167] in, Let be a rotation matrix. It is a translation vector.

[0168] For any camera system 3D point Its coordinates in the wave coordinate system are:

[0169]

[0170] A wavefront elevation field and 3D mapping are performed. Through the above transformation, each frame of depth map is converted into a point cloud set in the wave coordinate system. The elevation field is obtained by interpolating the data using a rule-based grid:

[0171]

[0172] Then filtered by a Gaussian filter:

[0173]

[0174] This removes discrete noise and yields a continuous and smooth three-dimensional wavefront.

[0175] like Figure 7 As shown, pseudo-color elevation data can be overlaid on the original grayscale image to obtain a three-dimensional mapping result of the wave field; simultaneously, a grid surface can be drawn in three-dimensional space to visually display the wave crest and trough structure. Figure 8 The wave field point cloudification result is shown.

[0176] The hydrodynamic parameter evaluation module transforms the elevation field obtained from visual reconstruction into wave dynamic parameters, including velocity field, wave height, and energy spectrum. Specifically:

[0177] Vertical velocity:

[0178]

[0179] Horizontal speed:

[0180]

[0181] in Let be the phase velocity.

[0182] Wave height is defined as:

[0183]

[0184] Energy density:

[0185]

[0186] in This is the density of water.

[0187] Perform a Fast Fourier Transform on the time series:

[0188]

[0189] in This represents a three-dimensional Fourier transform. The direction-frequency spectrum is obtained through polar coordinate transformation, used to characterize the energy propagation path and the direction of the dominant wave. It is then converted into a direction-frequency spectrum:

[0190]

[0191] in, , Representing wavenumber components in different directions, Indicates frequency, This is the wave propagation direction angle, used to extract the main propagation direction and energy distribution of the wave.

[0192] Figure 9 The results of wave time series assessment and probability density function comparison with the true values ​​under a certain working condition are presented. Taking a nearshore irregular wave measured condition as an example, the statistical results compared with the benchmark measurement (binocular stereo vision results) show that the root mean square error (RMSE) of the wavefront elevation field reconstructed by the method of this invention is approximately 0.015-0.08m within the observation area, and the error increases towards the boundary position. Significant wave height is indicated by [missing information]. 𝑠 The relative error was controlled within 5%, and the average wave height relative error did not exceed 7.5%. Further testing under typical complex sea conditions such as low light, strong wave reflection, and slight lens shake showed that the reconstructed RMSE increase was less than 20%, indicating that the method of this invention still has good robustness under conditions of varying illumination and attitude disturbances. In summary, the experiments demonstrate that this invention has significant advantages in accuracy, dynamic consistency, and environmental adaptability, and can achieve stable and reliable non-contact wave monitoring on shore-based, shipborne, and UAV platforms, showing promising engineering application prospects.

Claims

1. A three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system, characterized in that, It includes sequentially coupled image acquisition module, stereo vision gated perception module, temporal fusion module, camera extrinsic parameter analysis module, 3D reconstruction module and hydrodynamic parameter evaluation module. Each module achieves end-to-end collaboration through a unified spatiotemporal feature flow, forming an integrated process of dynamic perception and physical parameter inversion. The image acquisition module is used to simultaneously acquire multiple wave image sequences on shore-based, shipborne, or UAV platforms. It employs a high frame rate camera array with a fixed baseline and uses hardware triggering to achieve synchronous exposure of the left and right cameras and consecutive frames. The acquired images are jointly calibrated using a pinhole model and a radial distortion model, and then subjected to epipolar correction and adaptive illumination compensation to maintain spatiotemporal consistency. The stereo vision gated perception module extracts high-precision depth information based on depth feature matching and gated control mechanism. It utilizes a multi-layer feature pyramid and variable receptive field mechanism to achieve adaptive perception of waves at different scales, effectively distinguishing between reflection areas and real wave crests and troughs. By updating the correlation field of the left and right views through iterative perception units, a continuous depth optimization chain is established, ensuring that the wavefront reconstruction results maintain consistency between high-frequency changes and low-frequency fluctuations. The temporal fusion module achieves spatiotemporal dynamic consistency reconstruction of wave images through multi-frame feature fusion and cyclic feedback mechanism, automatically learns the spatiotemporal dependence of wave propagation, and suppresses matching drift caused by rapid motion or camera shake through cross-frame residual compensation, thereby improving the continuity and stability of reconstruction. The camera extrinsic parameter analysis module calculates the camera attitude change under the dynamic platform based on adaptive geometric constraints and feature space consistency estimation, updates intrinsic and extrinsic parameters in real time and unifies the coordinate system, and maintains the temporal consistency of the 3D reconstructed coordinate system under the condition of no calibration plate or external sensor. The 3D reconstruction module generates a dense point cloud based on the stereo matching and temporal fusion results, and restores wavefront details through multi-scale geometric optimization and feature-guided interpolation to obtain high-resolution, low-noise 3D wavefront temporal data. The hydrodynamic parameter evaluation module extracts wave height, wave velocity, energy spectrum, and hydrodynamic parameters of the main propagation direction based on the reconstructed three-dimensional wavefront time series data, and constructs a frequency-direction joint spectrum model. It combines the elevation changes extracted from the visual domain with the fluid dynamics model to realize the mapping from image space to physical quantity space, and can predict and analyze the short-term wave evolution trend based on energy distribution and spectral peak migration law.

2. The three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system according to claim 1, characterized in that, The stereo vision gated perception module is used to extract high-precision depth information based on binocular image sequences and to construct a deep network model through a feature encoder, a correlation pyramid and a multi-layer GRU update unit. Its deep optimization process satisfies the following definition: In the formula, These are pixel coordinates; These are the feature maps of the left and right images in the feature coding layer, respectively. Candidate values ​​for depth; Let be the feature similarity function; the network uses an iterative recursive mechanism to update the depth estimate, as expressed below: In the formula, s represents the GRU layer number; This represents the hidden state at layer s; It is a convolutional gate unit; For contextual features; For depth increment; through multi-scale feature pyramids, joint modeling of peaks and troughs at different scales is achieved; The network employs a multi-layer stacked GRU structure and achieves cross-frame deep consistency learning through hidden state propagation; the expression satisfied by the hidden state update is as follows: In the formula, To update the gate control quantity; subscript For time indexing; This is the candidate hidden state; This indicates the Hadamard product operation.

3. The three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system according to claim 2, characterized in that, The stereo vision gated perception module uses a multi-stage depth-supervised loss function, expressed as follows: In the formula, It is the attenuation factor; For the mask matrix, For smoothing loss; It is a constant; this loss function causes the depth estimation to converge gradually during the iterative phase. The expressions for spatial smoothing and edge preservation loss terms are also included, as follows: in, To reduce the constraints of textured areas and enhance the smoothness of flat areas; The expression for the timing consistency loss function is defined as follows: In the formula, The current frame depth; Displacement is estimated for optical flow to constrain the depth consistency of the same physical point at consecutive time steps; The expression for the KL divergence regularization based on wave height distribution in hydrodynamic statistical constraints is as follows: In the formula, and These are the probability density functions for the actual measured and predicted wave heights, respectively, used to improve the model's fit to the statistical characteristics of waves. Taking all constraints into account, the overall loss can be expressed as follows: In the formula, The weight coefficients for each loss term are adjusted using the validation set to obtain the optimal combination, ensuring that the reconstruction results achieve a balance between spatial accuracy and temporal consistency.

4. The three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system according to claim 3, characterized in that, The network uses the AdamW optimizer, with an initial learning rate set to... The single-cycle learning rate strategy is used to achieve rapid convergence in the early stages of training and smooth fine-tuning in the later stages, with training samples... The random cropped block input is set to Furthermore, random scale scaling, saturation perturbation, and right view vertical offset are added to the binocular image pairs to simulate alignment errors and illumination changes common in real observations, thereby improving the model's domain generalization ability. The training process lasted for 200 epochs, with each epoch containing 7200 image pairs.

5. The three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system according to claim 4, characterized in that, The temporal fusion module is used to establish spatiotemporal dynamic consistency among multiple frames of images. It updates the depth features of consecutive frames through multi-frame feature fusion and a recursive feedback mechanism. The update relationship is as follows: In the formula, As a weighting factor; These are the features of the current frame and the previous frame, respectively; To predict residuals, a cross-frame residual compensation mechanism is used to suppress matching drift caused by camera shake and high-frequency wavefront motion.

6. The three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system according to claim 5, characterized in that, The camera extrinsic parameter analysis module is used to dynamically estimate camera attitude changes and unify the three-dimensional coordinate system. The module estimates the geometric offset between point clouds in consecutive frames, using the following expression for the geometric consistency constraint function: In the formula, The wavefront normal vector; For planar offset terms; Point cloud coordinates; by minimizing The wavefront normal estimation and attitude calculation are performed, resulting in the following expression for the non-homogeneous rigid body transformation matrix from the camera to the wave coordinate system: In the formula, It is a rotation matrix; It is a translation vector; For any camera system 3D point Its position in the wave coordinate system The expression for the coordinates below is as follows: 。 7. The three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system according to claim 6, characterized in that, The three-dimensional reconstruction module expresses the wavefront elevation field as follows: The expression for Gaussian filtering is as follows: This removes discrete noise and yields a continuous and smooth three-dimensional wavefront.

8. The three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system according to claim 7, characterized in that, The hydrodynamic parameter evaluation module is used to calculate wave height, energy spectrum, and velocity field parameters based on time-series three-dimensional wavefront data. The expression for its vertical velocity field is as follows: The expression for horizontal velocity is as follows: In the formula, Wave number; It is the acceleration due to gravity; Phase velocity; The hydrodynamic parameter evaluation module obtains the wave energy spectrum expression through three-dimensional fast Fourier transform as follows: In the formula, Represents wavenumber components in different directions; The expression representing frequency, and converting it to a direction-frequency spectrum, is as follows: In the formula, This indicates the wave direction angle, used to extract the main propagation direction and energy distribution of the wave.

9. A three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system according to any one of claims 1-8, characterized in that, It also includes a visualization module, which generates a 3D pseudo-color elevation map and dynamic wavefront animation based on the reconstructed wavefront, and supports energy spectrum overlay display, main wave direction tracking and velocity vector field visualization output.

10. A three-dimensional wave field temporal visual intelligent reconstruction and dynamic parameter evaluation system according to any one of claims 1-8, characterized in that, The system achieves end-to-end training of the network through joint loss constraints. After attitude correction and wave surface fitting, the output depth field forms a continuous and high-fidelity three-dimensional wave model, which can realize the automatic extraction and predictive analysis of wave height, energy and velocity fields.