High arch dam construction vibration parameter identification method based on GNSS and binocular vision
By combining GNSS positioning, a pan-tilt head (PTZ) and a thermal infrared binocular camera, the problems of redundant equipment and low recognition accuracy during the vibration of high arch dams were solved, achieving simple, stable and accurate identification of high arch dam concrete vibration parameters.
Patent Information
- Application Number
- CN202510862230.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-19
AI Technical Summary
The existing technology has complicated equipment, difficult installation, large mechanical interference, and easy damage to sensors during the high arch dam concrete vibration process. In addition, binocular vision technology cannot fully identify vibration parameters, especially when the vibrating rod sticks to the concrete. The recognition accuracy is reduced.
Combining GNSS positioning, a gimbal, and a thermal infrared binocular camera, calibration and data processing are used to achieve three-dimensional reconstruction and real-time monitoring of the vibrator. A generative adversarial neural network and a multi-task prediction head are used to identify vibration parameters, and the gimbal controls the camera angle to keep the vibrator in the center of the image.
Simplify equipment, reduce mechanical interference, improve the accuracy and comprehensiveness of vibration parameter identification, and achieve high-quality detection during high arch dam construction.
Smart Images

Figure CN120672964A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for identifying vibration construction parameters of a high arch dam, and more particularly to a method for identifying vibration construction parameters of a high arch dam based on GNSS and binocular vision. Background Art
[0002] Concrete vibration of high arch dam is an extremely important part of its pouring construction. The quality of vibration will directly affect the concrete structure, and then affect the performance of the dam and even the safety of the dam. [1][2] Vibration is the process of transferring energy to concrete through the vibration of the vibrating rod. [3] , promoting the relative movement between the components in the concrete and making the components more evenly distributed, while at the same time expelling the gas retained in the concrete, ultimately achieving a more compact concrete and a more even distribution of sand and gravel aggregates. This improves the quality of concrete and enables the concrete structure to achieve its earthquake resistance, impermeability, and frost resistance properties. [4][5] . Therefore, as a concrete dam, it is necessary to study the concrete vibration of high arch dam. The current means of high arch dam vibration is mainly to use a vibrating trolley in combination with a manual vibrating rod to perform concrete vibration operations on high arch dams. The vibration quality is mainly controlled by the vibration parameters such as vibration position, vibration duration, insertion depth and insertion angle based on the subjective experience of the workers during the vibration process, and the concrete vibration quality is evaluated by core sampling afterwards. In order to solve the problems of this method, such as strong subjectivity and low accuracy, point detection of coring operations, inability to make comprehensive judgments, and sampling as a post-evaluation that cannot be remedied in real time, some scholars have proposed a real-time monitoring system for concrete vibration. [ 6 As the concrete vibration construction of high arch dams enters the intelligent era, various new technologies have been introduced into the real-time monitoring system of high arch dam concrete vibration operations, such as drones, high-precision GNSS positioning and other air-space-ground integrated technologies to identify the vibration position of high arch dam concrete, and use attitude sensors and laser ranging sensors to monitor the vibration time, insertion depth and insertion angle of the vibrating rod of the high arch dam concrete in real time. [7] However, these current technologies have drawbacks such as requiring too much equipment, redundant sensors, difficult on-site equipment installation, significant mechanical interference, and easily damaged sensors. Therefore, new technical means are urgently needed to update and upgrade the real-time monitoring system for high arch dam concrete vibration construction to address the current problems in the real-time monitoring system for high arch dam concrete vibration construction.
[0003] Currently, there is a significant amount of research on concrete vibration monitoring in similar civil engineering projects, focusing on the conventional areas of vibrator energy transfer, vibration parameters, and vibration quality identification. Compared to high arch dam concrete vibration, new technologies are being employed in concrete vibration quality identification, such as binocular vision, thermal infrared recognition, and sound recognition. Existing research on binocular vision primarily focuses on identifying the vibrator's position (vibration point), while limited research has focused on parameters such as vibration duration, insertion depth, and insertion angle. Therefore, current binocular vision technology has shortcomings in concrete vibration monitoring. Furthermore, limited research exists on thermal infrared recognition technology, primarily focusing on laboratory-based identification. Furthermore, identification is performed after concrete pouring, with limited research on the concrete during vibration. Therefore, all of these technologies can be applied to high arch dam concrete vibration construction monitoring, but there is room for improvement.
[0004] Currently, binocular vision technology has been combined with deep learning and has made great progress (references). Currently, there are patents that use a combination of binocular vision technology and deep learning technology to monitor the position of the vibrating rod. As mentioned above, there is a gap in the concrete vibration parameters, and it is impossible to capture all the identification parameters of conventional concrete vibration. Most of the existing studies that use binocular vision technology to identify vibrating rods use a binocular camera with a fixed viewing angle to shoot the vibration process, and the shooting angle of the binocular camera cannot be adjusted. At the same time, existing studies mainly use visible light binocular vision technology to identify vibrating rods. During the concrete vibration process, the vibrating rod will be wrapped by concrete, and the deep learning model has difficulty in identifying the vibrating rod in the vibration image. Therefore, there is still room for improvement in image recognition. At the same time, deep learning algorithms are developing rapidly, and the algorithm iteration is fast. Therefore, there is an urgent need to introduce new deep learning models into the vibrating rod recognition work to improve the model recognition accuracy.
[0005] In summary, existing research on binocular vision measurement of high arch dam concrete vibration parameters is insufficient, and it is unable to identify all parameters of high arch dam concrete vibration. Furthermore, traditional deep learning using visible light images is difficult due to the vibrating rod's adhesion to the concrete. Furthermore, the use of thermal infrared binocular vision technology for concrete parameter identification is limited, making it impossible to fully and accurately collect high arch dam concrete vibration construction parameters.
[0006] [1] Zhong Denghua, Ren Bingyu, Song Wenshuai, et al. Research on key technologies and applications of intelligent control of construction progress and quality of high arch dams[J].
[0007] Hydropower Technology, 2019, 50(08): 8-17.
[0008] [2] Wang Xiaoling, Wang Dong, Ren Bingyu, et al. Research and application of high arch dam concrete vibration robot system[J]. Journal of Hydraulic Engineering
[0009] Journal of Journal of Chinese Academy of Sciences, 2022, 53(06): 631-643+654.
[0010] [3]Li J,Tian Z,Sun X,et al.Modeling vibration energy transfer offresh concrete and energydistribution visualization system[J].Constructionand Building Materials,2022,354:129210.[4]Ma Y,Tian Z,Xu Process[J].Materials,2023,16(8):2958.
[0011] [5]Bang JS, Yim H J.Segregation evaluation of concrete pavementsunder excessive vibration using electrical resistivity measurement[J].CaseStudies in Construction Materials,2023,19:
[0012] e02300.
[0013] [6] Zhong Denghua, Shen Ziyang, Wang Jiajun, et al. Research on dynamic evaluation of concrete dam vibration construction quality based on real-time monitoring[J].
[0014] Journal of Chinese Academy of Sciences, 2018, 49(07): 775-786. DOI: 10.13243 / j.cnki.slxb.20180122.
[0015] [7] Wang Dong, Guan Tao, Yang Shuai, et al. Intelligent monitoring of concrete vibration quality under integrated air-space-ground perception[J].
[0016] Journal of Medical Engineering, 2023, 51(05): 1219-1227. DOI: 10.14062 / j.issn.0454-5648.20220881.
[0017] [8] Xiong Mudi, Zhao Yongjie, Qi Chao, et al. System and method for measuring ship height based on omnidirectional vision expansion of pan-tilt platform and binocular camera[P].
[0018] Liaoning Province: CN202310930381.2, 2023-11-10.
[0019] [9] Zhai Zhiqiang, Xiong Kun, Li Ran, et al. A camera pan-tilt system with active tracking function[P]. Beijing
[0020] City:CN202210621838.7,2024-12-03.
[0021]
[10] Li Bo, Ding Xia, He Runrun, et al. Vibrator positioning method based on binocular vision[P]. Shaanxi
[0022] Province:CN201910351691.2,2019-10-18.
[0023]
[11] Chen Yuntao, Huang Zhe, Li Xinru, et al. Real-time monitoring method of dynamic compaction settlement based on binocular vision and neural network model[P].
[0024] Jinshi:CN202310672126.2,2024-03-22.
[0025]
[12] Wang Hongbo, Zhang Yao, Zhang Jingrui, et al. A binocular vision position measurement system and method based on deep learning[P]. Beijing
[0026] City:CN202110550638.2,2023-03-24.
[0027]
[13] Xie Xiaohui. A method for object recognition and positioning of mobile robots based on binocular vision[P]. Guangdong
[0028] Province:CN202310938453.8,2023-11-10.
[0029]
[14] Chen Hongyue, Chen Qi, Yang Xinwei, et al. A measurement system and method for attitude angle of advanced hydraulic support based on binocular vision[P].
[0030] Liaoning Province: CN202310843553.2, 2023-10-10.
[0031]
[15] Yu Lili, Song Heng, Ouyang Huimin. A three-dimensional measurement method of crane swing angle based on binocular vision[P]. Jiangsu
[0032] Province:CN202310997802.3,2023-11-17. Summary of the Invention
[0033] The technical problem to be solved by the present invention is to provide a more simple, stable and accurate method for identifying vibration parameters of high arch dam construction based on GNSS and binocular vision in order to overcome the shortcomings of the existing technology.
[0034] The technical solution adopted by the present invention is: a method for identifying vibration parameters of high arch dam construction based on GNSS and binocular vision, comprising the following steps:
[0035] 1) Install the GNSS (satellite navigation system) positioning device, pan-tilt platform, and thermal infrared binocular camera on the top of the vibrating trolley. The thermal infrared binocular camera is installed on the pan-tilt platform. The installation position of the thermal infrared binocular camera on the pan-tilt platform is relatively fixed to the installation position of the GNSS antenna, and the coordinates of the two can be converted;
[0036] 2) Calibrate the thermal infrared binocular camera using the thermal infrared binocular camera calibration module installed on the industrial computer, and obtain the intrinsic parameter matrix, distortion coefficient, and extrinsic parameter matrix of the left and right cameras;
[0037] 3) Use a thermal infrared binocular camera to photograph the vibrating rod of the vibrating trolley during high arch dam construction. The left and right cameras shoot synchronously to obtain photos of the vibrating rod corresponding to the left and right cameras.
[0038] 4) Obtaining the coordinate information of the vibrator in the image by processing and identifying the captured photos;
[0039] 5) The control system installed on the industrial computer sends instructions to the pan-tilt head according to the coordinate information of the vibrator in the image, controls the pitch and yaw angles of the pan-tilt head, and keeps the vibrator in the center of the image.
[0040] The high arch dam construction vibration parameter identification method based on GNSS and binocular vision of the present invention has the following characteristics and beneficial effects:
[0041] The method of the present invention utilizes GNSS, a pan-tilt head (PTZ), and a thermal infrared binocular camera to monitor the concrete vibration process of high arch dams in real time. Compared to existing real-time monitoring systems for high arch dam concrete vibration, the method of the present invention features minimal equipment, simple and convenient installation, and minimal mechanical interference. This method effectively addresses several issues currently faced by these systems. Current research on using binocular vision technology to monitor the concrete vibration process primarily relies on fixed-position cameras, while a lack of research on using binocular vision technology for mobile vibrating trolleys has been reported. The method of the present invention combines GNSS positioning technology with thermal infrared binocular vision technology to enable monitoring of the vibration process using a binocular camera mounted on a mobile platform. Current research on monitoring vibrators using binocular vision technology primarily relies on fixed-viewing binocular cameras to capture the vibration process, making it impossible to adjust the camera's shooting angle. The method of the present invention incorporates a PTZ into binocular vision technology for monitoring the construction process of high arch dams, enabling high-quality inspection of the entire construction process. To address the issue of vibrator rod adhesion, which reduces rod recognition accuracy, a method using a thermal infrared binocular camera is proposed to enhance rod recognition accuracy. At the same time, the current technology of using binocular vision technology to identify vibrating rods is not comprehensive enough in identifying the vibration parameters. The method of the present invention performs parallax estimation based on the images taken by the left and right cameras, and then generates a three-dimensional point cloud of the vibrating rod, and performs three-dimensional reconstruction of the vibrating rod. It can comprehensively identify various parameters in the vibration construction process, thereby improving the engineering application effect of real-time monitoring of high arch dam vibration. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is the flow chart of the neural network for vibrator parameter identification;
[0043] Figure 2 It is the flow chart of the object recognition control neural network. DETAILED DESCRIPTION
[0044] The following describes in detail the method for identifying vibration parameters for high arch dam construction based on GNSS and binocular vision in conjunction with the embodiments and drawings.
[0045] As a concrete dam, high arch dams require real-time acquisition of concrete vibration parameters. In view of the problems of the current real-time monitoring system for high arch dam concrete vibration, such as too many devices required, redundant sensors, difficulty in on-site equipment installation, large mechanical interference, and easy damage to sensors, and the current binocular vision technology is still incomplete for vibration parameter identification and mainly uses visible light images for identification, while traditional visible light images have difficulty in identifying vibrating rods of sticky concrete. The present invention proposes to use thermal infrared binocular vision technology to identify the construction parameters of high arch dam concrete vibration. In view of the problems that the current binocular vision technology cannot adapt to the complex construction environment of high arch dams and the existing research on vibration parameter identification is still incomplete, the present invention proposes a method for identifying high arch dam construction vibration parameters that integrates GNSS, pan-tilt and thermal infrared binocular vision. A new idea is proposed for the identification of high arch dam concrete vibration construction parameters.
[0046] The method for identifying vibration parameters for high arch dam construction based on GNSS and binocular vision of the present invention comprises the following steps:
[0047] 1) Install the GNSS (satellite navigation system) positioning device, pan-tilt platform, and thermal infrared binocular camera on the top of the vibrating trolley. The thermal infrared binocular camera is installed on the pan-tilt platform. The installation position of the thermal infrared binocular camera on the pan-tilt platform is relatively fixed to the installation position of the GNSS positioning antenna, and the coordinates of the two can be converted;
[0048] The GNSS positioning device is composed of a GNSS receiver and a GNSS positioning antenna. The latitude, longitude and elevation data of the vibrating trolley are obtained in real time according to the GNSS positioning antenna and the GNSS receiver installed on the vibrating trolley. Real-time differential positioning is performed using a ground RTK differential base station to obtain the earth coordinate system positioning data of the GNSS positioning antenna. According to the relative positions of the pan-tilt platform, the thermal infrared binocular camera and the GNSS positioning antenna, the earth coordinate system information of the pan-tilt platform and the thermal infrared binocular camera is obtained, that is, the GNSS positioning antenna coordinate system can be converted to the pan-tilt platform base coordinate system through a fixed offset, and the thermal infrared binocular camera coordinate system can be converted from the pan-tilt platform base coordinate system through a fixed offset and the attitude deflection of the thermal infrared binocular camera.
[0049] The relative relationship between the GNSS positioning antenna coordinate system and the thermal infrared binocular camera coordinate system is:
[0050] O c =(O g +t c-g )·Z z
[0051] Z z =Z(α p )·Z(α f )
[0052] Among them, the center point O of the GNSS positioning antenna g O is the origin of the GNSS positioning antenna coordinate system. c is the coordinate origin of the thermal infrared binocular camera coordinate system, t c-g is the fixed offset of the GNSS positioning antenna and the thermal infrared binocular camera, obtained by physical measurement. The attitude of the thermal infrared binocular camera is determined by the deflection angle α of the gimbal. p and pitch angle α f Indicates that Z z is the gimbal attitude transformation matrix, = Z(α p ) is the gimbal deflection angle α p The transformation matrix, Z(α f ) is the gimbal pitch angle α f The transformation matrix.
[0053] The thermal infrared binocular camera is assembled by symmetrically fixing two left and right cameras on a fixed bracket, and the center points of the left and right camera connecting parts coincide with the center point of the pan / tilt platform.
[0054] 2) Calibrate the thermal infrared binocular camera using the thermal infrared binocular camera calibration module installed on the industrial computer, and obtain the intrinsic parameter matrix, distortion coefficient, and extrinsic parameter matrix of the left and right cameras;
[0055] The thermal infrared binocular camera calibration module calibrates the thermal infrared binocular camera using the Zhang Zhengyou calibration method to obtain the intrinsic parameter matrix, extrinsic parameter matrix, and distortion coefficient of the thermal infrared binocular camera. The conversion relationship between the pixel coordinate system and the world coordinate system of the image captured by the thermal infrared binocular camera is:
[0056]
[0057] Among them, Z C is the depth value, (u,v) is the image pixel coordinate, dx, dy is the pixel coordinate size, (u0,v0) is the coordinate of the intersection of the camera optical axis and the image plane, f is the focal length, R is the rotation matrix, T is the translation matrix, and (X,Y,Z) is the world coordinate.
[0058] 3) Use a thermal infrared binocular camera to photograph the vibrating rod of the vibrating trolley during high arch dam construction. The left and right cameras shoot synchronously to obtain photos of the vibrating rod corresponding to the left and right cameras.
[0059] 4) Obtaining the coordinate information of the vibrator in the image by processing and identifying the captured photos;
[0060] The processing is to extract frames from the images taken by the left and right cameras, extracting one frame every 10 seconds to obtain a one-to-one corresponding image data set for the left and right cameras. Specifically:
[0061] The thermal infrared image data obtained by the left and right cameras is enhanced by constructing a generative adversarial neural network (GAN), that is, restoring details and suppressing noise. The architecture of the generative adversarial neural network adopts an improved U-Net+Respath structure as the generator. The encoder includes 3 downsampling layers and 8 residual blocks. The decoder reuses the underlying features through skip connections and introduces void convolution to expand the receptive field. The generative adversarial neural network adopts a multi-scale PatchGAN structure as the discriminator. The training adopts a composite loss function. The loss function is as follows:
[0062] L=λ1L d +λ2L n +λ3L TV
[0063] Among them, L is the total loss, L d To combat the loss, L n is the content loss, L tv is TV regularization, λ1, λ2, λ3 are the loss weights.
[0064] The identification is performed by a vibrator parameter recognition neural network, which is constructed based on a dual-path EfficientNet-B3 and is used to simultaneously identify the vibrator depth information and the vibrator working status; the vibrator parameter recognition neural network is as follows: Figure 1 Shown, including:
[0065] (1) Dual-path feature extraction network. The backbone of the dual-path feature extraction network uses a dual-path EfficientNet-B3 to form a feature extraction framework. The feature extraction framework uses a compound scaling strategy to balance depth, width, and resolution, and uses a mobile inverted bottleneck structure to reduce computational complexity. At the same time, a parameter sharing strategy is adopted. In the first half of the stage, the shallow layer fully shares weights to extract image edge and texture features. In the second half of the stage, segment-independent weight parameters are used to adapt to the perspective difference between the left and right views. The dual-path feature extraction network uses a multi-scale output method to extract three-level features: C3 (56×56), C4 (28×28), and C5 (14×14), corresponding to details, semantics, and global information, respectively.
[0066] (2) a stereo feature fusion module, which implements efficient parallax perception through three-stage processing;
[0067] In the first stage, group correlation calculation is performed, which divides the feature channels into 8 groups. The pixel matching cost of the left and right views is calculated in each group, and the disparity search range is dynamically adjusted to output a cost volume with the dimension of [B, 8, D, H, W].
[0068] In the second stage, attention weighted fusion is performed, using the SE attention mechanism to calibrate the channel importance of the left and right branch features respectively;
[0069] In the third stage, 3D convolution fusion is performed, using a 1×3×3 3D convolution kernel to perform joint filtering in the spatial-parallax dimension. The 3D features are projected onto a 2D plane through channel compression, and a 128-channel fused feature map is output.
[0070] (3) The neural network for vibrator parameter recognition is designed with a multi-task prediction head, including:
[0071] The object detection prediction head uses an improved YOLO detection architecture to predict three anchor boxes for each detection layer. The loss function uses a combination of CIoU Loss and Focal Loss, and introduces a feature pyramid to fuse C3-C5 multi-scale features.
[0072] The depth estimation prediction head uses coarse-grained global prediction to restore the half-resolution depth map based on upsampling of the fused feature map. It uses the Huber loss function with edge constraints, adopts fine-grained local correction, performs depth residual regression within the detection box ROI, and jointly optimizes the geometric consistency of the detection box position and depth value to generate the depth map;
[0073] The attribute prediction head adopts a multi-label hierarchical prediction strategy, directly regressing the first-level attribute "vibrator" and the second-level attribute "work vs. non-work";
[0074] The labeled data set is divided into 70% training set, 15% validation set and 15% test set.
[0075] (4) The vibrator parameter recognition neural network adopts the strategy of adaptive multi-task loss balance to design the loss function, which is as follows:
[0076]
[0077] Among them, L0 is the total loss, L1 is the detection loss, L2 is the depth loss, L3 is the attribute loss, σ i It is a dynamically adjustable task weight;
[0078] (5) The depth map obtained by the depth estimation prediction head in the vibrating rod parameter recognition neural network is integrated with the coordinate system conversion formula to realize the measurement of the vibration parameters.
[0079] 5) The control system installed on the industrial computer sends instructions to the pan-tilt head according to the coordinate information of the vibrator in the image, controls the pitch and yaw angles of the pan-tilt head, and keeps the vibrator in the center of the image.
[0080] The control system adopts the object recognition control neural network, and the pan-tilt tracking system is designed based on the image captured by the left camera. The pan-tilt tracking system is divided into three parts: object recognition, tracking and pan-tilt control. Figure 2 As shown, specifically including:
[0081] The object recognition part uses the image captured by the left camera to create an object tracking dataset, and uses the LabelImg image annotation tool to annotate the vibrator in the image to create a labeled object tracking dataset; the labeled object tracking dataset is used to train the YOLOv8 model as an object recognition model, output the coordinates of the upper left corner and lower right corner of the object bounding box (u1, v1, u2, v2), and then calculate the coordinates of the object center point u, v, where the horizontal coordinate u = (u1 + u2) / 2 and the vertical coordinate v = (v1 + v2) / 2;
[0082] The object tracking part is based on the DeepSORT algorithm. First, the tracker is initialized, the coordinates of the upper left and lower right corners of the object bounding box output by the object recognition part (u1, v1, u2, v2) are received, the Kalman filter state vector is initialized, including position and velocity, the object area image is extracted, and a 128-dimensional feature vector is generated through the Re-ID model; then frame-by-frame tracking is performed, and the Kalman filter is used to predict the position of the target in the next frame as a priori estimation. The detection box of YOLOv8 and the predicted box are calculated for IoU and feature cosine similarity matching. The comprehensive matching score is: IoU weight 0.7, feature weight 0.3, and the Hungarian algorithm is used to assign the best match. If the match is successful, the Kalman filter state is updated. If the match fails, the object is marked as pending confirmation. If there is no match for three consecutive frames, the tracker is deleted; then the coordinates of the center point of the bounding box of the tracked object (u, v) are output, and the object velocity vector is passed to the gimbal control part for PID differential term calculation;
[0083] The gimbal control part calculates the desired gimbal deflection angle based on the output of the object tracking part, and then outputs control instructions through the PID controller to achieve gimbal motion control. The gimbal part also returns the actual angle for calibrating the prediction error. The gimbal deflection angle calculation formula is:
[0084]
[0085] Among them, θ s is the horizontal rotation angle of the gimbal, θ c is the vertical rotation angle of the PTZ, W is the width of the image, H is the height of the image, α p is the gimbal deflection angle, α f is the gimbal pitch angle.
[0086] Therefore, the overall logic of the gimbal tracking system is that the object recognition module triggers system initialization and abnormal recovery with high-precision detection, the tracking module maintains target lock through motion prediction and feature matching, and dynamically coordinates the rhythm of detection and control. The gimbal control module converts algorithm decisions into mechanical actions, and at the same time affects the perception input through feedback, ultimately realizing binocular gimbal tracking and shooting of objects.
[0087] The present invention's GNSS- and binocular-vision-based method for identifying vibration parameters during high arch dam construction, along with its data fusion and control system, integrates data from a GNSS system, a pan-tilt system, and a thermal infrared binocular camera. This data fusion enables real-time positioning and measurement of observed objects. The control system also controls the pan-tilt system to maintain the observed object in the center of the image captured by the left camera. This data fusion and binocular-vision technology enable comprehensive identification of vibration parameters during high arch dam construction.
Claims
1. A method for identifying vibration parameters in high arch dam construction based on GNSS and binocular vision, characterized in that: The steps include: 1) Install the GNSS (satellite navigation system) positioning device, pan-tilt platform, and thermal infrared binocular camera on the top of the vibrating trolley. The thermal infrared binocular camera is installed on the pan-tilt platform. The installation position of the thermal infrared binocular camera on the pan-tilt platform is relatively fixed to the installation position of the GNSS antenna, and the coordinates of the two can be converted; 2) Calibrate the thermal infrared binocular camera using the thermal infrared binocular camera calibration module installed on the industrial computer, and obtain the intrinsic parameter matrix, distortion coefficient, and extrinsic parameter matrix of the left and right cameras; 3) Use a thermal infrared binocular camera to photograph the vibrating rod of the vibrating trolley during high arch dam construction. The left and right cameras shoot synchronously to obtain photos of the vibrating rod corresponding to the left and right cameras. 4) Obtaining the coordinate information of the vibrator in the image by processing and identifying the captured photos; 5) The control system installed on the industrial computer sends instructions to the pan-tilt head according to the coordinate information of the vibrator in the image, controls the pitch and yaw angles of the pan-tilt head, and keeps the vibrator in the center of the image.
2. The method for identifying vibration parameters of high arch dam construction based on GNSS and binocular vision according to claim 1 is characterized in that: The GNSS positioning device described in step 1) is composed of a GNSS receiver and a GNSS positioning antenna. The latitude, longitude and elevation data of the vibrating trolley are obtained in real time according to the GNSS positioning antenna and the GNSS receiver installed on the vibrating trolley. Real-time differential positioning is performed using a ground RTK differential base station to obtain the earth coordinate system positioning data of the GNSS positioning antenna. According to the relative positions of the pan-tilt platform, the thermal infrared binocular camera and the GNSS positioning antenna, the earth coordinate system information of the pan-tilt platform and the thermal infrared binocular camera is obtained.
3. The method for identifying vibration parameters of high arch dam construction based on GNSS and binocular vision according to claim 1 is characterized in that: The relative relationship between the GNSS positioning antenna coordinate system and the thermal infrared binocular camera coordinate system in step 1) is: O c =(O g +t c-g )·Z z WITH z =Z(α p )·Z(α f ) Among them, the center point O of the GNSS positioning antenna g O is the origin of the GNSS positioning antenna coordinate system. c is the coordinate origin of the thermal infrared binocular camera coordinate system, t c-g is the fixed offset of the GNSS positioning antenna and the thermal infrared binocular camera, obtained by physical measurement. The attitude of the thermal infrared binocular camera is determined by the deflection angle α of the gimbal. p and pitch angle α f Indicates that Z z is the gimbal attitude transformation matrix, = Z(α p ) is the gimbal deflection angle α p The transformation matrix, Z(α f ) is the gimbal pitch angle α f The transformation matrix.
4. The method for identifying vibration parameters of high arch dam construction based on GNSS and binocular vision according to claim 1 is characterized in that: The thermal infrared binocular camera described in step 1) is assembled by symmetrically fixing two left and right cameras on a fixed bracket, and the center points of the left and right camera connecting parts coincide with the center point of the pan / tilt platform.
5. The method for identifying vibration parameters of high arch dam construction based on GNSS and binocular vision according to claim 1 is characterized in that: The thermal infrared binocular camera calibration module described in step 2) calibrates the thermal infrared binocular camera by using the Zhang Zhengyou calibration method to obtain the intrinsic parameter matrix, extrinsic parameter matrix and distortion coefficient of the thermal infrared binocular camera. The conversion relationship between the pixel coordinate system and the world coordinate system of the image captured by the thermal infrared binocular camera is: Among them, Z C is the depth value, (u,v) is the image pixel coordinate, dx, dy is the pixel coordinate size, (u0,v0) is the coordinate of the intersection of the camera optical axis and the image plane, f is the focal length, R is the rotation matrix, T is the translation matrix, and (X,Y,Z) is the world coordinate.
6. The method for identifying vibration parameters of high arch dam construction based on GNSS and binocular vision according to claim 1 is characterized in that: Step 4) is to perform frame extraction on the images captured by the left and right cameras, extracting one frame every 10 seconds to obtain a one-to-one corresponding image data set for the left and right cameras. Specifically: The thermal infrared image data obtained by the left and right cameras is enhanced by constructing a generative adversarial neural network, that is, restoring details and suppressing noise. The architecture of the generative adversarial neural network adopts an improved U-Net+Respath structure as the generator. The encoder includes 3 downsampling layers and 8 residual blocks. The decoder reuses the underlying features through skip connections and introduces void convolution to expand the receptive field. The generative adversarial neural network adopts a multi-scale PatchGAN structure as the discriminator. The training adopts a composite loss function. The loss function is as follows: L=λ1L d +λ2L n +λ3L TV Among them, L is the total loss, L d To combat the loss, L n is the content loss, L tv is TV regularization, λ1, λ2, λ3 are the loss weights.
7. The method for identifying vibration parameters of high arch dam construction based on GNSS and binocular vision according to claim 1 is characterized in that: Step 4) The identification is performed by a vibrator parameter recognition neural network, which is built based on a dual-channel EfficientNet-B3 and is used to simultaneously identify the vibrator depth information and the vibrator working status. The vibrator parameter recognition neural network includes: (1) Dual-path feature extraction network. The backbone of the dual-path feature extraction network uses a dual-path EfficientNet-B3 to form a feature extraction framework. The feature extraction framework uses a compound scaling strategy to balance depth, width, and resolution, and uses a mobile inverted bottleneck structure to reduce computational complexity. At the same time, a parameter sharing strategy is adopted. In the first half of the stage, the shallow layer fully shares weights to extract image edge and texture features. In the second half of the stage, segment-independent weight parameters are used to adapt to the perspective difference between the left and right views. The dual-path feature extraction network uses a multi-scale output method to extract three-level features: C3 (56×56), C4 (28×28), and C5 (14×14), corresponding to details, semantics, and global information, respectively. (2) a stereo feature fusion module, which implements efficient parallax perception through three-stage processing; In the first stage, group correlation calculation is performed, which divides the feature channels into 8 groups. The pixel matching cost of the left and right views is calculated in each group, and the disparity search range is dynamically adjusted to output a cost volume with the dimension of [B, 8, D, H, W]. In the second stage, attention weighted fusion is performed, using the SE attention mechanism to calibrate the channel importance of the left and right branch features respectively; In the third stage, 3D convolution fusion is performed, using a 1×3×3 3D convolution kernel to perform joint filtering in the spatial-parallax dimension. The 3D features are projected onto a 2D plane through channel compression, and a 128-channel fused feature map is output. (3) The neural network for vibrator parameter recognition is designed with a multi-task prediction head, including: The object detection prediction head uses an improved YOLO detection architecture to predict three anchor boxes for each detection layer. The loss function uses a combination of CIoU Loss and Focal Loss, and introduces a feature pyramid to fuse C3-C5 multi-scale features. The depth estimation prediction head uses coarse-grained global prediction to restore the half-resolution depth map based on upsampling of the fused feature map. It uses the Huber loss function with edge constraints, adopts fine-grained local correction, performs depth residual regression within the detection box ROI, and jointly optimizes the geometric consistency of the detection box position and depth value to generate the depth map; The attribute prediction head adopts a multi-label hierarchical prediction strategy, directly regressing the vibrator as the first-level attribute and work and non-work as the second-level attributes; (4) The vibrator parameter recognition neural network adopts the strategy of adaptive multi-task loss balance to design the loss function, which is as follows: Among them, L0 is the total loss, L1 is the detection loss, L2 is the depth loss, L3 is the attribute loss, σ i It is a dynamically adjustable task weight; (5) The depth map obtained by the depth estimation prediction head in the vibrating rod parameter recognition neural network is integrated with the coordinate system conversion formula to realize the measurement of the vibration parameters.
8. The method for identifying vibration parameters of high arch dam construction based on GNSS and binocular vision according to claim 1 is characterized in that: The control system described in step 5) uses an object recognition control neural network, and the pan / tilt tracking system is designed based on the image captured by the left camera. The pan / tilt tracking system is divided into three parts: object recognition, tracking, and pan / tilt control, specifically including: The object recognition part uses the image captured by the left camera to create an object tracking dataset, and uses the LabelImg image annotation tool to annotate the vibrator in the image to create a labeled object tracking dataset; the labeled object tracking dataset is used to train the YOLOv8 model as an object recognition model, output the coordinates of the upper left corner and lower right corner of the object bounding box (u1, v1, u2, v2), and then calculate the coordinates of the object center point u, v, where the horizontal coordinate u = (u1 + u2) / 2 and the vertical coordinate v = (v1 + v2) / 2; The object tracking part is based on the DeepSORT algorithm. First, the tracker is initialized, the coordinates of the upper left and lower right corners of the object bounding box output by the object recognition part (u1, v1, u2, v2) are received, the Kalman filter state vector is initialized, including position and velocity, the object area image is extracted, and a 128-dimensional feature vector is generated through the Re-ID model; then frame-by-frame tracking is performed, and the Kalman filter is used to predict the position of the target in the next frame as a priori estimation. The detection box of YOLOv8 and the predicted box are calculated for IoU and feature cosine similarity matching. The comprehensive matching score is: IoU weight 0.7, feature weight 0.3, and the Hungarian algorithm is used to assign the best match. If the match is successful, the Kalman filter state is updated. If the match fails, the object is marked as pending confirmation. If there is no match for three consecutive frames, the tracker is deleted; then the coordinates of the center point of the bounding box of the tracked object (u, v) are output, and the object velocity vector is passed to the gimbal control part for PID differential term calculation; The gimbal control part calculates the desired gimbal deflection angle based on the output of the object tracking part, and then outputs control instructions through the PID controller to achieve gimbal motion control. The gimbal part also returns the actual angle for calibrating the prediction error. The gimbal deflection angle calculation formula is: Among them, θ s is the horizontal rotation angle of the gimbal, θ c is the vertical rotation angle of the PTZ, W is the width of the image, H is the height of the image, α p is the gimbal deflection angle, α f is the gimbal pitch angle.