Dynamic spatial data set construction system and method for bridge structural member

The dynamic spatial data set of bridge structural parts is constructed through an algorithm based on the Transformer model, which solves the timing fracture and noise interference problems of the dynamic construction process in the traditional method, real-time detection and continuous tracking of the bridge erection process are realized, and safety monitoring accuracy and accuracy are improved.

CN120544009AActive Publication Date: 2025-08-26NANJING UNIV OF SCI & TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511040596.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-08-26
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

During the traditional bridge erection process, existing data set construction methods cannot capture the dynamic construction process, timing fracture and noise interference lead to point cloud distortion, target detection and tracking are poor real-time and low accuracy, making it difficult to solve the problem of space-time consistency in dynamic scenarios.

Method used

Using an algorithm based on the Transformer model, combined with the image acquisition and processing unit and remote analysis terminal, dynamic spatial data sets of bridge structural components are generated through depth alignment, object detection, spatial topological relationship modeling, 3D coordinate calculation, distance calculation, calibration optimization and space-time fusion, to achieve high-speed and high-precision object detection and tracking.

Benefits of technology

Real-time detection and continuous tracking of the bridge erection process are realized, the training accuracy of the bridge erection operation status monitoring and safety warning model is improved, the timing fracture problem of traditional methods in dynamic construction scenarios is solved, the accuracy and accuracy are improved, and the safety of the bridge erection process is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544009A_ABST
    Figure CN120544009A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic spatial data set construction system and method for a bridge structural member, and the method specifically comprises the steps: collecting a depth map and a color map of a bridge construction scene, and recording the timestamp of each frame of image; carrying out depth alignment and registration, and carrying out real-time target detection by adopting a detection model based on a Transform architecture; spatial topological relation modeling is carried out, 3D coordinates are calculated, and mechanical vibration compensation is carried out based on IMU data; the Euclidean distance between the two target points, the distance between planes where the target points are located and perpendicular to the optical axis and the coordinate axis component are calculated; the space consistency of continuous frames is judged through a finite-state machine, tracking reference calibration is triggered, and track calibration and optimization are carried out; constructing a time-space fusion data sequence, and associating the timestamp and the spatial transformation matrix of each point cloud frame; and marking a dynamic construction label, and generating a dynamic parameter label of the bridge structural member in the construction process. According to the invention, the training precision of the bridging operation state monitoring and safety early warning model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a system and method for constructing a dynamic spatial data set of bridge structural components. Background Art

[0002] During the bridge construction process, excessive local stress can cause structural fracture, while excessive bending deformation can lead to local and overall instability. These problems directly affect the safety of the bridge construction process. Therefore, to ensure construction safety, it is necessary to pay attention to its safety. Strengthening safety monitoring during the bridge construction process has very important social and economic significance.

[0003] Traditional bridge construction monitoring methods have three major flaws: static data limitations, where existing point cloud datasets cannot capture the dynamic construction process; time series interruptions, where a single-frame point cloud struggles to reflect progressive changes such as beam displacement and bolt loosening; and noise interference, where mechanical vibrations cause point cloud distortion, with a maximum error of ±8.7mm.

[0004] Existing multidimensional dataset construction methods mostly rely on BP neural networks, traditional RCNN or YOLO architectures, which are difficult to solve the spatiotemporal consistency problem in dynamic scenes, and the target detection and tracking are poor in real time and low in accuracy. Summary of the Invention

[0005] The purpose of the present invention is to provide a system and method for constructing a dynamic spatial dataset of bridge structural components, which adopts an algorithm based on the Transformer model to achieve high-speed and high-precision target detection and tracking, and improve the training accuracy of the bridge operation status monitoring and safety warning model.

[0006] The technical solution for achieving the purpose of the present invention is: a dynamic spatial data set construction system for bridge structural components, including an image acquisition and processing unit and a remote analysis terminal, wherein the remote analysis terminal is provided with a depth alignment module, a target detection module, a spatial topology relationship modeling module, a 3D coordinate calculation module, a distance calculation module, a calibration optimization module, a spatiotemporal fusion module, and a labeling module;

[0007] The image acquisition and processing unit is used to acquire image data and then send it to the remote analysis terminal via wireless routing;

[0008] The depth alignment module is used for depth alignment data registration;

[0009] The target detection module includes a preprocessing submodule and a model processing submodule; the preprocessing submodule receives the registered depth map processed by the depth alignment module and converts it into the format required by the detection model; the model processing submodule uses a detection model based on the Transformer architecture to process the received image stream, perform real-time target detection, and obtain the target's bounding box and category information, as well as the coordinates of the center point of the bounding box;

[0010] The spatial topology relationship modeling module is used to build a spatial index according to the number of targets, determine the orientation of a single target through the offset, and reset the opposite side mark;

[0011] The 3D coordinate calculation module is used to calculate the coordinates of the target in the 3D world coordinate system and output stable coordinates after compensating for mechanical vibration;

[0012] The distance calculation module is used to calculate the straight-line distance between two target points, the distance between the target points in the plane perpendicular to the optical axis in the 3D world coordinate system, and the coordinate axis components;

[0013] The calibration optimization module is used to smooth the trajectory and output it;

[0014] The spatiotemporal fusion module is used to construct a point cloud temporal sequence and associate a spatial transformation matrix;

[0015] The label marking module is used to generate dynamic parameter labels of bridge structural members during the bridge construction process, including beam center coordinates, depth information, and relative distances between targets.

[0016] Furthermore, the image acquisition and processing unit includes a pan-tilt platform, a mounting assembly, a visual acquisition module and a wireless routing module, wherein:

[0017] The pan-tilt platform is in the shape of a cuboid;

[0018] The mounting assembly is fixed to the top surface of the pan / tilt head;

[0019] The visual acquisition module includes a binocular camera and an inertial measurement unit, which realizes the vertical displacement freedom through the installation component and fixes the spatial posture through a locking mechanism; the binocular camera collects the original depth map and color map,

[0020] The wireless routing module is used to send the collected image data and the inertial data output by the inertial measurement unit to the remote analysis terminal;

[0021] The image acquisition and processing unit is installed upside down below the bridge-building machine and fixed at the mid-span position of the main beam of the bridge-building machine, so that the pan-tilt head faces upward and the binocular camera faces downward, and the optical axis of the binocular camera is set perpendicular to the bridge deck.

[0022] Furthermore, the binocular camera is a binocular imaging unit based on Intel RealSense D435i, which outputs synchronized video stream data with a resolution of 1280×720 pixels and a sampling frequency of 30Hz; the inertial measurement unit uses a BMI085 IMU sensor; the carrying shell and mounting assembly of the visual acquisition module form a rigid frame, so that the inertial measurement unit and the binocular camera form a co-base rigid coupling structure, and the motion state of the inertial measurement unit and the binocular camera are synchronized;

[0023] The wireless routing module is coupled to the data output interface of the image acquisition and processing unit to establish an encrypted communication channel between the image acquisition and processing unit and the remote analysis terminal; the communication channel achieves real-time synchronous transmission of image data from the visual acquisition module and inertial data output by the inertial measurement unit to the remote analysis terminal and cloud storage with a transmission delay of less than 50ms.

[0024] Furthermore, the depth alignment module receives the original depth map and color map from the binocular camera. The depth pixel coordinate system and the color pixel coordinate system are 2D coordinates on the image plane. The function of the depth alignment module is:

[0025] Iterates through each pixel to determine whether the depth value is valid, and then establishes a mapping relationship between the depth camera coordinate system and the color camera coordinate system through a multi-level coordinate space transformation chain; performs pixel-level reprojection based on this mapping relationship to generate a registered depth map that is aligned with the color image space, and outputs the registered depth map to the object detection module;

[0026] The multi-level coordinate space transformation chain, the first level is the conversion from depth pixels to depth camera 3D points, the formula is as follows:

[0027]

[0028] in, is the pixel coordinate of the depth map, is the depth value, is the focal length of the depth camera, is the principal point of the depth camera, The 3D point coordinates in the depth camera coordinate system after conversion;

[0029] The second level is the conversion from depth camera 3D points to color camera 3D points. The formula is as follows:

[0030]

[0031] in, 3 The depth of 3 is converted to a color rotation matrix, 3 1 translation vector, is the 3D point coordinate in the converted color camera coordinate system;

[0032] The third level is the conversion of the color camera 3D points to color pixels. The formula is as follows:

[0033]

[0034] in, is the focal length of the color camera, is the principal point of the color camera, is the pixel coordinate projected onto the color image plane;

[0035] The fourth level is boundary processing and assignment: the floating-point coordinates obtained by projection are converted into integer pixel coordinates, where any pixel coordinate value is set to 0 if it is less than 0, and is set to the maximum value if it exceeds the maximum value of the color pixel coordinate.

[0036] Furthermore, the target detection module has the following specific processing flow:

[0037] The preprocessing submodule is used to convert the input color image into the format required by the detection model, including: adjusting the image size to 800*600 pixels, normalizing the pixel values, converting the image data into the BLOB format required for detection model input, and inputting the BLOB format data into the model processing module;

[0038] The model processing submodule includes the following processing steps ① to ③:

[0039] ① Feature extraction, which is used to extract feature maps from the preprocessed input image through the CNN backbone network;

[0040] ② Transformer processing, which is used to perform the following processing on the feature map in sequence:

[0041] (a) Adding position coding to the feature map: The position coding is generated using sine / cosine functions to encode spatial position information. The formula is as follows:

[0042]

[0043]

[0044] in Represents the position index, Represents the dimension index, represents the feature dimension of the detection model, is the even-dimensional component of the positional encoding, It is the odd-dimensional component of position coding; the position coding dimension is 256;

[0045] (b) Through the Transformer encoder, a multi-head self-attention mechanism is used to calculate the correlation weights between pixels in the feature map to extract and enhance the global feature representation;

[0046] The multi-head self-attention mechanism is as follows:

[0047] Assume that the input of the head attention layer is , firstly transform the input sequence into Mapping to query matrix , key matrix Sum Matrix , the formula is as follows:

[0048]

[0049]

[0050]

[0051] in, is the query weight matrix, is the key weight matrix, is the value weight matrix;

[0052] right 、 and The attention weights are obtained by performing a scaled dot product operation. This process is performed independently in each layer and the calculation process is as follows:

[0053]

[0054] in, is the bond matrix Dimensions, is a normalized exponential function, which normalizes the self-attention score; Compute the output matrix for attention;

[0055] The attention scores of each layer will be spliced ​​in the channel dimension. The result of the layer is the final output of the multi-head attention layer, and the formula is as follows:

[0056]

[0057]

[0058] in, is a linear mapping matrix used to transform The attention scores of the layers are concatenated into a whole; For the The output of the attention head is is the attention head index, is the total number of attention heads, For the The query matrix of each head, For the The key matrix of the head, For the The value matrix of each head, is the final output of the multi-head attention layer;

[0059] (c) Through the Transformer decoder, the object query is used to interact with the enhanced features output by the encoder;

[0060] (d) Through the prediction head, based on the interaction results, a tensor containing bounding box predictions and category predictions is output;

[0061] (e) Outputting the tensor containing the prediction results to the post-processing module;

[0062] ③ Post-processing module, used to parse the input tensor to obtain the detection results, including:

[0063] (a) Extract confidence scores for each prediction result;

[0064] (b) Setting the target confidence threshold to 0.5, if the confidence score is lower than the preset target confidence threshold of 0.5, the corresponding prediction result is discarded;

[0065] (c) for a prediction result whose confidence score satisfies the threshold condition, determining the category ID corresponding to the maximum category score of the prediction result;

[0066] (d) determining whether the target category is a bridge deck or a side of a bridge based on the category ID;

[0067] (e) If the target category is a bridge deck or bridge side, the corresponding normalized bounding box parameters are parsed, including the center point coordinates and width and height, and the normalized bounding box parameters are multiplied by the original image size to obtain the bounding box in actual pixel coordinates;

[0068] (f) recording the center point coordinates of the bounding box in the actual pixel coordinates;

[0069] The target detection module ultimately outputs the bounding box and category information of the bridge deck or bridge side that meets the conditions, as well as the recorded center point coordinates, which are used to construct a dynamic spatial dataset.

[0070] Furthermore, the spatial topology relationship modeling module receives the information output by the target detection module and performs the following processing:

[0071] When there are two targets, an ordered spatial index is constructed based on the horizontal coordinates;

[0072] When there is a single target, compare the offset of the current target horizontal coordinates with the historical benchmark: if the offset is less than the dynamic tolerance threshold, it is determined to be the left topological target and the right target identifier is set to zero; if the offset is not less than the dynamic tolerance threshold, it is determined to be the right topological target and the left target identifier is set to zero.

[0073] Furthermore, the 3D coordinate calculation module includes a coordinate calculation submodule and a mechanical vibration compensation submodule, as follows:

[0074] The coordinate calculation submodule is set as:

[0075] Taking the target center as the reference, extract the depth statistics of the N×N pixel area, where N is a positive integer and N≥6;

[0076] Use back projection to convert the depth data of the registered depth map into the 3D world coordinate system. The formula is as follows:

[0077]

[0078] in, is the pixel coordinate of the depth map, Is the 3D world coordinate system The 3D point coordinates under is the focal length of the depth camera, is the principal point of the depth camera, is the physical depth value;

[0079] The mechanical vibration compensation submodule is set as:

[0080] Receive acceleration and gyroscope data from the inertial measurement unit (IMU). The IMU coordinate system and the color camera coordinate system are related through rigid body transformation to calculate the vibration offset. The formula is:

[0081]

[0082] in, and is the IMU sensor coordinate system 、 Axial acceleration, and is the color camera coordinate system 、 axis angular velocity, is the IMU sensor coordinate system Axis acceleration sensitivity, is the IMU sensor coordinate system Axis rotation sensitivity, is the IMU sensor coordinate system Axis acceleration sensitivity, is the IMU sensor coordinate system Axis rotation sensitivity, is the color camera coordinate system Axis direction compensation offset, is the color camera coordinate system Axis direction compensation offset;

[0083] The vibration offset and the coordinates processed by the coordinate calculation submodule are integrated to perform position compensation. The formula is:

[0084]

[0085] in, is the original 3D world coordinate system Axis coordinates, is the original 3D world coordinate system Axis coordinates, and is the compensated 3D world coordinate system axis, axis coordinates;

[0086] Apply an IIR filter to the compensated position for filtering, the formula is:

[0087]

[0088] in, The coordinates after compensation for the current frame, is the filter position of the previous frame, is the current frame filtering position;

[0089] is the filter coefficient, which is determined using the adaptive filter coefficient formula:

[0090]

[0091] in, is the modulus of the acceleration vector;

[0092] Finally, the 3D coordinate calculation module outputs the stable coordinates to the distance calculation module.

[0093] Furthermore, the distance calculation module is configured as follows:

[0094] The straight-line distance between two target points is determined by the Euclidean distance decomposition algorithm. The formula is:

[0095]

[0096] in Represents the straight-line distance between two target points. The origin of the 3D world coordinate system is the optical center of the camera. The axis is along the optical axis, and the coordinates of the 3D coordinate system of the two target points are and ;

[0097] The 3D world coordinate system xy plane distance and coordinate axis components are calculated by depth difference. The formula is:

[0098]

[0099]

[0100]

[0101]

[0102] in, is the depth difference, Is the 3D world coordinate system Axis pixel difference, Is the 3D world coordinate system Axis pixel difference, is the distance between the two target points on the plane perpendicular to the optical axis, The target point is in the horizontal direction, that is, the 3D world coordinate system The relative displacement of the axis, The target point is in the vertical direction, that is, the 3D world coordinate system The relative displacement of the axis, is the pixel plane distance.

[0103] Furthermore, the calibration optimization module includes a state calibration finite state machine, a trajectory optimizer, and a multimodal output interface, wherein:

[0104] (1) The state calibration finite state machine receives the coordinates output by the 3D coordinate calculation unit and triggers the tracking reference calibration only when the detection results of consecutive frames all meet the spatial consistency threshold, specifically:

[0105] According to the number of detected targets, the detection is divided into three situations: no target, single target, and dual targets:

[0106] If there is no target, the global coordinates are reset to zero;

[0107] If it is a single target, when the target is close to the left, it is marked as the left target, and setting the right target is invalid; when the target is close to the right, it is marked as the right target, and setting the left target is invalid; maintain the 3D world coordinate system The axis coordinates are consistent;

[0108] If there are two targets, the forced sorting ensures that the 3D world coordinate system on the left The value target is smaller than the right side, which will be smaller in the 3D world coordinate system. The target of the coordinate is marked as the left target, and an ordered target pair is generated and output;

[0109] Calculate target distance based on sorted coordinates, verify the validity of coordinate data, and output target location and quantity information;

[0110] (2) The trajectory optimizer receives the calibration signal output by the state calibration finite state machine, builds a timing cache queue to store the target historical position, and calculates the sliding mean based on the window length, as follows:

[0111] Read the sliding window length parameter N and initialize two first-in-first-out queues qx and qy, which are used to store the target in the 3D world coordinate system. Axis and Historical position data in the axis direction;

[0112] Get the current relative position offset output by the distance calculation module frame by frame , remove the oldest historical coordinate value of the queue and add the relative position offset of the current frame to update the queue;

[0113] The updated queue is traversed and accumulated, and a sliding average is calculated. The calculation formula is:

[0114]

[0115] in, is the smoothed 3D world coordinate system Axis coordinates, is the smoothed 3D world coordinate system Axis coordinates, N is the sliding window length, For the historical 3D world coordinate system Coordinate offset queue value, For the historical 3D world coordinate system Coordinate offset queue value;

[0116] (3) The multimodal output interface is coupled to the spatiotemporal fusion module, and outputs the number of targets, spatial position coordinates, real-time spatial offset vectors, and a three-dimensional distance matrix between targets;

[0117] The spatiotemporal fusion module fuses the spatial transformation and the temporal sequence to generate 3D spatiotemporal information for control or decision-making, and outputs it to the visualization module and the storage module for storage in the remote terminal;

[0118] The label marking module receives the target position data output by the target detection module, draws a measurement bounding box in the corresponding detection area in the real-time video stream image, and marks the target's spatial coordinate information and depth measurement results in the detection area.

[0119] A method for constructing a dynamic spatial data set of a bridge structure is provided. The method is based on the dynamic spatial data set construction system for a bridge erection control system and comprises the following steps:

[0120] Step 1: Install the entire hardware module upside down below the bridge crane, with the gimbal facing up and the binocular camera facing down, ensuring that the camera optical axis is perpendicular to the beam below. Continuously capture synchronized depth and color frames of the bridge construction scene at a preset 30Hz sampling frequency, and record the timestamp of each frame. Acceleration and angular velocity data are collected using the inertial measurement unit.

[0121] Step 2: Establish an encrypted communication channel through the wireless routing module to send the image and inertial data to the remote analysis terminal in real time with a transmission delay of less than 50ms;

[0122] Step 3: The depth alignment module of the remote analysis terminal receives the original depth map and color map output by the image acquisition and processing unit, traverses the depth map pixels, and filters the valid depth values; establishes a mapping relationship between the depth camera coordinate system and the color camera coordinate system through a multi-level coordinate space transformation chain, including: depth pixels to depth camera 3D points, depth camera 3D points to color camera 3D points, and color camera 3D points to color pixels; performs boundary processing on the floating-point coordinates and converts them into integer pixel coordinates to generate a registered depth map that is spatially aligned with the color map; and outputs the registered depth map to the target detection module;

[0123] Step 4: The object detection module processes the registered depth map from the depth alignment module using a model based on the Transformer architecture. Preprocessing: resize the image to 800×600 pixels, normalize the pixel values, and convert it to BLOB format. Model processing: extract feature maps through the CNN backbone network, add position encoding, enhance global features through the Transformer encoder based on the multi-head self-attention mechanism, interact with the object query through the decoder, and output the bounding box and category prediction tensor. Post-processing: filter prediction results with a confidence level ≥ 0.5, identify the target category of the bridge structure component, parse and record the pixel coordinates of the center point of the bounding box, and output them to the spatial topology relationship modeling module.

[0124] Step 5: The spatial topology relationship modeling module constructs a spatial index based on the target detection results. If there are two targets, an ordered index is constructed according to the horizontal coordinates. If there is a single target, the current horizontal coordinate is compared with the historical benchmark. If the offset is less than the dynamic tolerance threshold, it is determined to be the left topology target and the right target identifier is set to zero. If the offset is not less than the dynamic tolerance threshold, it is determined to be the right topology target and the left target identifier is set to zero. The constructed spatial index is output to the 3D coordinate calculation module.

[0125] Step 6: The 3D coordinate calculation module calculates 3D coordinates based on the distinguished left and right target coordinates, extracts depth statistics of an N×N region with the target center, where N≥6; converts the depth data to a 3D world coordinate system through back-projection; calculates the vibration offset based on the IMU data, compensates the vibration offset with the coordinates, and smoothes the trajectory using an IIR filter with adaptive filter coefficients; and outputs the stabilized coordinates to the distance calculation module.

[0126] Step 7: The distance calculation module determines the straight-line distance between the two target points based on the compensated 3D coordinates using the Euclidean distance decomposition algorithm, and calculates the xy plane distance and coordinate axis components of the 3D world coordinate system using the depth difference; the calculated data is output to the calibration optimization module;

[0127] Step 8: The calibration optimization module uses a finite state machine to determine the spatial consistency of consecutive frames and trigger the tracking benchmark calibration; performs trajectory optimization, initializes the sliding window queue to store historical positions; removes the oldest data and adds the current frame offset; calculates the sliding mean; and outputs the smoothed data to the spatiotemporal fusion module;

[0128] Step 9: In the spatiotemporal fusion module, the spatial transformation information and the temporal sequence information are fused to generate 3D spatiotemporal information for control or decision-making; the 3D spatiotemporal information is output to the visualization module and the storage module, and the dynamic spatial data set including the number of targets, spatial position coordinates, real-time offset vectors, and three-dimensional distance matrix is ​​stored in the remote terminal;

[0129] Step 10: In the labeling module, a measurement bounding box is drawn in the corresponding detection area of ​​the real-time video stream image, and the spatial coordinate information and depth measurement results of the target are marked in the detection area.

[0130] Compared with the existing technology, the present invention has the following significant advantages: (1) It can be used to detect key structural components in real time during the bridge erection process and generate 4D data sets for continuous detection and tracking, thereby improving the training accuracy of the operation status monitoring and safety warning model of the bridge erection process and solving the time series discontinuity problem of traditional methods in dynamic construction scenarios; (2) It adopts an algorithm based on the Transformer model, which can achieve high-speed and high-precision target detection and tracking; (3) It collects acceleration and gyroscope data through sensors, compensates and eliminates mechanical vibration, thereby improving precision and accuracy; (4) It improves the safety of the bridge erection process construction, which has important social and economic significance for the safety monitoring of the bridge erection process. BRIEF DESCRIPTION OF THE DRAWINGS

[0131] Figure 1 It is a schematic diagram of the image acquisition and processing unit of the dynamic spatial data set construction system for bridge structural members of the present invention.

[0132] Figure 2 It is a schematic diagram of a remote analysis terminal of a dynamic spatial data set construction system for bridge structural members of the present invention.

[0133] Figure 3 This is a schematic diagram of the installation position of the image acquisition and processing unit in the present invention. DETAILED DESCRIPTION

[0134] The present invention provides a system and method for constructing a dynamic spatial dataset for bridge structural components. The system includes an image acquisition and processing module, a depth alignment module, a target detection module, a spatial topology relationship modeling module, a 3D coordinate calculation module, a distance calculation module, a calibration optimization module, a spatiotemporal fusion module, and a labeling module. The method comprises the following steps: first, a synchronized original depth map and color image sequence of the bridge construction scene are acquired, and the timestamp of each frame is recorded; then, depth alignment and registration are performed on the original depth map and color image, real-time target detection is performed, and key structural components are identified; spatial topology relationship modeling is performed, 3D coordinates are calculated, vibration offset is calculated based on IMU acceleration and gyroscope data, mechanical vibration compensation is performed, and compensated 3D coordinates are output; then, based on the compensated coordinates, the Euclidean distance, xy plane distance, and components between two points are calculated; a finite state machine is used to determine the spatial consistency of consecutive frames, trigger tracking benchmark calibration, and perform trajectory calibration and optimization; a spatiotemporal fusion data sequence is constructed, and the timestamp and spatial transformation matrix of each point cloud frame are associated; finally, dynamic construction labels are annotated to generate a time-series point cloud dataset with a timestamp.

[0135] The present invention provides a dynamic spatial data set construction system and method that can continuously detect and track the bridge erection process, improve the training accuracy of the bridge erection machine operation status monitoring and safety warning model, and solve the time series discontinuity problem of traditional methods in dynamic construction scenarios.

[0136] Combine Figures 1 to 3 The present invention provides a dynamic spatial data set construction system for bridge structural components, characterized in that it includes an image acquisition and processing unit and a remote analysis terminal, wherein the remote analysis terminal is provided with a depth alignment module, a target detection module, a spatial topology relationship modeling module, a 3D coordinate calculation module, a distance calculation module, a calibration optimization module, a spatiotemporal fusion module, and a label annotation module;

[0137] The image acquisition and processing unit is used to acquire image data and then send it to the remote analysis terminal via wireless routing;

[0138] The depth alignment module is used for depth alignment data registration;

[0139] The target detection module includes a preprocessing submodule and a model processing submodule; the preprocessing submodule receives the registered depth map processed by the depth alignment module and converts it into the format required by the detection model; the model processing submodule uses a detection model based on the Transformer architecture to process the received image stream, perform real-time target detection, and obtain the target's bounding box and category information, as well as the coordinates of the center point of the bounding box;

[0140] The spatial topology relationship modeling module is used to build a spatial index according to the number of targets, determine the orientation of a single target through the offset, and reset the opposite side mark;

[0141] The 3D coordinate calculation module is used to calculate the coordinates of the target in the 3D world coordinate system and output stable coordinates after compensating for mechanical vibration;

[0142] The distance calculation module is used to calculate the straight-line distance between two target points, the distance between the target points in the plane perpendicular to the optical axis in the 3D world coordinate system, and the coordinate axis components;

[0143] The calibration optimization module is used to smooth the trajectory and output it;

[0144] The spatiotemporal fusion module is used to construct a point cloud temporal sequence and associate a spatial transformation matrix;

[0145] The label marking module is used to generate dynamic parameter labels of bridge structural members during the bridge construction process, including beam center coordinates, depth information, and relative distances between targets.

[0146] As a specific example, the image acquisition and processing unit includes a pan / tilt platform, a mounting component, a visual acquisition module, and a wireless routing module, wherein:

[0147] The pan-tilt platform is in the shape of a cuboid;

[0148] The mounting assembly is fixed to the top surface of the pan / tilt head;

[0149] The visual acquisition module includes a binocular camera and an inertial measurement unit, which realizes the vertical displacement freedom through the installation component and fixes the spatial posture through a locking mechanism; the binocular camera collects the original depth map and color map,

[0150] The wireless routing module is used to send the collected image data and the inertial data output by the inertial measurement unit to the remote analysis terminal;

[0151] The image acquisition and processing unit is installed upside down below the bridge-building machine and fixed at the mid-span position of the main beam of the bridge-building machine, so that the pan-tilt head faces upward and the binocular camera faces downward, and the optical axis of the binocular camera is set perpendicular to the bridge deck.

[0152] As a specific example, the binocular camera is a binocular imaging unit based on Intel RealSense D435i, which outputs synchronized video stream data with a resolution of 1280×720 pixels and a sampling frequency of 30Hz; the inertial measurement unit uses a BMI085 IMU sensor; the carrying housing and mounting assembly of the visual acquisition module form a rigid frame, so that the inertial measurement unit and the binocular camera form a co-base rigid coupling structure, and the motion state of the inertial measurement unit and the binocular camera are synchronized;

[0153] The wireless routing module is coupled to the data output interface of the image acquisition and processing unit to establish an encrypted communication channel between the image acquisition and processing unit and the remote analysis terminal; the communication channel achieves real-time synchronous transmission of image data from the visual acquisition module and inertial data output by the inertial measurement unit to the remote analysis terminal and cloud storage with a transmission delay of less than 50ms.

[0154] As a specific example, the depth alignment module receives the original depth map and color map from the binocular camera, and the depth camera coordinate system and color camera coordinate system , respectively, with the depth camera optical center and color camera optical center is the origin, 、 Axis along the optical axis, 、 The axis is parallel to the horizontal direction of the imaging plane, 、 The axis is parallel to the vertical direction of the imaging plane, Plane and Axis vertical, Plane and The axes are vertical, representing the ideal position of the imaging plane; the depth pixel coordinate system and the color pixel coordinate system are 2D coordinates on the image plane. The functions of the depth alignment module are:

[0155] Iterates through each pixel to determine whether the depth value is valid, and then establishes a mapping relationship between the depth camera coordinate system and the color camera coordinate system through a multi-level coordinate space transformation chain; performs pixel-level reprojection based on this mapping relationship to generate a registered depth map that is aligned with the color image space, and outputs the registered depth map to the object detection module;

[0156] The multi-level coordinate space transformation chain, the first level is the conversion from depth pixels to depth camera 3D points, the formula is as follows:

[0157]

[0158] in, is the pixel coordinate of the depth map, is the depth value, is the focal length of the depth camera, is the principal point of the depth camera, The 3D point coordinates in the depth camera coordinate system after conversion;

[0159] The second level is the conversion from depth camera 3D points to color camera 3D points. The formula is as follows:

[0160]

[0161] in, 3 The depth of 3 is converted to a color rotation matrix, 3 1 translation vector, is the 3D point coordinate in the converted color camera coordinate system;

[0162] The third level is the conversion of the color camera 3D points to color pixels. The formula is as follows:

[0163]

[0164] in, is the focal length of the color camera, is the principal point of the color camera, is the pixel coordinate projected onto the color image plane;

[0165] The fourth level is boundary processing and assignment: the floating-point coordinates obtained by projection are converted into integer pixel coordinates, where any pixel coordinate value is set to 0 if it is less than 0, and is set to the maximum value if it exceeds the maximum value of the color pixel coordinate.

[0166] As a specific example, the target detection module has a specific processing flow as follows:

[0167] The preprocessing submodule is used to convert the input color image into the format required by the detection model, including: adjusting the image size to 800*600 pixels, normalizing the pixel values, converting the image data into the BLOB format required for detection model input, and inputting the BLOB format data into the model processing module;

[0168] The model processing submodule includes the following processing steps ① to ③:

[0169] ① Feature extraction, which is used to extract feature maps from the preprocessed input image through the CNN backbone network;

[0170] ② Transformer processing, which is used to perform the following processing on the feature map in sequence:

[0171] (a) Adding position coding to the feature map: The position coding is generated using sine / cosine functions to encode spatial position information. The formula is as follows:

[0172]

[0173]

[0174] in Represents the position index, Represents the dimension index, represents the feature dimension of the detection model, is the even-dimensional component of the positional encoding, It is the odd-dimensional component of position coding; the position coding dimension is 256;

[0175] (b) Through the Transformer encoder, a multi-head self-attention mechanism is used to calculate the correlation weights between pixels in the feature map to extract and enhance the global feature representation;

[0176] The multi-head self-attention mechanism is as follows:

[0177] Assume that the input of the head attention layer is , firstly transform the input sequence into Mapping to query matrix , key matrix Sum Matrix , the formula is as follows:

[0178]

[0179]

[0180]

[0181] in, is the query weight matrix, is the key weight matrix, is the value weight matrix;

[0182] right 、 and The attention weights are obtained by performing a scaled dot product operation. This process is performed independently in each layer and the calculation process is as follows:

[0183]

[0184] in, is the bond matrix Dimensions, is a normalized exponential function, which normalizes the self-attention score; Compute the output matrix for attention;

[0185] The attention scores of each layer will be spliced ​​in the channel dimension. The result of the layer is the final output of the multi-head attention layer, and the formula is as follows:

[0186]

[0187]

[0188] in, is a linear mapping matrix used to transform The attention scores of the layers are concatenated into a whole; For the The output of the attention head is is the attention head index, is the total number of attention heads, For the The query matrix of each head, For the The key matrix of the head, For the The value matrix of each head, is the final output of the multi-head attention layer;

[0189] (c) Through the Transformer decoder, the object query is used to interact with the enhanced features output by the encoder;

[0190] (d) Through the prediction head, based on the interaction results, a tensor containing bounding box predictions and category predictions is output;

[0191] (e) Outputting the tensor containing the prediction results to the post-processing module;

[0192] ③ Post-processing module, used to parse the input tensor to obtain the detection results, including:

[0193] (a) Extract confidence scores for each prediction result;

[0194] (b) If the confidence score is lower than the preset target confidence threshold of 0.5, the corresponding prediction result is discarded;

[0195] (c) for a prediction result whose confidence score satisfies the threshold condition, determining the category ID corresponding to the maximum category score of the prediction result;

[0196] (d) determining whether the target category is a bridge deck or a side of a bridge based on the category ID;

[0197] (e) If the target category is a bridge deck or bridge side, the corresponding normalized bounding box parameters are parsed, including the center point coordinates and width and height, and the normalized bounding box parameters are multiplied by the original image size to obtain the bounding box in actual pixel coordinates;

[0198] (f) recording the center point coordinates of the bounding box in the actual pixel coordinates;

[0199] The target detection module ultimately outputs the bounding box and category information of the bridge deck or bridge side that meets the conditions, as well as the recorded center point coordinates, which are used to construct a dynamic spatial dataset.

[0200] As a specific example, the spatial topology relationship modeling module receives the information output by the target detection module and performs the following processing:

[0201] When there are two targets, an ordered spatial index is constructed based on the horizontal coordinates;

[0202] When there is a single target, compare the offset of the current target horizontal coordinates with the historical benchmark: if the offset is less than the dynamic tolerance threshold, it is determined to be the left topological target and the right target identifier is set to zero; if the offset is not less than the dynamic tolerance threshold, it is determined to be the right topological target and the left target identifier is set to zero.

[0203] As a specific example, the 3D coordinate calculation module includes a coordinate calculation submodule and a mechanical vibration compensation submodule, as follows:

[0204] The coordinate calculation submodule is set as:

[0205] Taking the target center as the reference, extract the depth statistics of the N×N pixel area, where N is a positive integer and N≥6;

[0206] Use back projection to convert the depth data of the registered depth map into the 3D world coordinate system. The formula is as follows:

[0207]

[0208] in, is the pixel coordinate of the depth map, Is the 3D world coordinate system The 3D point coordinates under is the focal length of the depth camera, is the principal point of the depth camera, is the physical depth value;

[0209] The mechanical vibration compensation submodule is set as:

[0210] Receive acceleration and gyroscope data from the inertial measurement unit (IMU). The IMU coordinate system and the color camera coordinate system are related through rigid body transformation to calculate the vibration offset. The formula is:

[0211]

[0212] in, and is the IMU sensor coordinate system 、 Axial acceleration, and is the color camera coordinate system 、 axis angular velocity, is the IMU sensor coordinate system Axis acceleration sensitivity, is the IMU sensor coordinate system Axis rotation sensitivity, is the IMU sensor coordinate system Axis acceleration sensitivity, is the IMU sensor coordinate system Axis rotation sensitivity, is the color camera coordinate system Axis direction compensation offset, is the color camera coordinate system Axis direction compensation offset;

[0213] The vibration offset and the coordinates processed by the coordinate calculation submodule are integrated to perform position compensation. The formula is:

[0214]

[0215] in, is the original 3D world coordinate system Axis coordinates, is the original 3D world coordinate system Axis coordinates, and is the compensated 3D world coordinate system axis, axis coordinates;

[0216] Apply an IIR filter to the compensated position for filtering, the formula is:

[0217]

[0218] in, The coordinates after compensation for the current frame, is the filter position of the previous frame, is the current frame filtering position;

[0219] is the filter coefficient, which is determined using the adaptive filter coefficient formula:

[0220]

[0221] in, is the modulus of the acceleration vector;

[0222] Finally, the 3D coordinate calculation module outputs the stable coordinates to the distance calculation module.

[0223] As a specific example, the distance calculation module is configured as follows:

[0224] The straight-line distance between two target points is determined by the Euclidean distance decomposition algorithm. The formula is:

[0225]

[0226] in Represents the straight-line distance between the two target points. The origin of the 3D world coordinate system is the optical center of the camera, the z-axis is along the optical axis, and the coordinates of the 3D coordinate system of the two target points are and ;

[0227] The 3D world coordinate system xy plane distance and coordinate axis components are calculated by depth difference. The formula is:

[0228]

[0229]

[0230]

[0231]

[0232] in, is the depth difference, Is the 3D world coordinate system Axis pixel difference, Is the 3D world coordinate system Axis pixel difference, is the distance between the two target points on the plane perpendicular to the optical axis, The target point is in the horizontal direction, that is, the 3D world coordinate system The relative displacement of the axis, The target point is in the vertical direction, that is, the 3D world coordinate system The relative displacement of the axis, is the pixel plane distance.

[0233] As a specific example, the calibration optimization module includes a state calibration finite state machine, a trajectory optimizer, and a multimodal output interface, wherein:

[0234] (1) The state calibration finite state machine receives the coordinates output by the 3D coordinate calculation unit and triggers the tracking reference calibration only when the detection results of consecutive frames all meet the spatial consistency threshold, specifically:

[0235] According to the number of detected targets, the detection is divided into three situations: no target, single target, and dual targets:

[0236] If there is no target, the global coordinates are reset to zero;

[0237] If it is a single target, when the target is close to the left, it is marked as the left target, and setting the right target is invalid; when the target is close to the right, it is marked as the right target, and setting the left target is invalid; keep the y-axis coordinate of the 3D world coordinate system consistent;

[0238] If there are two targets, the forced sorting ensures that the 3D world coordinate system on the left If the value target is smaller than the right side, mark the target with the smaller x coordinate in the 3D world coordinate system as the left target, generate and output an ordered target pair;

[0239] Calculate target distance based on sorted coordinates, verify the validity of coordinate data, and output target location and quantity information;

[0240] (2) The trajectory optimizer receives the calibration signal output by the state calibration finite state machine, builds a timing cache queue to store the target historical position, and calculates the sliding mean based on the window length, as follows:

[0241] Read the sliding window length parameter , initialize two first-in-first-out queues and , respectively used to store the target in the 3D world coordinate system Axis and Historical position data in the axis direction;

[0242] Get the current relative position offset output by the distance calculation module frame by frame , remove the oldest historical coordinate value of the queue and add the relative position offset of the current frame to update the queue;

[0243] The updated queue is traversed and accumulated, and a sliding average is calculated. The calculation formula is:

[0244]

[0245] in, is the smoothed 3D world coordinate system Axis coordinates, is the smoothed 3D world coordinate system Axis coordinates, is the sliding window length, For the historical 3D world coordinate system Axis coordinate offset queue value, For the historical 3D world coordinate system Axis coordinate offset queue value;

[0246] (3) The multimodal output interface is coupled to the spatiotemporal fusion module, and outputs the number of targets, spatial position coordinates, real-time spatial offset vectors, and a three-dimensional distance matrix between targets;

[0247] The spatiotemporal fusion module fuses the spatial transformation and the temporal sequence to generate 3D spatiotemporal information for control or decision-making, and outputs it to the visualization module and the storage module for storage in the remote terminal;

[0248] The label marking module receives the target position data output by the target detection module, draws a measurement bounding box in the corresponding detection area in the real-time video stream image, and marks the target's spatial coordinate information and depth measurement results in the detection area.

[0249] The present invention also provides a method for constructing a dynamic spatial data set for a bridge erection control system. The method is based on the dynamic spatial data set construction system for a bridge erection control system, and the method comprises the following steps:

[0250] Step 1: Install the entire hardware module upside down below the bridge crane, with the gimbal facing up and the binocular camera facing down, ensuring that the camera optical axis is perpendicular to the beam below. Continuously capture synchronized depth and color frames of the bridge construction scene at a preset 30Hz sampling frequency, and record the timestamp of each frame. Acceleration and angular velocity data are collected using the inertial measurement unit.

[0251] Step 2: Establish an encrypted communication channel through the wireless routing module to send the image and inertial data to the remote analysis terminal in real time with a transmission delay of less than 50ms;

[0252] Step 3: The depth alignment module of the remote analysis terminal receives the original depth map and color map output by the image acquisition and processing unit, traverses the depth map pixels, and filters the valid depth values; establishes a mapping relationship between the depth camera coordinate system and the color camera coordinate system through a multi-level coordinate space transformation chain, including: depth pixels to depth camera 3D points, depth camera 3D points to color camera 3D points, and color camera 3D points to color pixels; performs boundary processing on the floating-point coordinates and converts them into integer pixel coordinates to generate a registered depth map that is spatially aligned with the color map; and outputs the registered depth map to the target detection module;

[0253] Step 4: The object detection module processes the registered depth map from the depth alignment module using a model based on the Transformer architecture. Preprocessing: resize the image to 800×600 pixels, normalize the pixel values, and convert it to BLOB format. Model processing: extract feature maps through the CNN backbone network, add position encoding, enhance global features through the Transformer encoder based on the multi-head self-attention mechanism, interact with the object query through the decoder, and output the bounding box and category prediction tensor. Post-processing: filter prediction results with a confidence level ≥ 0.5, identify the target category of the bridge structure component, parse and record the pixel coordinates of the center point of the bounding box, and output them to the spatial topology relationship modeling module.

[0254] Step 5: The spatial topology relationship modeling module constructs a spatial index based on the target detection results. If there are two targets, an ordered index is constructed according to the horizontal coordinates. If there is a single target, the current horizontal coordinate is compared with the historical benchmark. If the offset is less than the dynamic tolerance threshold, it is determined to be the left topology target and the right target identifier is set to zero. If the offset is not less than the dynamic tolerance threshold, it is determined to be the right topology target and the left target identifier is set to zero. The constructed spatial index is output to the 3D coordinate calculation module.

[0255] Step 6: The 3D coordinate calculation module calculates 3D coordinates based on the distinguished left and right target coordinates, extracts depth statistics of an N×N region with the target center, where N≥6; converts the depth data to a 3D world coordinate system through back-projection; calculates the vibration offset based on the IMU data, compensates the vibration offset with the coordinates, and smoothes the trajectory using an IIR filter with adaptive filter coefficients; and outputs the stabilized coordinates to the distance calculation module.

[0256] Step 7: The distance calculation module determines the straight-line distance between the two target points based on the compensated 3D coordinates using the Euclidean distance decomposition algorithm, and calculates the xy plane distance and coordinate axis components of the 3D world coordinate system using the depth difference; the calculated data is output to the calibration optimization module;

[0257] Step 8: The calibration optimization module uses a finite state machine to determine the spatial consistency of consecutive frames and trigger the tracking benchmark calibration; performs trajectory optimization, initializes the sliding window queue to store historical positions; removes the oldest data and adds the current frame offset; calculates the sliding mean; and outputs the smoothed data to the spatiotemporal fusion module;

[0258] Step 9: In the spatiotemporal fusion module, the spatial transformation information and the temporal sequence information are fused to generate 3D spatiotemporal information for control or decision-making; the 3D spatiotemporal information is output to the visualization module and the storage module, and the dynamic spatial data set including the number of targets, spatial position coordinates, real-time offset vectors, and three-dimensional distance matrix is ​​stored in the remote terminal;

[0259] Step 10: In the labeling module, a measurement bounding box is drawn in the corresponding detection area of ​​the real-time video stream image, and the spatial coordinate information and depth measurement results of the target are marked in the detection area.

[0260] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0261] Example 1

[0262] This embodiment provides a dynamic spatial data set construction system for a bridge erection control system, including an image acquisition and processing unit, a depth alignment module, a target detection module, a spatial topology relationship modeling module, a 3D coordinate calculation module, a distance calculation module, a calibration optimization module, a spatiotemporal fusion module, and a labeling module;

[0263] The image acquisition and processing unit is used to acquire image data and then send it to the remote analysis terminal via wireless routing;

[0264] The depth alignment module is used for depth alignment data registration;

[0265] The target detection module is used to perform real-time target detection and trajectory generation;

[0266] The spatial topology relationship modeling module is used to build a spatial index according to the number of targets, determine the orientation of a single target through the offset, and reset the opposite side mark;

[0267] The 3D coordinate calculation module is used to calculate the coordinates of the target in the three-dimensional coordinate system and output stable coordinates after compensating for mechanical vibration;

[0268] The distance calculation module is used to calculate the straight-line distance between two target points, the distance on the xy plane where the target points are located, and the coordinate axis components;

[0269] The calibration optimization module is used to smooth the trajectory and output it;

[0270] The spatiotemporal fusion module is used to construct a point cloud temporal sequence and associate a spatial transformation matrix;

[0271] The label marking module is used to generate dynamic parameter labels of bridge structural members during the bridge construction process, including beam center coordinates, depth information, and relative distances between targets.

[0272] Furthermore, the dynamic parameter labels of the construction process include beam coordinates, displacement vectors, and the rate of change of relative distances between targets.

[0273] Furthermore, the image acquisition and processing unit includes: a gimbal, which is in the shape of a rectangular parallelepiped; an adjustable mounting assembly, fixed to the top surface of the gimbal; a visual acquisition unit, including a binocular camera and an inertial measurement unit, which realizes vertical displacement freedom through the adjustable mounting assembly and fixes the spatial posture through a locking mechanism; and a wireless routing module for remote connection.

[0274] Furthermore, the hardware module is fixedly installed at the mid-span of the main beam of the bridge-building machine, and the optical axis of the binocular camera is set perpendicular to the bridge deck; the unit is configured as a binocular imaging unit based on Intel RealSense D435i, which outputs synchronous video stream data with a resolution of 1280×720 pixels and a sampling frequency of 30Hz, including original depth maps and color maps.

[0275] Furthermore, the inertial measurement unit adopts a BMI085 IMU sensor; the carrying shell and the adjustable mounting assembly form a rigid frame, so that the inertial measurement unit and the binocular vision acquisition unit form a common base rigid coupling structure to ensure that the motion states of the two are synchronized.

[0276] Furthermore, the wireless routing module is operatively coupled to the data output interface of the point cloud processing module; the wireless routing module establishes an encrypted communication channel between the point cloud processing module and the remote analysis terminal; the channel is configured to achieve real-time synchronous transmission and cloud storage of the video stream data of the visual acquisition unit and the inertial data of the inertial measurement unit to the remote analysis terminal with a transmission delay of less than 50ms.

[0277] Furthermore, the original depth map, color map and IMU data are transmitted to the remote analysis terminal through the wireless routing module, and the depth map pixels are traversed by the preset algorithm in the remote analysis terminal to screen the valid depth values; a mapping relationship between the depth coordinate system and the color coordinate system is established through a multi-level coordinate space transformation chain, including: depth pixels to depth camera 3D points, depth camera 3D points to color camera 3D points, color camera 3D points, to color pixels; floating-point coordinates are boundary processed and converted into integer pixel coordinates to generate a registered depth map aligned with the color map space; and the registered depth map is output to the target detection module.

[0278] Furthermore, the target detection module processes the registered depth map from the depth alignment module using a model based on the Transformer architecture; preprocessing: adjust the image size to 800×600 pixels, normalize the pixel values, convert it to BLOB format, and add position encoding; model processing: extract feature maps through the CNN backbone network, enhance global features through the Transformer encoder (multi-head self-attention mechanism), interact with the object query through the decoder, and output the bounding box and category prediction tensor; post-processing: filter prediction results with confidence ≥ 0.5, identify the target category of bridge structural components, parse and record the pixel coordinates of the center point of the bounding box; and output to the spatial topological relationship modeling module.

[0279] Furthermore, the spatial topological relationship modeling module constructs a spatial index based on the target detection results. If there are two targets, an ordered index is constructed according to the horizontal coordinates; if it is a single target, the current horizontal coordinates are compared with the historical benchmark. If the offset is less than the dynamic tolerance threshold, it is determined to be a left topological target and the right target identifier is set to zero; if the offset is greater than or equal to the dynamic tolerance threshold, it is determined to be a right topological target and the left target identifier is set to zero; and the result is output to the 3D coordinate calculation module.

[0280] Furthermore, the 3D coordinate calculation module performs 3D coordinate calculation based on the distinguished left and right target coordinates, extracts depth statistics of an N×N area (N≥6) with the target center; converts the depth data to a 3D world coordinate system through back projection; calculates the vibration offset based on the IMU data, compensates the vibration offset in conjunction with the coordinates, and smoothes the trajectory using an IIR filter with adaptive filter coefficients; and outputs the stable coordinates to the distance calculation module.

[0281] Furthermore, the distance calculation module determines the straight-line distance between two target points based on the compensated 3D coordinates through the Euclidean distance decomposition algorithm, calculates the xy plane distance and coordinate axis components through the depth difference, and outputs the calculated data to the calibration optimization module.

[0282] Furthermore, the calibration optimization module determines the spatial consistency of consecutive frames through a finite state machine, triggers tracking benchmark calibration, performs trajectory optimization, initializes a sliding window queue to store historical positions, removes the oldest data and adds the current frame offset, calculates the sliding mean, and outputs the smoothed data to the spatiotemporal fusion module.

[0283] Furthermore, the spatiotemporal fusion module fuses spatial transformation information and time sequence information to generate 3D spatiotemporal information for control or decision-making; outputs the 3D spatiotemporal information to the visualization module and the storage module, and stores the dynamic spatial data set containing the target number, spatial position coordinates, real-time offset vector and three-dimensional distance matrix in the remote terminal.

[0284] Furthermore, the label marking module draws a measurement bounding box in the corresponding detection area of ​​the real-time video stream image, and marks the spatial coordinate information and depth measurement result of the target in the detection area.

[0285] Example 2

[0286] This embodiment provides a method for constructing a dynamic spatial data set for a bridge erection control system, comprising the following steps:

[0287] Step 1: Install the entire hardware module upside down below the bridge crane, with the gimbal facing up and the binocular camera facing down, ensuring that the camera optical axis is perpendicular to the beam below. Continuously capture synchronized depth and color frames of the bridge construction scene at a preset 30Hz sampling frequency, and record the timestamp of each frame. Acceleration and angular velocity data are collected using the inertial measurement unit.

[0288] Step 2: Establish an encrypted communication channel through the wireless routing module to send the image and inertial data to the remote analysis terminal in real time with a transmission delay of less than 50ms;

[0289] Step 3: The depth alignment module of the remote analysis terminal receives the original depth map and color map output by the image acquisition and processing unit, traverses the depth map pixels, and filters the valid depth values; establishes a mapping relationship between the depth camera coordinate system and the color camera coordinate system through a multi-level coordinate space transformation chain, including: depth pixels to depth camera 3D points, depth camera 3D points to color camera 3D points, and color camera 3D points to color pixels; performs boundary processing on the floating-point coordinates and converts them into integer pixel coordinates to generate a registered depth map that is spatially aligned with the color map; and outputs the registered depth map to the target detection module;

[0290] Step 4: The object detection module processes the registered depth map from the depth alignment module using a model based on the Transformer architecture. Preprocessing: resize the image to 800×600 pixels, normalize the pixel values, and convert it to BLOB format. Model processing: extract feature maps through the CNN backbone network, add position encoding, enhance global features through the Transformer encoder based on the multi-head self-attention mechanism, interact with the object query through the decoder, and output the bounding box and category prediction tensor. Post-processing: filter prediction results with a confidence level ≥ 0.5, identify the target category of the bridge structure component, parse and record the pixel coordinates of the center point of the bounding box, and output them to the spatial topology relationship modeling module.

[0291] Step 5: The spatial topology relationship modeling module constructs a spatial index based on the target detection results. If there are two targets, an ordered index is constructed according to the horizontal coordinates. If there is a single target, the current horizontal coordinate is compared with the historical benchmark. If the offset is less than the dynamic tolerance threshold, it is determined to be the left topology target and the right target identifier is set to zero. If the offset is not less than the dynamic tolerance threshold, it is determined to be the right topology target and the left target identifier is set to zero. The constructed spatial index is output to the 3D coordinate calculation module.

[0292] Step 6: The 3D coordinate calculation module calculates 3D coordinates based on the distinguished left and right target coordinates, extracts depth statistics of an N×N region with the target center, where N≥6; converts the depth data to a 3D world coordinate system through back-projection; calculates the vibration offset based on the IMU data, compensates the vibration offset with the coordinates, and smoothes the trajectory using an IIR filter with adaptive filter coefficients; and outputs the stabilized coordinates to the distance calculation module.

[0293] Step 7: The distance calculation module determines the straight-line distance between the two target points based on the compensated 3D coordinates using the Euclidean distance decomposition algorithm, and calculates the xy plane distance and coordinate axis components of the 3D world coordinate system using the depth difference; the calculated data is output to the calibration optimization module;

[0294] Step 8: The calibration optimization module uses a finite state machine to determine the spatial consistency of consecutive frames and trigger the tracking benchmark calibration; performs trajectory optimization, initializes the sliding window queue to store historical positions; removes the oldest data and adds the current frame offset; calculates the sliding mean; and outputs the smoothed data to the spatiotemporal fusion module;

[0295] Step 9: In the spatiotemporal fusion module, the spatial transformation information and the temporal sequence information are fused to generate 3D spatiotemporal information for control or decision-making; the 3D spatiotemporal information is output to the visualization module and the storage module, and the dynamic spatial data set including the number of targets, spatial position coordinates, real-time offset vectors, and three-dimensional distance matrix is ​​stored in the remote terminal;

[0296] Step 10: In the labeling module, a measurement bounding box is drawn in the corresponding detection area of ​​the real-time video stream image, and the spatial coordinate information and depth measurement results of the target are marked in the detection area.

[0297] This invention can provide a basis for the training and verification of machine learning models. Through real-time processing and analysis of data streams, the platform can intuitively detect the current status of key structural components of the bridge during the construction process, providing on-site engineers with immediate and comprehensive structural health status awareness, greatly improving the safety of the construction process, preventing accidents, and realizing unmanned and rapid operations.

[0298] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A dynamic spatial data set construction system for bridge structural components, characterized in that: It includes an image acquisition and processing unit and a remote analysis terminal, wherein the remote analysis terminal is provided with a depth alignment module, a target detection module, a spatial topology relationship modeling module, a 3D coordinate calculation module, a distance calculation module, a calibration optimization module, a spatiotemporal fusion module, and a labeling module; The image acquisition and processing unit is used to acquire image data and then send it to the remote analysis terminal via wireless routing; The depth alignment module is used for depth alignment data registration; The target detection module includes a pre-processing sub-module and a model processing sub-module; The preprocessing submodule receives the registered depth map processed by the depth alignment module and converts it into the format required by the detection model; The model processing submodule uses a detection model based on the Transformer architecture to process the received image stream and perform real-time object detection to obtain the bounding box and category information of the target, as well as the coordinates of the center point of the bounding box; The spatial topology relationship modeling module is used to build a spatial index according to the number of targets, determine the orientation of a single target through the offset, and reset the opposite side mark; The 3D coordinate calculation module is used to calculate the coordinates of the target in the 3D world coordinate system and output stable coordinates after compensating for mechanical vibration; The distance calculation module is used to calculate the straight-line distance between two target points, the distance between the target points in the plane perpendicular to the optical axis in the 3D world coordinate system, and the coordinate axis components; The calibration optimization module is used to smooth the trajectory and output it; The spatiotemporal fusion module is used to construct a point cloud temporal sequence and associate a spatial transformation matrix; The label marking module is used to generate dynamic parameter labels of bridge structural members during the bridge construction process, including beam center coordinates, depth information, and relative distances between targets.

2. The dynamic spatial dataset construction system for bridge structural components according to claim 1, characterized in that: The image acquisition and processing unit includes a pan / tilt platform, a mounting assembly, a visual acquisition module, and a wireless routing module, wherein: The pan-tilt platform is in the shape of a cuboid; The mounting assembly is fixed to the top surface of the pan / tilt head; The visual acquisition module includes a binocular camera and an inertial measurement unit, which realizes the vertical displacement freedom through the installation component and fixes the spatial posture through a locking mechanism; the binocular camera collects the original depth map and color map, The wireless routing module is used to send the collected image data and the inertial data output by the inertial measurement unit to the remote analysis terminal; The image acquisition and processing unit is installed upside down below the bridge-building machine and fixed at the mid-span position of the main beam of the bridge-building machine, so that the pan-tilt head faces upward and the binocular camera faces downward, and the optical axis of the binocular camera is set perpendicular to the bridge deck.

3. The dynamic spatial dataset construction system for bridge structural components according to claim 2, characterized in that: The binocular camera is a binocular imaging unit based on Intel RealSense D435i, which outputs synchronized video stream data with a resolution of 1280×720 pixels and a sampling frequency of 30Hz. The inertial measurement unit uses a BMI085 IMU sensor. The carrying housing and mounting assembly of the visual acquisition module form a rigid frame, so that the inertial measurement unit and the binocular camera form a co-base rigid coupling structure, and the motion state of the inertial measurement unit and the binocular camera are synchronized. The wireless routing module is coupled to the data output interface of the image acquisition and processing unit to establish an encrypted communication channel between the image acquisition and processing unit and the remote analysis terminal; the communication channel achieves real-time synchronous transmission of image data from the visual acquisition module and inertial data output by the inertial measurement unit to the remote analysis terminal and cloud storage with a transmission delay of less than 50ms.

4. The dynamic spatial data set construction system for bridge structural components according to claim 3, characterized in that: The depth alignment module receives the original depth map and color map from the binocular camera. The depth pixel coordinate system and the color pixel coordinate system are 2D coordinates on the image plane. The functions of the depth alignment module are: Iterates through each pixel to determine whether the depth value is valid, and then establishes a mapping relationship between the depth camera coordinate system and the color camera coordinate system through a multi-level coordinate space transformation chain; performs pixel-level reprojection based on this mapping relationship to generate a registered depth map that is aligned with the color image space, and outputs the registered depth map to the object detection module; The multi-level coordinate space transformation chain, the first level is the conversion from depth pixels to depth camera 3D points, the formula is as follows: ; in, is the pixel coordinate of the depth map, is the depth value, is the focal length of the depth camera, is the principal point of the depth camera, The 3D point coordinates in the depth camera coordinate system after conversion; The second level is the conversion from depth camera 3D points to color camera 3D points. The formula is as follows: ; in, 3 The depth of 3 is converted to a color rotation matrix, 3 1 translation vector, is the 3D point coordinate in the converted color camera coordinate system; The third level is the conversion of the color camera 3D points to color pixels. The formula is as follows: ; in, is the focal length of the color camera, is the principal point of the color camera, is the pixel coordinate projected onto the color image plane; The fourth level is boundary processing and assignment: the floating-point coordinates obtained by projection are converted into integer pixel coordinates, where any pixel coordinate value is set to 0 if it is less than 0, and is set to the maximum value if it exceeds the maximum value of the color pixel coordinate.

5. The dynamic spatial data set construction system for bridge structural components according to claim 4, characterized in that: The target detection module has the following specific processing flow: The preprocessing submodule is used to convert the input color image into the format required by the detection model, including: adjusting the image size to 800*600 pixels, normalizing the pixel values, converting the image data into the BLOB format required for detection model input, and inputting the BLOB format data into the model processing module; The model processing submodule includes the following processing steps ① to ③: ① Feature extraction, which is used to extract feature maps from the preprocessed input image through the CNN backbone network; ② Transformer processing, which is used to perform the following processing on the feature map in sequence: (a) Adding position coding to the feature map: The position coding is generated using sine / cosine functions to encode spatial position information. The formula is as follows: ; ; in Represents the position index, Represents the dimension index, represents the feature dimension of the detection model, is the even-dimensional component of the positional encoding, It is the odd-dimensional component of position coding; the position coding dimension is 256; (b) Through the Transformer encoder, a multi-head self-attention mechanism is used to calculate the correlation weights between pixels in the feature map to extract and enhance the global feature representation; The multi-head self-attention mechanism is as follows: Assume that the input of the head attention layer is , firstly transform the input sequence into Mapping to query matrix , key matrix Sum Matrix , the formula is as follows: ; ; ; in, is the query weight matrix, is the key weight matrix, is the value weight matrix; right 、 and The attention weights are obtained by performing a scaled dot product operation. This process is performed independently in each layer and the calculation process is as follows: ; in, is the bond matrix Dimensions, is a normalized exponential function, which normalizes the self-attention score; Compute the output matrix for attention; The attention scores of each layer will be spliced ​​in the channel dimension. The result of the layer is the final output of the multi-head attention layer, and the formula is as follows: ; ; in, is a linear mapping matrix used to transform The attention scores of the layers are concatenated into a whole; For the The output of the attention head is is the attention head index, is the total number of attention heads, For the The query matrix of each head, For the The key matrix of the head, For the The value matrix of each head, is the final output of the multi-head attention layer; (c) Through the Transformer decoder, the object query is used to interact with the enhanced features output by the encoder; (d) Through the prediction head, based on the interaction results, a tensor containing bounding box predictions and category predictions is output; (e) Outputting the tensor containing the prediction results to the post-processing module; ③ Post-processing module, used to parse the input tensor to obtain the detection results, including: (a) Extract confidence scores for each prediction result; (b) Setting the target confidence threshold to 0.5, if the confidence score is lower than the preset target confidence threshold of 0.5, the corresponding prediction result is discarded; (c) for a prediction result whose confidence score satisfies the threshold condition, determining the category ID corresponding to the maximum category score of the prediction result; (d) determining whether the target category is a bridge deck or a side of a bridge based on the category ID; (e) If the target category is a bridge deck or bridge side, the corresponding normalized bounding box parameters are parsed, including the center point coordinates and width and height, and the normalized bounding box parameters are multiplied by the original image size to obtain the bounding box in actual pixel coordinates; (f) recording the center point coordinates of the bounding box in the actual pixel coordinates; The target detection module ultimately outputs the bounding box and category information of the bridge deck or bridge side that meets the conditions, as well as the recorded center point coordinates, which are used to construct a dynamic spatial dataset.

6. The dynamic spatial data set construction system for bridge structural components according to claim 5, characterized in that: The spatial topology relationship modeling module receives the information output by the target detection module and performs the following processing: When there are two targets, an ordered spatial index is constructed based on the horizontal coordinates; When there is a single target, compare the offset of the current target horizontal coordinates with the historical benchmark: if the offset is less than the dynamic tolerance threshold, it is determined to be the left topological target and the right target identifier is set to zero; if the offset is not less than the dynamic tolerance threshold, it is determined to be the right topological target and the left target identifier is set to zero.

7. The dynamic spatial data set construction system for bridge structural components according to claim 6, characterized in that: The 3D coordinate calculation module includes a coordinate calculation submodule and a mechanical vibration compensation submodule, as follows: The coordinate calculation submodule is set as: Taking the target center as the reference, extract the depth statistics of the N×N pixel area, where N is a positive integer and N≥6; Use back projection to convert the depth data of the registered depth map into the 3D world coordinate system. The formula is as follows: ; in, is the pixel coordinate of the depth map, Is the 3D world coordinate system The 3D point coordinates under is the focal length of the depth camera, is the principal point of the depth camera, is the physical depth value; The mechanical vibration compensation submodule is set as: Receive acceleration and gyroscope data from the inertial measurement unit (IMU). The IMU coordinate system and the color camera coordinate system are related through rigid body transformation to calculate the vibration offset. The formula is: ; in, and is the IMU sensor coordinate system 、 Axial acceleration, and is the color camera coordinate system 、 axis angular velocity, is the IMU sensor coordinate system Axis acceleration sensitivity, is the IMU sensor coordinate system Axis rotation sensitivity, is the IMU sensor coordinate system Axis acceleration sensitivity, is the IMU sensor coordinate system Axis rotation sensitivity, is the color camera coordinate system Axis direction compensation offset, is the color camera coordinate system Axis direction compensation offset; The vibration offset and the coordinates processed by the coordinate calculation submodule are integrated to perform position compensation. The formula is: ; in, is the original 3D world coordinate system Axis coordinates, is the original 3D world coordinate system Axis coordinates, and is the compensated 3D world coordinate system axis, axis coordinates; Apply an IIR filter to the compensated position for filtering, the formula is: ; in, The coordinates after compensation for the current frame, is the filter position of the previous frame, is the current frame filtering position; is the filter coefficient, which is determined using the adaptive filter coefficient formula: ; in, is the modulus of the acceleration vector; Finally, the 3D coordinate calculation module outputs the stable coordinates to the distance calculation module.

8. The dynamic spatial data set construction system for bridge structural components according to claim 7, characterized in that: The distance calculation module is configured as follows: The straight-line distance between two target points is determined by the Euclidean distance decomposition algorithm. The formula is: ; in Represents the straight-line distance between two target points. The origin of the 3D world coordinate system is the optical center of the camera. The axis is along the optical axis, and the coordinates of the 3D coordinate system of the two target points are and ; The 3D world coordinate system xy plane distance and coordinate axis components are calculated by depth difference. The formula is: ; ; ; ; in, is the depth difference, Is the 3D world coordinate system Axis pixel difference, Is the 3D world coordinate system Axis pixel difference, is the distance between the two target points on the plane perpendicular to the optical axis, The target point is in the horizontal direction, that is, the 3D world coordinate system The relative displacement of the axis, The target point is in the vertical direction, that is, the 3D world coordinate system The relative displacement of the axis, is the pixel plane distance.

9. The dynamic spatial data set construction system for bridge structural components according to claim 8, characterized in that: The calibration optimization module includes a state calibration finite state machine, a trajectory optimizer, and a multimodal output interface, wherein: (1) The state calibration finite state machine receives the coordinates output by the 3D coordinate calculation unit and triggers the tracking reference calibration only when the detection results of consecutive frames all meet the spatial consistency threshold, specifically: According to the number of detected targets, the detection is divided into three situations: no target, single target, and dual targets: If there is no target, the global coordinates are reset to zero; If it is a single target, when the target is close to the left, it is marked as the left target, and setting the right target is invalid; when the target is close to the right, it is marked as the right target, and setting the left target is invalid; maintain the 3D world coordinate system The axis coordinates are consistent; If there are two targets, force the sorting to ensure the 3D world coordinate system on the left If the value target is smaller than the right side, mark the target with the smaller x coordinate in the 3D world coordinate system as the left target, generate and output an ordered target pair; Calculate the target distance based on the sorted coordinates, verify the validity of the coordinate data, and output the target position and quantity information; (2) The trajectory optimizer receives the calibration signal output by the state calibration finite state machine, builds a timing cache queue to store the target historical position, and calculates the sliding mean based on the window length, as follows: Read the sliding window length parameter N and initialize two first-in-first-out queues qx and qy, which are used to store the target in the 3D world coordinate system. Axis and Historical position data in the axis direction; Get the current relative position offset output by the distance calculation module frame by frame , remove the oldest historical coordinate value of the queue and add the relative position offset of the current frame to update the queue; The updated queue is traversed and accumulated, and a sliding average is calculated. The calculation formula is: ; in, is the smoothed 3D world coordinate system Axis coordinates, is the smoothed 3D world coordinate system Axis coordinates, N is the sliding window length, For the historical 3D world coordinate system Coordinate offset queue value, For the historical 3D world coordinate system Coordinate offset queue value; (3) The multimodal output interface is coupled to the spatiotemporal fusion module, and outputs the number of targets, spatial position coordinates, real-time spatial offset vectors, and a three-dimensional distance matrix between targets; The spatiotemporal fusion module fuses the spatial transformation and the temporal sequence to generate 3D spatiotemporal information for control or decision-making, and outputs it to the visualization module and the storage module for storage in the remote terminal; The label marking module receives the target position data output by the target detection module, draws a measurement bounding box in the corresponding detection area in the real-time video stream image, and marks the target's spatial coordinate information and depth measurement results in the detection area.

10. A method for constructing a dynamic spatial data set of a bridge structure, characterized in that: The method is based on the dynamic spatial data set construction system for bridge structural members according to any one of claims 1 to 9, and comprises the following steps: Step 1: Install the entire hardware module upside down below the bridge crane, with the gimbal facing up and the binocular camera facing down, ensuring that the camera optical axis is perpendicular to the beam below. Continuously capture synchronized depth and color frames of the bridge construction scene at a preset 30Hz sampling frequency, and record the timestamp of each frame. Acceleration and angular velocity data are collected using the inertial measurement unit. Step 2: Establish an encrypted communication channel through the wireless routing module to send the image and inertial data to the remote analysis terminal in real time with a transmission delay of less than 50ms; Step 3: The depth alignment module of the remote analysis terminal receives the original depth map and color map output by the image acquisition and processing unit, traverses the depth map pixels, and filters the valid depth values; establishes a mapping relationship between the depth camera coordinate system and the color camera coordinate system through a multi-level coordinate space transformation chain, including: depth pixels to depth camera 3D points, depth camera 3D points to color camera 3D points, and color camera 3D points to color pixels; performs boundary processing on the floating-point coordinates and converts them into integer pixel coordinates to generate a registered depth map that is spatially aligned with the color map; and outputs the registered depth map to the target detection module; Step 4: The object detection module processes the registered depth map from the depth alignment module using a model based on the Transformer architecture. Preprocessing: resize the image to 800×600 pixels, normalize the pixel values, and convert it to BLOB format. Model processing: extract feature maps through the CNN backbone network, add position encoding, enhance global features through the Transformer encoder based on the multi-head self-attention mechanism, interact with the object query through the decoder, and output the bounding box and category prediction tensor. Post-processing: filter prediction results with a confidence level ≥ 0.5, identify the target category of the bridge structure component, parse and record the pixel coordinates of the center point of the bounding box, and output them to the spatial topology relationship modeling module. Step 5: The spatial topology relationship modeling module constructs a spatial index based on the target detection results. If there are two targets, an ordered index is constructed according to the horizontal coordinates. If there is a single target, the current horizontal coordinate is compared with the historical benchmark. If the offset is less than the dynamic tolerance threshold, it is determined to be the left topology target and the right target identifier is set to zero. If the offset is not less than the dynamic tolerance threshold, it is determined to be the right topology target and the left target identifier is set to zero. The constructed spatial index is output to the 3D coordinate calculation module. Step 6: The 3D coordinate calculation module calculates 3D coordinates based on the distinguished left and right target coordinates, extracts depth statistics of an N×N region with the target center, where N≥6; converts the depth data to a 3D world coordinate system through back-projection; calculates the vibration offset based on the IMU data, compensates the vibration offset with the coordinates, and smoothes the trajectory using an IIR filter with adaptive filter coefficients; and outputs the stabilized coordinates to the distance calculation module. Step 7: The distance calculation module determines the straight-line distance between the two target points based on the compensated 3D coordinates using the Euclidean distance decomposition algorithm, and calculates the xy plane distance and coordinate axis components of the 3D world coordinate system using the depth difference; the calculated data is output to the calibration optimization module; Step 8: The calibration optimization module uses a finite state machine to determine the spatial consistency of consecutive frames and trigger the tracking benchmark calibration; performs trajectory optimization, initializes the sliding window queue to store historical positions; removes the oldest data and adds the current frame offset; calculates the sliding mean; and outputs the smoothed data to the spatiotemporal fusion module; Step 9: In the spatiotemporal fusion module, the spatial transformation information and the temporal sequence information are fused to generate 3D spatiotemporal information for control or decision-making; the 3D spatiotemporal information is output to the visualization module and the storage module, and the dynamic spatial data set including the number of targets, spatial position coordinates, real-time offset vectors, and three-dimensional distance matrix is ​​stored in the remote terminal; Step 10: In the labeling module, a measurement bounding box is drawn in the corresponding detection area of ​​the real-time video stream image, and the spatial coordinate information and depth measurement results of the target are marked in the detection area.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method, device and equipment based on Transform model and storage medium

    CN116721207A

  • Real-time transmission method and system based on internet-of-things perception data in digital twinborn scene

    CN120143775A

  • Target identification tracking method and system based on multi-source fusion imaging

    CN120182323A