Cable-stayed bridge progress identification method based on point cloud segmentation
Through the improved OctFormer network and three-dimensional reconstruction technology, high-precision segmentation and construction progress identification of each component of the cable-stayed bridge are achieved, which solves the difficult identification problems in traditional methods and improves the automation and safety of the construction process.
Patent Information
- Application Number
- CN202510302320.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to identify the construction progress of each component of a cable-stayed bridge with high accuracy, especially at high altitudes or water components. The traditional point cloud segmentation algorithm has low accuracy, making it difficult to meet the actual needs of the construction process.
The improved OctFormer network is used for point cloud segmentation, combining drone data acquisition and three-dimensional reconstruction, and the precise segmentation and identification of each component of the cable-stayed bridge is achieved through the improved KAN network activation function and residual connection block (G-RCB).
It improves the accuracy and automation level of cable-stayed bridge construction progress identification, reduces manual intervention, ensures construction quality and safety, and provides real-time monitoring and feedback.
Smart Images

Figure CN120279441A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of point cloud segmentation and intelligent construction, and particularly to a method for identifying the progress of a cable-stayed bridge based on point cloud segmentation. Background Art
[0002] Currently, the progress identification of each component construction process level of a cable-stayed bridge is a complex engineering task, involving multiple aspects of difficulties. First, a cable-stayed bridge includes various types of structural components, such as bridge towers, main girders, stay cables, anchoring systems, etc. The construction characteristics and progress indicators of each component are different, and need to be distinguished and processed during identification. Second, many components of a cable-stayed bridge are located at high altitudes or in water, which brings difficulties to on-site measurement and data collection. Especially during the construction process, the issues of safety and accessibility are particularly prominent. Third, many components of a cable-stayed bridge have complex geometric shapes, such as curved stay cables and irregularly shaped bridge towers. These shapes are difficult to represent with traditional two-dimensional drawings, increasing the difficulty of identification.
[0003] Point cloud segmentation technology is widely used in the field of computer vision. Applying it to the progress identification of a cable-stayed bridge is an ideal intelligent detection approach, which can reduce safety risks and enable the progress identification during the construction process. However, although traditional point cloud networks can segment the point cloud data of each component of the cable-stayed bridge uploaded, the accuracy is relatively low, and the guiding significance for the actual construction process in the project is limited. Therefore, for the identification of the progress of each component of the cable-stayed bridge, how to use a high-precision point cloud segmentation algorithm for identification is an urgent need in the field. Summary of the Invention
[0004] The object of the present invention is to overcome the above-mentioned technical drawbacks and provide a method for identifying the progress of a cable-stayed bridge based on point cloud segmentation. This method can accurately distinguish different components of the cable-stayed bridge, such as bridge towers, main girders, stay cables, etc., which helps to identify the construction progress of each component targeted. Moreover, through automated point cloud segmentation, manual intervention is reduced, the automation level of construction progress monitoring is improved, and human errors are reduced.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] A method for identifying the progress of a cable-stayed bridge based on point cloud segmentation, the identification method includes the following steps:
[0007] Using a drone to take pictures to obtain three-dimensional point cloud data of each part of the cable-stayed bridge during the construction process;
[0008] Build an improved OctFormer network, which includes a conversion layer, an embedding module, an OctFormer module, a downsampling module, and a G-RCB module. The embedding module is composed of 5 mutually stacked octree convolution modules, and each octree convolution module is sequentially connected by an octree convolution, a batch normalization layer, and a ParametricReLU activation function;
[0009] The OctFormer module includes a layer normalization layer, an octree attention, a layer normalization layer, and a KAN network layer connected in sequence. The input of the first layer normalization layer is connected with the output of the octree attention through a residual connection to obtain a feature Y. The feature Y is used as the input of the second layer normalization layer and is connected with the output of the KAN network layer through a residual connection to obtain the output of the OctFormer module;
[0010] The activation function of the KAN network layer consists of a base function b(x) and a learnable function as follows:
[0011]
[0012] where, x is the input; is the learnable function, φ(x) is the scale function, expressed as β is the coefficient of the scale function; h0[k] is the coefficient of the scale function, taking values in [1 / 16, 1 / 4, 3 / 8, 1 / 4, 1 / 16], k is the index; N = 5; φ(2x - k) represents the scaled and translated version of the scale function at different positions; H is a preset learning parameter; C is a fixed coefficient; the silu neural network activation function, also known as the swish function;
[0013] By traversing all k, each coefficient h0[k] is multiplied by the corresponding scale function φ(2x - k) and summed to obtain the final scale function φ(x);
[0014] The point cloud data enters the conversion layer, is converted into an octree, and then after being processed by the embedding module, it enters the alternating processing of the OctFormer module and the downsampling module. The output of the last OctFormer module is input into the G-RCB module to obtain the output, thus completing the construction of the improved OctFormer network;
[0015] Use the three-dimensional point cloud data of each part of the cable-stayed bridge to train the improved OctFormer network for accurate segmentation of each part of the cable-stayed bridge, and output the bounding box of the target detection object and its corresponding class label;
[0016] Perform 3D reconstruction on the point cloud data of each part of the cable-stayed bridge after precise segmentation, and then compare it with the pre-planned BMI model to finally obtain the real-time progress of each part at present.
[0017] Furthermore, the structure of the G-RCB module is expressed as:
[0018] Y G-RCB = Y end + λ × Y G
[0019] where Y G is the output of the linear gating unit GLU, Y end is the input of the linear gating unit GLU, and λ is a trainable weight; the linear gating unit GLU consists of an input layer, a middle layer, and an output layer. The middle layer is composed of two convolutional layers and a sigmoid function. The following is the formula for the output h of the middle layer:
[0020]
[0021] where W and V are trainable parameter matrices; b and c are the corresponding bias terms, and σ represents the sigmoid function, is element-wise multiplication; finally, the output layer uses the softmax function. Input h into the output layer to obtain Y G .
[0022] Furthermore, the kernel sizes and strides of the octree convolutions in the five octree convolution modules are {3, 2, 3, 2, 3} and {1, 2, 1, 2, 1} respectively
[0023] The downsampling module consists of an octree convolution with a kernel size of 2 and a stride of 2 and a batch normalization layer.
[0024] The present invention also protects a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method can be implemented.
[0025] Furthermore, the storage medium is provided with:
[0026] A point cloud data collection module for collecting 3D point cloud data of each part during the construction of the cable-stayed bridge;
[0027] An improved OctFormer network: used to achieve real-time segmentation of each part of the cable-stayed bridge and segment the categories of each part of the cable-stayed bridge;
[0028] 3D reconstruction and comparison module: It is used to perform 3D reconstruction on the segmented point cloud data, connect to the improved OctFormer network, perform 3D reconstruction on the point cloud data of each part of the cable-stayed bridge segmented by the improved OctFormer network module, and then compare it with the pre-planned BIM model to identify the real-time progress.
[0029] Compared with the prior art, the main beneficial effects and innovation points of the present invention are:
[0030] The present invention effectively applies the point cloud segmentation algorithm to the progress recognition of cable-stayed bridges, can be combined with real-time data collection, provides real-time monitoring and feedback of the construction progress, helps to adjust the construction plan in a timely manner, can improve the measurement accuracy of the dimensions and positions of cable-stayed bridge components, and ensure the construction quality.
[0031] In the present invention, the improved OctFormer network proposes a new activation function for the KAN network in combination with the working scenario and segmentation requirements of the cable-stayed bridge, and improves the accuracy and efficiency through multi-level training, with higher accuracy. At the same time, an embedding module composed of the ParametricReLU activation function is adopted to improve the generalization ability of the model, allowing the network to learn within a wider range of inputs and accelerating the convergence speed of the model. At the same time, in cooperation with the residual connection block (G-RCB), it is more robust in the face of noise and outliers, enabling the model to have both the dual effects of segmentation accuracy and segmentation speed, and improving the performance of the model. Description of the Drawings
[0032] Figure 1 It is a schematic structural diagram of the improved OctFormer network in the present invention.
[0033] Figure 2 It is a schematic structural diagram of the downsampling module in the present invention.
[0034] Figure 3 It is a schematic structural diagram of the OctFormer module in the present invention.
[0035] Figure 4 It is a schematic structural diagram of the embedding module in the present invention.
[0036] Figure 5 It is a schematic structural diagram of the KAN network layer in the present invention. Detailed Embodiments
[0037] In order to more clearly describe the technical problems, technical solutions and advantages of the present invention, the following will be described in detail with reference to the drawings and embodiments. It should be noted that these embodiments are only used to illustrate the principles and application scopes of the present invention and should not be regarded as limitations on the present invention.
[0038] The present invention combines the improved OctFormer network with 3D reconstruction. Before performing 3D reconstruction, segmenting the point cloud can remove background noise and non-target objects, improve the accuracy of reconstruction, and also help identify and extract key features, which are crucial for reconstructing complex structures (such as the edges of buildings). The model after 3D reconstruction is compared with the construction progress in the preset BIM model, and then the progress of the cable-stayed bridge under the current construction can be determined. Here, the comparison method can adopt the degree of similarity.
[0039] Embodiment 1
[0040] A cable-stayed bridge progress recognition system based on point cloud segmentation in this embodiment includes:
[0041] Point cloud data collection module: Composed of a lidar, the lidar is installed on a drone and is used to collect 3D point cloud data of various parts of the cable-stayed bridge.
[0042] Improved OctFormer network: Used to achieve real-time segmentation of various parts of the cable-stayed bridge; connected to the point cloud data collection module, its main function is to perform real-time processing on the collected point cloud data and segment the categories of various parts of the cable-stayed bridge.
[0043] 3D reconstruction comparison module: Used to achieve 3D reconstruction of the segmented point cloud data; connected to the improved OctFormer network, mainly completed by Autodesk ReCap software; Autodesk ReCap software will perform 3D reconstruction on the point cloud data of various parts of the cable-stayed bridge segmented by the improved OctFormer network, and then compare it with the BIM model of each part of the cable-stayed bridge planned in advance to identify the real-time progress.
[0044] The improved OctFormer network includes a conversion layer, an embedding module, an OctFormer module, a downsampling module, and a G-RCB module. The specific process is as follows:
[0045] First, the point cloud data enters the conversion layer, is normalized, and is converted into an octree to obtain initial features. The initial features include the numerical values of the average point position, color, and point normal stored in the non-empty octree leaf nodes.
[0046] Then, it enters the embedding module, which contains 5 adjacent octree convolution modules. The embedding module projects the initial features into a high-dimensional space and enhances these features.
[0047] Next, the OctFormer module further extracts and integrates the features from the embedding module through the octree attention mechanism, enabling the features of each point to contain not only local information but also broader context information through the cross-window attention mechanism, forming a high-level feature representation.
[0048] Subsequently, the high-level feature representation passes through the downsampling module, which reduces the spatial resolution and increases the number of feature channels for multi-scale feature extraction.
[0049] Finally, after the alternating application of the OctFormer module and the downsampling module, the result is input into the G-RCB module with residual connections. After being processed by the G-RCB module, the output of the improved OctFormer network is obtained, which is the bounding box of the object to be detected and its corresponding class label.
[0050] Embodiment 2
[0051] In this embodiment, the specific composition of the improved OctFormer network is as follows:
[0052] The embedding module is as Figure 4 shown: This embedding module consists of 5 octree convolution modules. Each octree convolution module contains an octree convolution, a batch normalization layer, and a ParametricReLU activation function. The kernel sizes and strides of these octree convolutions are {3, 2, 3, 2, 3} and {1, 2, 1, 2, 1} respectively. When the stride is 2, the spatial resolution of the octree convolution downsampling tensor is 2 times. The embedding module is used to map the initial features output by the conversion layer into a high-dimensional feature space through convolution.
[0053] The ParametricReLU activation function allows non-zero gradients for negative input values, solving the "dying ReLU" problem of ReLU in the negative input value region by introducing learnable parameters. ParametricReLU helps improve the generalization ability of the model by adaptively adjusting the negative slope, allowing the network to learn over a wider range of input values. Since ParametricReLU also has non-zero gradients in the negative input region, this helps to accelerate the convergence speed of the model.
[0054] The OctFormer module is as Figure 3 shown: It mainly consists of octree attention, the KAN network layer, and the layer normalization layer. The high-dimensional features output from the embedding module first enter the layer normalization layer in the OctFormer module. The layer normalization layer normalizes each channel of the high-dimensional features to ensure that the input distributions of each layer are roughly the same. Then, the high-dimensional features x that have undergone layer normalization are input into the octree attention for the following operations: First, positional encoding is performed on the feature X to obtain X':
[0055] X’ = X + depth_wise_conv(x)
[0056] Where X is the input feature, with shape (N, C), N is the number of points, and C is the feature dimension of each point. depth_wise_conv is a depthwise convolution operation that independently applies a convolutional kernel to each input channel. This type of convolution can reduce the computational cost and the number of parameters while maintaining the ability to capture spatial features.
[0057] Then, calculate the number of empty grids N Z , to ensure that the feature X’ can be evenly divided into windows containing P points:
[0058] N Z = (P × D) - N % (P × D)
[0059] Where P is the number of points in each window, usually set to 32, and D is the dilation rate of the attention, default set to 1 or 4.
[0060] After calculating the number of empty grids N Z , pad the empty grids so that X’ can be divided, and then reshape X’ to create windows.
[0061] Next, create a mask m to ignore the padded zeros in the self-attention calculation. The shape of the mask matches X’ after window division, and then perform the self-attention calculation:
[0062] X” = attentin(query = X', key = X', value = X', key_padding_mask = m)
[0063] Where attentin represents self-attention;
[0064] Finally, reshape the result X” to restore it to the original feature shape. Remove the padded elements to obtain the final output feature X”’.
[0065] The feature X”’ output by the octree attention mechanism is combined with X through a residual connection to form a feature Y, and the feature Y is first layer-normalized and then input into the KAN network layer.
[0066] The KAN network layer (see Figure 5):A structure including two activation function layers. Among them, the input of the first layer is two identical features, and the input of the second layer is five. For the two inputs of the first layer, ten activation functions of the first layer are used for processing, and each input gets five duplicates. The sum of the first duplicate of the first input and the first duplicate of the second input is used as the first input of the second layer. The remaining four inputs of the second layer are obtained in the same way. The five inputs of the second layer are processed by five activation functions of the second layer, and then the five obtained results are added together to get the output of the KAN network layer. The formulas of each activation function in the two activation function layers are the same, but the parameters are not shared.
[0067] When the number of inputs of the first layer of the KAN network is n, the output is 2n + 1, the input of the second layer is 2n + 1, and the final output is 1. In this embodiment, the input is two, so the input of the second layer is 2×2 + 1 = 5. The activation function for each edge, that is, the non - linear transformation. l is the layer number, l = 0 represents the first layer, and l = 1 represents the second layer. Each input has 5 duplicates and then they are combined separately. Among them, i is used to mark the nodes of the current layer, and j is used to mark the nodes of the next layer. The output of each node X l,i is processed by the activation function and contributes to the calculation of all X l+1,j in the next layer.
[0068] The activation function is composed of the basic function b(x) and the learnable function as follows:
[0069]
[0070] Among them, x is the input; is the learnable function, φ(x) is the scale function, expressed as β is the coefficient of the scale function; h0[k] is the coefficient of the scale function, taking values in [1 / 16, 1 / 4, 3 / 8, 1 / 4, 1 / 16], k is the index; N = 5; φ(2x - k) represents the scaled and translated version of the scale function at different positions; H is the preset learning parameter; C is the fixed coefficient; the silu neural network activation function, also known as the swish function;
[0071] By traversing all k, each coefficient h0[k] is multiplied by the corresponding scale function φ(2x - k) and summed to obtain the final scale function φ(x);
[0072] The function has the characteristics of multi-scale and multi-resolution, which can effectively capture the local and global features of the signal, making the KAN network layer have better symmetry and linear phase characteristics while maintaining orthogonality.
[0073] The KAN network layer will try its best to fit the objective function. During the training process, the generalization ability of the model will be improved and the risk of overfitting will be reduced through the sparsification method. After sparsification, the pruning technique is further used to remove those unimportant connections and neurons. The KAN network will output the feature Y', which is then connected with the feature X''' through a residual connection to obtain the output feature Y'' of the OctFormer module.
[0074] The downsampling module is as Figure 2 shown: It consists of an octree convolution with a kernel size of 2 and a stride of 2 and a batch normalization layer. This module can reduce the spatial resolution of the feature Y'' and increase the number of channels of the feature map to 2 times. The spatial resolution of the feature output by the first OctFormer module is S / 4, where S is the spatial resolution. Every time a downsampling module is passed through, the spatial resolution is reduced by two times and the number of channels is increased by two times. C represents the number of channels.
[0075] The structure of the G-RCB module is as Figure 1 shown in, and the G-RCB is a residual structure as a whole:
[0076] Y G-RCB = Y end + λ × Y G
[0077] Among them, Y G is the output of the linear gated unit GLU, Y end is the input of the linear gated unit GLU, and λ is a trainable weight.
[0078] In the present invention, the linear gated unit GLU is added, which mainly consists of an input layer, an intermediate layer, and an output layer. The intermediate layer is composed of two convolutional layers and a sigmoid function. The following is the output formula of the intermediate layer:
[0079]
[0080] Among them, Y end is the input, W, V are trainable parameter matrices; b, c are the corresponding bias terms, σ represents the sigmoid function, is the element-wise multiplication. Finally, the output layer uses the softmax function. The h is input into the output layer to obtain Y G .
[0081] The improved OctFormer module of the present invention can learn which information is important and which can be ignored as a whole. It is more robust in the face of noise and outliers and can still show good recognition results at high altitudes or in water, achieving effective and high-precision segmentation of various parts of the cable-stayed bridge. At the same time, it improves the parameter efficiency when achieving similar or better performance, helps reduce the overfitting phenomenon of the model, and improves the performance of the model.
[0082] Embodiment 3
[0083] A cable-stayed bridge progress recognition method based on point cloud segmentation in this embodiment includes the following steps:
[0084] (1) Use a drone to take three-dimensional point cloud data of various parts of the cable-stayed bridge during the construction process and perform enhancement processing on the point cloud data.
[0085] (1.1) Use the point cloud data collection module to collect point cloud data. The collection scenario is various parts during the cable-stayed bridge construction process in daily life. The point cloud data collected in one scenario forms a point cloud set, and all the collected point cloud sets constitute a point cloud data set. In this embodiment, the point cloud data of various parts during the construction process is collected, and all the collected point cloud sets in all scenarios are randomly divided into a training set, a test set, and a validation set. Among them, the number of point cloud sets in the training set in this embodiment is 1000, the number of point cloud sets in the validation set is 300, and the number of point cloud sets in the test set is 400. On average, each point cloud set has 150 points.
[0086] (1.2) Center-normalize the collected point cloud data set, that is, subtract the mean value and divide by the maximum distance from the point to the origin.
[0087] (1.3) Then rotate the point cloud data, perform random rotation within the range of [-180°, 180°] about the x, y, and z axes, which can simulate the scenarios of various components in different directions.
[0088] (1.4) Then scale and translate the point cloud data to simulate the observation effects at different distances. Use a scaling factor within the range of [0.75, 1.25] for global scaling and a random translation within [-0.1, 0.1].
[0089] (1.5) Finally, add random noise to the point cloud data to simulate the sensor noise in the real world and obtain the enhanced data set.
[0090] (2) Construct an improved OctFormer network, train it, and then perform testing and validation. The trained improved OctFormer network can achieve precise segmentation of various parts of the cable-stayed bridge.
[0091] The training process is as follows: The training set is input into the network, and the network is trained for 600 epochs with a batch size of 16 and a weight decay of 0.05. The initial learning rate is set to 0.006 and is reduced by a factor of 10 after 360 and 480 epochs respectively.
[0092] The optimal training parameters are obtained when the training gradient value of the training set is within the range of ±1e-5. Then, it is tested and verified on the test set and the validation set. If the results of the two are the same and the object detection accuracy is not lower than 95%, the training ends and the trained improved OctFormer network is obtained.
[0093] (3) Input the point cloud data to be recognized into the trained improved OctFormer network for precise segmentation of each part of the cable-stayed bridge.
[0094] (4) After obtaining the point cloud data of each part of the cable-stayed bridge after precise segmentation, input it into the Autodesk ReCap software for 3D reconstruction, and then compare it with the pre-planned BMI model to finally obtain the real-time progress of each part at present.
[0095] The present invention performs point cloud segmentation through a more precise and fast improved OctFormer network and combines it with 3D reconstruction to create a more precise and detailed 3D model, making the recognition of the construction progress of the cable-stayed bridge more precise, convenient, time-saving and labor-saving. Through automated point cloud segmentation, manual intervention is reduced, the automation level of construction progress monitoring is improved, and human errors are reduced.
[0096] To further highlight the advantages of the present application, the following uses the enhanced dataset above to compare the segmentation effects of the segmentation network of the present application with the current mainstream point cloud-based segmentation methods. The comparison results are shown in Table 1. Val. and Test respectively represent the mIoU on the validation and test sets. '-' indicates that the results are not reported or the source code is not public, and the best results are marked in bold.
[0097] Table 1 Comparison of overall performance
[0098] Method VL. Test. 3DMV - 48.4 PanopticFusion - 52.9 PointNet++ 53.5 55.7 SegGCN - 58.9 PointConv 61.0 66.6 PointTransformer 70.6 - StratifiedTransformer 74.3 73.7 PointTransformerV2 85.2 86.6 The method of the present invention 90.5 90.6
[0099] Those not described in the present invention are applicable to the prior art.
Claims
1. A method for identifying the progress of a cable-stayed bridge based on point cloud segmentation, characterized in that The recognition method includes the following steps: Using a drone to take pictures to obtain the three-dimensional point cloud data of each part of the cable-stayed bridge during the construction process; Constructing an improved OctFormer network, the improved OctFormer network includes a conversion layer, an embedding module, an OctFormer module, a downsampling module and a G-RCB module. The embedding module is composed of 5 stacked octree convolution modules, and the octree convolution module is composed of an octree convolution, a batch normalization layer and a ParametricReLU activation function in series; The OctFormer module includes a layer normalization layer, an octree attention, a layer normalization layer and a KAN network layer connected in sequence. The input of the first layer normalization layer is connected with the output of the octree attention through a residual connection to obtain a feature Y. The feature Y is used as the input of the second layer normalization layer and is connected with the output of the KAN network layer through a residual connection to obtain the output of the OctFormer module; The activation function of the KAN network layer composed of the basic function b(x) and the learnable function is formed according to the following formula: Among them, x is the input; is a learnable function, φ(x) is the scaling function, expressed as β is the coefficient of the scaling function; h0[k] is the coefficient of the scaling function, taking values in [1 / 16, 1 / 4, 3 / 8, 1 / 4, 1 / 16], k is the index; N = 5; φ(2x - k) represents the scaled and translated versions of the scaling function at different positions; H is a preset learning parameter; C is a fixed coefficient; the silu neural network activation function, also known as the swish function; By traversing all k, multiplying each coefficient h0[k] by the corresponding scaling function φ(2x - k) and summing, the final scaling function φ(x) is obtained; The point cloud data enters the conversion layer, is converted into an octree, and then processed by the embedding module, and then enters the alternating processing of the OctFormer module and the downsampling module. The output of the last OctFormer module is input into the G-RCB module to obtain the output, thus completing the construction of the improved OctFormer network; Training the improved OctFormer network with the three-dimensional point cloud data of each part of the cable-stayed bridge for precise segmentation of each part of the cable-stayed bridge, and outputting the bounding box of the target detection object and its corresponding class label; Performing three-dimensional reconstruction on the point cloud data of each part of the cable-stayed bridge after precise segmentation, and then comparing it with the pre-planned BIM model to finally obtain the real-time progress of each part at present.
2. The cable-stayed bridge progress identification method according to claim 1, wherein The structure of the G-RCB module is expressed as: Y G-RCB = Y end + λ × Y G Among them, Y G is the output of the linear gating unit GLU, and Y end is the input of the linear gating unit GLU, and λ is a trainable weight; the linear gating unit GLU is composed of an input layer, an intermediate layer, and an output layer. The intermediate layer is composed of two convolutional layers and a sigmoid function. The following is the formula for the output h of the intermediate layer: where W and V are trainable parameter matrices; b and c are corresponding bias terms, and σ represents the sigmoid function. denotes element-wise multiplication; finally, the output layer uses the softmax function. Input h into the output layer to obtain Y. G .
3. The cable-stayed bridge progress recognition method according to claim 1, characterized in that The kernel sizes and strides of the octree convolutions in the five octree convolution modules are {3, 2, 3, 2, 3} and {1, 2, 1, 2, 1} respectively The downsampling module is composed of an octree convolution with a kernel size of 2 and a stride of 2 and a batch normalization layer.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it can implement the steps of the method according to any one of claims 1-3.
5. The computer-readable storage medium according to claim 4, wherein The storage medium is provided with: A point cloud data collection module for collecting the three-dimensional point cloud data of each part during the construction process of the cable-stayed bridge; An improved OctFormer network: used to realize the real-time segmentation of each part of the cable-stayed bridge and segment the categories of each part of the cable-stayed bridge; A three-dimensional reconstruction comparison module: used to perform three-dimensional reconstruction on the segmented point cloud data; connected to the improved OctFormer network, perform three-dimensional reconstruction on the point cloud data of each part of the cable-stayed bridge segmented by the improved OctFormer network module, and then compare it with the pre-planned BIM model to identify the real-time progress.