A method and system for identifying construction processes in a tunnel based on visual information

Through visual information, the tunnel construction process is identified, and the sparse sampling and multi-scale feature extraction network combined with optical flow data is used to solve the problem of unreasonable organization of processes in tunnel construction, and efficient construction management is achieved.

CN119007077BActive Publication Date: 2025-07-18SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411092754.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2025-07-18
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

In tunnel construction, the unreasonable organization of the construction process leads to inefficient construction, extended construction period and increased costs, especially when railway lines are constructed in parallel.

Method used

The construction process identification method based on visual information is adopted to achieve efficient identification of construction processes through uniform sparse sampling strategy, multi-scale spatiotemporal feature extraction network and optical flow data feature extraction, combined with spatial attention mechanism and residual network.

Benefits of technology

It improves the identification efficiency and accuracy of construction processes in the tunnel, ensures intelligent and automated management of the construction process, and reduces the impact of construction costs and construction periods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119007077B_ABST
    Figure CN119007077B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for identifying construction processes in a tunnel based on visual information. A uniform sparse sampling strategy is adopted to dynamically sample the acquired video; a multi-scale spatio-temporal feature extraction network is constructed, and the constructed extraction network is used to extract and fuse multi-scale spatio-temporal features of the sampled video images. Optical flow data information is introduced, and an optical flow feature extraction network is used to process the optical flow data information to extract optical flow data features. The optical flow feature extraction network is added with a spatial attention mechanism, and key information is screened out through the spatial attention mechanism, and then the key information is sent to a residual network for feature extraction; the multi-scale spatio-temporal features and optical flow data features are mapped to a feature fusion space, and based on the mapped and fused data, classification is performed to obtain the result of process identification. The present invention can improve the efficiency and accuracy of identifying construction processes in a tunnel and is beneficial to the control of construction processes in a tunnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of in-tunnel construction, and particularly relates to a method and system for identifying in-tunnel construction processes based on visual information. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] With the increasing amount of traffic infrastructure construction, the situation of parallel line construction is quite common. During construction, the construction workface in tunnels is small, and the space in single-track tunnels is limited. Especially when facing parallel construction with a railway line that is about to undergo joint commissioning and testing, the construction period pressure is even greater. If the implementation and connection of construction processes cannot be reasonably organized, after the line to be opened enters joint commissioning and testing, the tunnel under construction will directly face construction on the existing line, which will inevitably significantly reduce the on-site construction efficiency, have a huge impact on the project's construction period, and also greatly increase the subsequent construction costs. Summary of the Invention

[0004] In order to solve the above problems, the present invention proposes a method and system for identifying in-tunnel construction processes based on visual information. The present invention can improve the efficiency and accuracy of identifying in-tunnel construction processes and is beneficial to the control of in-tunnel construction processes.

[0005] According to some embodiments, the present invention adopts the following technical solutions:

[0006] A method for identifying in-tunnel construction processes based on visual information, comprising the following steps:

[0007] Obtain the construction monitoring video in the tunnel;

[0008] Perform dynamic sampling on the obtained video using a uniform sparse sampling strategy;

[0009] Construct a multi-scale spatio-temporal feature extraction network, and use the constructed extraction network to extract and fuse multi-scale spatio-temporal features of the sampled video images. After extracting high-level and low-level spatio-temporal features, the multi-scale spatio-temporal feature extraction network introduces the semantic information of high-level spatio-temporal features into low-level spatio-temporal features using a semantic feature embedding fusion method;

[0010] Introduce optical flow data information, and use an optical flow feature extraction network to process the optical flow data information to extract optical flow data features. The optical flow feature extraction network is added with a spatial attention mechanism to screen out key information through the spatial attention mechanism, and then send the key information into a residual network for feature extraction;

[0011] Map the multi-scale spatio-temporal features and optical flow data features to the feature fusion space, and classify based on the fused data to obtain the process recognition result.

[0012] As an alternative implementation, the specific process of dynamically sampling the acquired video using a uniform sparse sampling strategy includes: Suppose there are N feature maps in the current video clip, then the current sampling value S is expressed as:

[0013] S = [N / L]

[0014] In the formula, S represents the sampling value of the current video, N represents the number of image frames of the current video after data preprocessing, and L represents the amount of data input to the network. Index the position of the feature map according to the obtained sampling value S to get the model input L = [L(0), L(1)S,..., L(l - 1)S, L(l)S].

[0015] As an alternative implementation, the multi-scale spatio-temporal feature extraction network includes a multi-scale convolution module, a basic backbone network, and a multi-feature information aggregation module. The multi-scale convolution module is used to obtain the global features of the image, the basic backbone network is used to generate high-level and low-level spatio-temporal features, and the multi-feature information aggregation module is used to introduce the semantic information of the high-level spatio-temporal features into the low-level spatio-temporal features by using the semantic feature embedding fusion method, enhance the semantic expression of the low-level spatio-temporal features, and make the context spatio-temporal information and scale information complement each other.

[0016] Furthermore, the multi-scale convolution module includes 3D convolution kernels of 3 different sizes connected in sequence.

[0017] Furthermore, the basic backbone network includes 8 3D convolutional layers, pooling layers, and fully connected layers. Among them, the 8 3D convolutional layers are connected in sequence, and a pooling layer is connected after each 3D convolutional layer or every two convolutional layers. Finally, two fully connected layers are connected. The first 3D convolutional layer is the multi-scale convolution module, and all pooling layers use the max pooling operation.

[0018] Furthermore, in the process of multi-scale spatio-temporal feature extraction, the 3D convolution kernel moves simultaneously in the x, y, and z directions. The calculation process of the j-th feature map of the i-th layer at (x, y, z) is as follows:

[0019] ;

[0020] In the formula, m represents the feature map in the (i - 1)-th layer connected to the current feature map; is the value of the j-th feature map of the i-th layer at (x, y, z), f is the Relu activation function, usually expressed as f(x) = max(0, x), is the sum of all M feature maps in the (i - 1)-th layer connected to the current feature map, For summation within the range of the size H(i) of the convolutional kernel in the time dimension (z direction), For summation within the range of the size W(i) of the convolutional kernel in the width (y direction), For summation within the range of the size L(i) of the convolutional kernel in the length (x direction), Denotes the convolutional kernel weight connected to the m-th feature map of the (i - 1)-th layer; Denotes the value of the m-th feature map of the (i - 1)-th layer at (x + l, y + w, z + h), Denotes the bias of the j-th feature map of the i-th layer.

[0021] As a further aspect, the multi-feature information aggregation module includes 4 parallel 1×1×1 convolutional kernels, which are used to set the channel values of the high and low-level spatio-temporal features to set values. Through semantic embedding, the high-level features are resampled and fused with the sub-high-level features in a top-down manner. The high-level semantic information is used to improve the low-level detailed information, and then the fused features are resampled and fused with the next layer of features to enhance the semantic expression of the low-level spatio-temporal features;

[0022] Convolutional kernels with different strides are used to map the spatio-temporal feature maps into feature maps with the same dimension, and then the high and low-level spatio-temporal features after embedding semantic information are fused.

[0023] As an alternative implementation, the optical flow feature extraction network uses the deep residual network ResNet101 model as the basic structure, and adds the Transformer attention mechanism to the basic structure. Specifically, two feature maps of size 1×H×W are obtained by using a max pooling layer and an average pooling layer, and then point-to-point spatial information is obtained through a convolutional layer of size 7×7. The SIGMOID function is used to activate the spatial information, and finally the spatial attention activation map is obtained.

[0024] As an alternative implementation, the SoftMax function is used to map the feature vectors into a probability sequence, the corresponding results generated by the encoder of the classification sequence are extracted, and the results are classified to obtain the final classification result.

[0025] A tunnel construction process recognition system based on visual information, comprising the following steps:

[0026] A video acquisition module, configured to acquire a tunnel construction monitoring video;

[0027] A dynamic sampling module, configured to perform dynamic sampling on the acquired video using a uniform sparse sampling strategy;

[0028] The multi-scale spatio-temporal feature extraction module is configured to construct a multi-scale spatio-temporal feature extraction network, and use the constructed extraction network to extract and fuse multi-scale spatio-temporal features from the sampled video images. After extracting high-level and low-level spatio-temporal features, the multi-scale spatio-temporal feature extraction network introduces the semantic information of high-level spatio-temporal features into low-level spatio-temporal features by means of semantic feature embedding fusion;

[0029] The optical flow data feature extraction module is configured to introduce optical flow data information, and use the optical flow feature extraction network to process the optical flow data information to extract optical flow data features. The optical flow feature extraction network is added with a spatial attention mechanism to filter out key information through the spatial attention mechanism, and then send the key information to the residual network for feature extraction;

[0030] The fusion recognition module is configured to map the multi-scale spatio-temporal features and optical flow data features to the feature fusion space, and perform classification based on the mapped and fused data to obtain the process recognition result.

[0031] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the steps in the above method are completed.

[0032] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the above method are completed.

[0033] Compared with the prior art, the beneficial effects of the present invention are:

[0034] Based on the full-video time-domain modeling, the present invention extracts global spatio-temporal features adaptable to perspective changes through a feature extraction network based on multi-column convolution, and obtains deep features of optical flow data through a feature extraction network guided by the Transformer attention mechanism; then, feature fusion is performed in the fully connected layer and the final behavior recognition is completed using a SoftMax classifier, and the final recognition result is accurate and efficient.

[0035] When obtaining the long-term time-domain information of the full video segment and establishing a video-level feature extraction network, the present invention introduces a uniform sparse sampling strategy to complete the time-domain modeling of the full video segment, and fully retains the long-time sequence information on the premise of reducing the redundancy of video frames.

[0036] The present invention constructs a multi-feature information aggregation module for aggregating high-level and low-level spatio-temporal features, and solves the problem that the lack of high-level semantic information or low-level spatial detail information will affect the final event recognition result and lead to a decrease in accuracy.

[0037] When extracting optical flow information, the present invention can effectively locate key information in the image through the spatial attention mechanism, effectively improving the network performance.

[0038] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following presents preferred embodiments in conjunction with the accompanying drawings and provides a detailed description as follows. Description of the Drawings

[0039] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0040] Figure 1 It is a schematic diagram of a process identification framework;

[0041] Figure 2 It is a schematic diagram of a multi-scale spatio-temporal feature extraction network;

[0042] Figure 3 It is a schematic diagram of a multi-scale module convolution structure;

[0043] Figure 4 It is a schematic diagram of a multi-scale convolution block structure;

[0044] Figure 5 It is a schematic diagram of a C3D network structure;

[0045] Figure 6 It is a schematic diagram of an optical flow feature extraction network structure;

[0046] Figure 7 It is a schematic diagram of a spatial attention module. Detailed Embodiments

[0047] The present invention will be further described below in conjunction with the drawings and embodiments.

[0048] It should be noted that the following detailed description is illustrative and is intended to provide a further description of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0049] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. Additionally, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0050] Embodiment 1

[0051] A construction process identification method based on visual information, as Figure 1 shown, includes the following steps:

[0052] S1. Use a uniform sparse sampling strategy to perform data sampling, and then complete the time-domain modeling of the entire video segment;

[0053] S2. Obtain multi-scale spatio-temporal features through multi-column convolution to weaken the interference brought by perspective changes to video images;

[0054] S3. Introduce optical flow data information, and obtain deep features of the optical flow data through a feature extraction network guided by a spatial attention mechanism;

[0055] S4. Fuse the obtained multi-scale spatio-temporal features and optical flow information in the fully connected layer of the network to achieve end-to-end long video behavior recognition;

[0056] S5. Input the continuous video frame data processed by the data conversion layer into the SM classifier for classification.

[0057] The following specifically introduces each step. First, step S1. In this embodiment, recognition is mainly performed through video monitoring. However, considering that the frame interval of long-time series events is relatively small, random sampling is likely to introduce a large amount of redundant information and eliminate the correlation of video images in the time dimension. Therefore, when obtaining the long-time domain information of the entire video segment and establishing a video-level feature extraction network, this embodiment introduces a uniform sparse sampling strategy to complete the time-domain modeling of the entire video segment, and fully retains the long-time series information on the premise of reducing the redundancy of video frames. Assume that the current video clip has N feature maps, then the current sampling value S can be expressed as:

[0058] S = [N / L]

[0059] In the formula, S represents the sampling value of the current video, N represents the number of image frames of the current video after data preprocessing, and L represents the amount of data input to the network. According to the obtained sampling value S, the position index of the feature map is performed to obtain the model input L = [L(0), L(1)S,..., L(l - 1)S, L(l)S].

[0060] This embodiment ensures the consistency of the amount of data input to the network through the dynamic sampling value, without adjusting other parameters, and can be applied to video data of different durations.

[0061] In step S2, the multi-scale spatio-temporal feature extraction network, such as Figure 2As shown in the figure, it mainly includes three parts: a multi-scale convolution module, a basic backbone network C3D, and multi-feature information aggregation. First, the global features of the original image are obtained through the multi-scale convolution module, and then the high-level and low-level spatio-temporal features are generated by using the basic backbone network. Finally, through the semantic feature embedding and fusion method of the multi-feature aggregation module, the semantic information contained in the high-level spatio-temporal features is introduced into the low-level spatio-temporal features, enhancing the semantic expression of the low-level spatio-temporal features, making the context spatio-temporal information and scale information complement each other, and improving the network's representation ability of spatio-temporal features.

[0062] The multi-scale convolution is a multi-scale convolution module based on a multi-column structure for spatio-temporal feature extraction. The specific structure is as Figure 3 shown. Since there are often dynamic switches in the viewing angle when shooting videos, resulting in large-scale changes in video images, and single-column convolution is difficult to handle the scale change problem in video images. In the multi-scale convolution module, three 3D convolution kernels of different sizes are used to learn scale-related features from the original input image block to effectively obtain multi-scale information. The multi-scale convolution block structure adopted in this paper is as Figure 4 shown. Using convolution kernels of 3×3×3, 5×5×5, and 7×7×7 can effectively aggregate global spatio-temporal information.

[0063] Feature extraction is performed using the aforementioned basic backbone network C3D. This network takes stacked video RGB frames as input data and then uses 3D convolution kernels for feature extraction. The size of the convolution kernel determines the effectiveness of extracting video features. Due to problems such as dynamic occlusion and viewing angle changes in video images, it is required that the features extracted by the network must be general and effective. At the same time, in the time dimension, the connection between video features should be compact. Based on this, all the convolution kernels in the 8 3D convolution layers included in the C3D network are set to 3×3×3; the pooling layers all use max pooling operations. Among them, the kernel of pool1 is 1×2×2, and the kernels of the remaining pooling layers are all 2×2×2; the network has 2 fully connected layers, which are mainly used for dimensionality reduction of the feature vectors. The network structure is as Figure 5 shown. In order to handle the scale change problem in video images, the first convolution layer of this network is replaced with a multi-scale convolution module, and the global features of the original image are obtained through multi-scale convolution.

[0064] Regarding the aforementioned multi-feature fusion, during the multi-scale spatio-temporal feature extraction process, the 3D convolution kernel moves simultaneously in the x, y, and z directions. The calculation process of the j-th feature map of the i-th layer of the neural network at (x, y, z) is as follows:

[0065] ;

[0066] In the formula, m represents the feature map in the (i - 1)-th layer that is connected to the current feature map; is the value of the j-th feature map in the i-th layer at (x, y, z), and f is the Relu activation function, usually expressed as f(x) = max(0, x). is the sum of all M feature maps in the (i - 1)-th layer connected to the current feature map. is the sum within the range of the size H(i) of the convolutional kernel in the time dimension (z direction). is the sum within the range of the size W(i) of the convolutional kernel in the width (y direction). is the sum within the range of the size L(i) of the convolutional kernel in the length (x direction). represents the convolutional kernel weight connected to the m-th feature map in the (i - 1)-th layer; represents the value of the m-th feature map in the (i - 1)-th layer at (x + l, y + w, z + h). represents the bias of the j-th feature map in the i-th layer.

[0067] As the convolutional layer network deepens, some feature information is lost during the convolution process. Since the receptive field of the high-level spatio-temporal feature network is relatively large, the high-level spatio-temporal features extracted contain more semantic information and fewer spatial detail features; the receptive field of the low-level spatio-temporal feature network is relatively small, and the low-level spatio-temporal features extracted contain more spatial detail information and less high-level semantic information. If high-level semantic information or low-level spatial detail information is missing, it will affect the final event recognition result and lead to a decrease in accuracy.

[0068] To address this problem, in this embodiment, a multi-feature information aggregation module is constructed to aggregate high-level and low-level spatio-temporal features. First, 4 parallel 1×1×1 convolutional kernels are used to set the channel values of the high-level and low-level spatio-temporal features to 512; then, through semantic embedding, the high-level features are resampled and fused with the sub-high-level features in a top-down manner, and the high-level semantic information is used to improve the low-level detail information. Then, the fused features are resampled and fused with the next layer of features to enhance the semantic expression of the low-level spatio-temporal features. After that, 3×3×3 convolutional kernels with different strides are used to map the spatio-temporal feature maps into feature maps with the same dimension; finally, the high-level and low-level spatio-temporal features embedded with semantic information are fused, and the calculation formula for the fused high-level and low-level spatio-temporal features is as follows:

[0069]

[0070] Among them, F hl : represents within the set range (from l min to l max the sum of all feature values F l and finally calculates a new feature valueF hl 。

[0071] l min : It is the starting point of the accumulation operation, indicating from which eigenvalue the accumulation starts.

[0072] l max : It is the termination point of the accumulation operation, indicating until which eigenvalue the accumulation ends.

[0073] F l : They are the respective eigenvalues for accumulation. As l from l min varies to l max , all these eigenvalues will be accumulated in sequence to form the final eigenvalue F hl 。

[0074] In step S3, optical flow data is introduced as another input modality of the model, and an optical flow feature extraction network is used.

[0075] For the feature map F, first, it passes through a max pooling layer and an average pooling layer to obtain two feature maps of size 1×H×W, then obtains point-to-point spatial information through a convolutional layer of size 7×7, and then uses the SIGMOID function to activate the spatial information to obtain the finally obtained spatial attention activation map. The specific structure is as Figure 6 shown.

[0076] F is a comprehensive feature obtained through the hierarchical fusion of multiple layers of features, combining high-level semantic information and low-level detail information. The introduction of optical flow data and the use of a deep residual network further improve the effect of feature extraction. The introduction of optical flow features reduces the interference of illumination changes and complex backgrounds, and the spatial attention mechanism effectively locates the significant regions where actions occur, improving the network performance.

[0077] The main reason for adopting optical flow information is that: optical flow is the instantaneous velocity of pixel motion of a spatial moving object on the observation plane, which can reflect information such as the speed and direction of the moving object in the video image; optical flow has apparent invariance, which is manifested in that the complex background and the differences of the moving object itself in the video will not affect the manifestation form of optical flow. Taking the optical flow map as another input modality of the network to reduce the interference of factors such as illumination changes and complex backgrounds.

[0078] In the past, the network structures used for extracting optical flow features were relatively shallow. The extraction of optical flow information focused more on shallow detail information, while ignoring the deeper high-level semantic information in the optical flow. To fully exploit the potential features of optical flow data, in this embodiment, the deep residual network ResNet101 model is used as the basic structure. Considering that the key information in the optical flow map often aggregates in the area where the action occurs, a spatial attention mechanism is added to the basic network in the present invention. The key information is selected through spatial attention and then sent to the residual network for feature extraction. The essence of the attention mechanism is to locate the information relevant to the current task and suppress irrelevant information. Since the content presented in the optical flow map is the area where significant changes in the action occur, the key information in the image can be effectively located through the spatial attention mechanism, effectively improving the network performance.

[0079] In S4, the 4096-dimensional multi-scale spatio-temporal features and the 4096-dimensional optical flow data features are mapped to a 4096-dimensional feature fusion space. The advantage of this feature fusion method is mainly that the model can learn the respective feature parameters of the two parallel networks during the training stage and complete the coordinated feedback autonomously, realizing the end-to-end training of the model.

[0080] In S5, the SoftMax function is selected to map the feature vector into a probability sequence to retain more original information of the features. The classification sequence corresponding to the result generated by the encoder is extracted. The obtained result is input into the MLP module, and the probabilities of each behavior are compared. The event with the highest probability is the result of the final event classification.

[0081] The MLP module and the SoftMax classifier work together. The former is used for feature learning and processing, and the latter is used for the final probability mapping and classification decision. The MLP module further processes the feature vector output by the encoder, learns more complex feature relationships, and generates the final feature vector for classification.

[0082] Finally, through the application of feature fusion, encoder processing, the MLP module, and the SoftMax function, the real-time detection and classification of tunnel construction processes can be achieved. The pre-stored event categories correspond to each specific construction process. The feature representations of various processes are learned during the model training stage. When applied to the actual scenario, the current construction process can be accurately identified and classified, ensuring the intelligent and automated management of the construction process.

[0083] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented on one or more computer-usable storage media (including but not limited to disk storage, CD - ROM, in the form of a computer program product implemented on an optical memory or the like).

[0084] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0085] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operating steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0087] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made by those skilled in the art without creative efforts within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for identifying construction processes in a tunnel based on visual information, characterized in that It includes the following steps: Obtain the construction monitoring video in the tunnel; Perform dynamic sampling on the obtained video using a uniform sparse sampling strategy; Construct a multi-scale spatio-temporal feature extraction network, and use the constructed extraction network to extract and fuse multi-scale spatio-temporal features from the sampled video images. The multi-scale spatio-temporal feature extraction network includes a multi-scale convolution module, a basic backbone network, and a multi-feature information aggregation module. The multi-scale convolution module includes 3 3D convolution kernels with different sizes connected in sequence; the basic backbone network includes 8 3D convolution layers, pooling layers, and fully connected layers. Among them, the 8 3D convolution layers are connected in sequence, and a pooling layer is connected after each 3D convolution layer or every two convolution layers. Finally, two fully connected layers are connected. The first 3D convolution layer is the multi-scale convolution module, and each pooling layer uses a max pooling operation; The multi-feature information aggregation module includes 4 parallel 1×1×1 convolution kernels, which are used to set the channel values of the high-level and low-level spatio-temporal features to set values. Through semantic embedding, the high-level features are resampled and fused with the sub-high-level features in a top-down manner. The high-level semantic information is used to improve the low-level detail information, and then the fused features are resampled and fused with the next-level features to enhance the semantic expression of the low-level spatio-temporal features; Use convolution kernels with different strides to map the spatio-temporal feature map into a feature map with the same dimension, and then fuse the high-level and low-level spatio-temporal features embedded with semantic information; Introduce optical flow data information, and use an optical flow feature extraction network to process the optical flow data information to extract optical flow data features. The optical flow feature extraction network uses the deep residual network ResNet101 model as the basic structure, and adds a spatial attention mechanism to the basic structure. Specifically, two feature maps of size 1×H×W are obtained by using a max pooling layer and an average pooling layer, and then point-to-point spatial information is obtained through a 7×7 convolution layer. The SIGMOID function is used to activate the spatial information, and finally the spatial attention activation map is obtained; Map the multi-scale spatio-temporal features and optical flow data features to the feature fusion space, and perform classification based on the mapped and fused data to obtain the process recognition result.

2. The method for identifying construction processes in a tunnel based on visual information according to claim 1, characterized in that, The specific process of performing dynamic sampling on the obtained video using a uniform sparse sampling strategy includes: assuming that the current video clip has N feature maps, then the current sampling value S is expressed as: S = [N / L] In the formula, S represents the sampling value of the current video, N represents the number of image frames of the current video after data preprocessing, and L represents the amount of data input to the network. According to the obtained sampling value S, the position index of the feature map is obtained, and the model input L = [L(0), L(1)S,..., L(l - 1)S, L(l)S] is obtained.

3. The method for identifying construction processes in a tunnel based on visual information according to claim 1, characterized in that The multi-scale spatio-temporal feature extraction network includes a multi-scale convolution module, a basic backbone network, and a multi-feature information aggregation module. The multi-scale convolution module is used to obtain the global features of the image, the basic backbone network is used to generate high-level and low-level spatio-temporal features, and the multi-feature information aggregation module is used to introduce the semantic information of the high-level spatio-temporal features into the low-level spatio-temporal features by means of semantic feature embedding fusion, enhance the semantic expression of the low-level spatio-temporal features, and make the context spatio-temporal information and scale information complement each other.

4. The method for identifying construction processes in a tunnel based on visual information according to claim 3, characterized in that, in During the multi-scale spatio-temporal feature extraction process, the 3D convolution kernel moves simultaneously in the x, y, and z directions. The calculation process of the j-th feature map of the i-th layer at (x, y, z) is as follows: ; Wherein, m represents the feature map connected to the current feature map in the (i-1)-th layer; is the value of the j-th feature map in the i-th layer at (x, y, z), f is the Relu activation function, usually expressed as f(x) = max(0, x), is the sum of all M feature maps connected to the current feature map in the (i-1)-th layer, is the sum within the size H(i) of the convolutional kernel in the time dimension, i.e., the z direction, is the sum within the size W(i) of the convolutional kernel in the width, i.e., the y direction, is the sum within the size L(i) of the convolutional kernel in the length, i.e., the x direction, represents the convolutional kernel weight connected to the m-th feature map in the (i-1)-th layer; represents the value of the m-th feature map in the (i-1)-th layer at (x + l, y + w, z + h), represents the bias of the j-th feature map in the i-th layer.

5. The method for identifying construction processes in a tunnel based on visual information according to claim 1, characterized in that, The SoftMax function is used to map the feature vector into a probability sequence, the corresponding result generated by the encoder is extracted from the classification sequence, and the result is classified to obtain the final classification result.

6. A tunnel construction process identification system based on visual information, characterized in that, It includes the following steps: A video acquisition module, configured to acquire a construction monitoring video in the tunnel; A dynamic sampling module, configured to perform dynamic sampling on the acquired video using a uniform sparse sampling strategy; A multi-scale spatio-temporal feature extraction module, configured to construct a multi-scale spatio-temporal feature extraction network, and use the constructed extraction network to extract and fuse multi-scale spatio-temporal features of the sampled video images. The multi-scale spatio-temporal feature extraction network includes a multi-scale convolution module, a basic backbone network, and a multi-feature information aggregation module. The multi-scale convolution module includes 3 3D convolution kernels of different sizes connected in sequence; the basic backbone network includes 8 3D convolution layers, pooling layers, and fully connected layers. Among them, the 8 3D convolution layers are connected in sequence, and a pooling layer is connected after each 3D convolution layer or every two convolution layers. Finally, two fully connected layers are connected. The first 3D convolution layer is the multi-scale convolution module, and each pooling layer uses the maximum pooling operation; The multi-feature information aggregation module includes 4 parallel 1×1×1 convolution kernels, which are used to set the channel values of the high-level and low-level spatio-temporal features to set values. Through semantic embedding, the high-level features are resampled and fused with the sub-high-level features from top to bottom. The high-level semantic information is used to improve the low-level detail information, and then the fused features are resampled and fused with the next layer of features to enhance the semantic expression of the low-level spatio-temporal features; Convolution kernels with different strides are used to map the spatio-temporal feature maps into feature maps with the same dimension, and then the high-level and low-level spatio-temporal features after embedding semantic information are fused; An optical flow data feature extraction module, configured to introduce optical flow data information, and use an optical flow feature extraction network to process the optical flow data information to extract optical flow data features. The optical flow feature extraction network uses the deep residual network ResNet101 model as the basic structure, and a spatial attention mechanism is added to the basic structure. Specifically, two feature maps of size 1×H×W are obtained by using a maximum pooling layer and an average pooling layer, and then point-to-point spatial information is obtained through a 7×7 convolution layer. The SIGMOID function is used to activate the spatial information, and finally the spatial attention activation map is obtained; The fusion recognition module is configured to map the multi-scale spatio-temporal features and the optical flow data features to a feature fusion space, and classify based on the fused data after mapping to obtain a process recognition result.