Method and system for predicting particle size distribution of rock slag of full-face tunnel boring machine
Through multi-view noise enhancement and frequency domain bias decomposition technology, robust features are extracted from unlabeled rock slag images, solving the problem of predicting rock slag particle size distribution under harsh geological conditions, achieving high-precision particle size identification and parameter adjustment, and improving the safety and efficiency of tunnel boring machines.
Patent Information
- Application Number
- CN202510778204.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies are unable to effectively identify and predict the particle size distribution of rock debris from full-section tunnel boring machines under harsh geological conditions, resulting in increased tool wear and construction risks, and an inability to adjust tunneling parameters in a timely manner, affecting construction safety and efficiency.
A multi-view noise-enhanced contrast pre-training method and frequency domain bias decomposition technology are adopted to extract robust features from unlabeled rock slag images through a self-supervised learning model. The rock slag particle size distribution is fitted with the Rosin-Rammler curve to achieve high-precision particle size prediction.
It significantly reduces the workload of manual labeling, improves the robustness and prediction accuracy of the model, can accurately identify particle size distribution and reflect the mechanical properties of rock mass, helps adjust excavation parameters, and improves the automation and intelligence level of tunnel boring machines.
Smart Images

Figure CN120673415A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of intelligent construction and image recognition of tunnel boring machines, and in particular to a method and system for predicting the particle size distribution of rock slag of a full-section tunnel boring machine. Background Art
[0002] Tunnel boring machines (TBMs) are widely used in tunnel construction for transportation, water conservancy, and mining applications due to their advantages, including high excavation speeds, minimal disturbance to surrounding rock, good tunnel formation, and minimal impact on the environment. However, TBMs have poor adaptability to geological conditions, making excavation difficult and potentially posing significant construction risks in extremely harsh geological conditions, such as fault zones. During TBM excavation, the cutterhead and shield are in close contact with the tunnel face, preventing construction workers from observing the rock conditions ahead. Unintentional excavation in adverse geological conditions can lead to severe cutter wear and cutterhead deformation at best, and serious accidents such as machine jams, rockbursts, and tunnel collapses at worst. These can seriously threaten worker safety, delay construction progress, and cause significant economic losses. By monitoring the size of rock fragments, it is possible to infer the degree of cutter wear, detect adverse geological conditions, and adjust TBM excavation parameters accordingly, thereby optimizing support methods to improve excavation efficiency and safety.
[0003] Patent document CN117876776A discloses a TBM rock slag classification method based on a global attention mechanism and a residual neural network, comprising: collecting rock slag discharge images; constructing an image segmentation model and training it using the rock slag discharge images; using the trained image segmentation model to regularly segment the collected rock slag discharge images to obtain a number of local rock slag binary images; analyzing and preprocessing the several local rock slag binary images to obtain preprocessed rock slag binary images; calculating rock slag parameters on the rock slag binary images, and classifying the TBM rock slag.
[0004] Patent document CN116682010A discloses a real-time prediction method for surrounding rock classification based on TBM rock slag images, including: in the current cycle segment, TBM rock slag images are regularly acquired, and the surrounding rock grade classification results are determined using a trained surrounding rock classification prediction model; during the cycle segment, based on the judgment results of multiple rock slag images acquired in the current cycle segment, the probability Pi of the current cycle segment belonging to each type of surrounding rock is calculated in real time, and the surrounding rock classification corresponding to the highest probability is selected as the surrounding rock classification identification output result.
[0005] The aforementioned patent only considers rock slag identification and classification under normal ideal working conditions, but does not take into account the harsh on-site environment, and cannot fully utilize the large number of unlabeled rock slag images. However, the present invention uses a multi-view noise-enhanced contrast pre-training phase to learn robust feature representations of rock slag under harsh working conditions from a large number of unlabeled slag images. It also uses frequency domain transformation and Gaussian window mask filtering to achieve separation and spatial reconstruction of high- and low-frequency components, and uses jump splicing to transfer high-frequency features containing edge contour information, allowing the model to focus more on important features and achieve better recognition results. Summary of the Invention
[0006] In view of the defects in the prior art, the purpose of the present invention is to provide a method and system for predicting the particle size distribution of rock slag of a full-section tunnel boring machine.
[0007] According to the present invention, a method for predicting particle size distribution of rock slag from a full-face tunnel boring machine is provided, comprising:
[0008] Step S1: collecting an image of rock slag on a conveyor belt during tunnel boring machine construction, marking the rock slag contour to obtain a binary mask;
[0009] Step S2: Preprocess the slag image and binary mask to obtain the model input required for different tasks;
[0010] Step S3: Building a self-supervised learning model and pre-training it using an unlabeled dataset; building a semantic segmentation model based on the self-supervised learning model, loading the pre-trained weights and training it on the labeled dataset;
[0011] Step S4: Perform closed contour traversal and pixel statistics on the segmentation mask output by the semantic segmentation model, and fit the two key distribution parameters of the Rosin-Rammler curve based on the nonlinear least squares method;
[0012] Step S5: Evaluate the prediction effect of the slag particle size distribution based on the fitting results.
[0013] Preferably, the step S1 includes:
[0014] An image acquisition device is installed above the conveyor belt to obtain the original rock slag image covering the rock block area; the obtained image dataset is divided into a pre-training set, a training set, and a test set; the rock slag contours of the training set and test set images are manually annotated to obtain the corresponding binary masks.
[0015] Preferably, step S2 includes the following sub-steps:
[0016] Step S2.1: Preprocess the images in the pre-training set by randomly cropping them to 224*224 images from the original images and randomly adding one of four types of engineered noise for data augmentation, and converting them into tensors;
[0017] Step S2.2: Use
[0018]
[0019] Simulate the motion blur noise caused by camera dynamic blur caused by conveyor belt jitter, where (i, j) represents the point coordinates in the kernel function K, L is the kernel size, θ is the random rotation angle, Line(θ,L) represents the set of all points on a line defined by parameters θ and L, and (x,y) represents the pixel coordinates of image I;
[0020] use
[0021] FI illum (x,y)=I(x,y)·[1-α·exp(-((x c -x) 2 +(y c -y) 2 ))]
[0022] Simulates the uneven illumination noise caused by brightness gradient caused by local lighting loss or equipment occlusion, where (x c ,y c ) is the image center coordinate, α is the attenuation intensity, and (x, y) represents the pixel coordinate of image I;
[0023] use
[0024] I haze (x,y)=I(x,y)·(1-β)+A·β
[0025] Simulate rock dust suspension and dust mist noise caused by water spraying, where A = 255 is the atmospheric light intensity, β is the fog concentration, and (x, y) represents the pixel coordinates of image I;
[0026] use
[0027]
[0028] Simulate image compression artifact noise, where DCT represents discrete cosine transform, I represents the uncompressed image matrix to be transmitted, which is usually 8*8, round is the rounding function, and the quantization matrix Q is a diagonal matrix that determines the compression strength of each frequency component and is related to the compression quality factor q;
[0029] Step S2.3: Perform the same random cropping on the training set images and masks, interpolate to obtain 224*224 images and convert them into tensors;
[0030] Step S2.4: Center-crop the test set images and masks to obtain 1440 × 1440 pixel images. Add the four types of noise mentioned above to simulate the harsh working conditions on site to generate four additional test sets. Interpolate to obtain 224 × 224 images and convert them into tensors.
[0031] Preferably, step S3 includes the following sub-steps:
[0032] Step S3.1: Use the PyTorch framework to build a self-supervised learning model and perform multi-view noise enhancement contrast pre-training on a certain number of unlabeled image datasets to learn a robust feature representation of rock debris;
[0033] Step S3.2: Use the Pytorch framework to build a semantic segmentation model based on noise-resistant self-supervision combined with frequency domain bias decomposition, load pre-trained weights and train it on a manually annotated dataset.
[0034] Preferably, the step S3.1 includes the following sub-steps:
[0035] Step S3.1.1: Use the SimCLR framework-based contrastive self-supervised pre-training method with the core network architecture of ResNet34;
[0036] Step S3.1.2: Use a dual-branch network structure to generate positive sample pairs for the original rock slag image under multi-class noise data enhancement. By maximizing the similarity of different enhanced views of the same image, the driven model automatically captures the inherent semantic features of rock slag images.
[0037] Preferably, the step S3.2 includes the following sub-steps:
[0038] Step S3.2.1: Use fast Fourier transform to convert the input to the frequency domain, and then move the high-frequency area to the center of the spectrum;
[0039] Step S3.2.2: Generate a learnable Gaussian frequency domain mask using the following formula
[0040]
[0041] Among them, r uv Indicates the distance from the frequency domain coordinate (u, v) to the center, H, W represent the height and width of the image respectively, α c ,β w ,γ,b are learnable parameters, which are used to control the center position of the Gaussian mask, control the width of the Gaussian mask, balance the weights of high and low frequency components, and limit the effective frequency range. S represents the Sigmoid function, G low , G highand G represent the low-frequency component selection mask, high-frequency component selection mask, and enhancement mask of the image, respectively;
[0042] Step S3.2.3: Use
[0043]
[0044]
[0045] Where ⊙ represents matrix dot product; decompose the amplitude spectrum into high frequency M h With low frequency M l Component, input amplitude mapping network Enhanced features are obtained
[0046] Step S3.2.4: Use inverse Fourier transform to preserve the phase and inversely transform the high-frequency features and low-frequency features to achieve spatial feature reconstruction to obtain the high-frequency reconstructed image X h and low-frequency reconstructed image X l ;
[0047] Step S3.2.5: Set X h and X l Splicing to get X c And input the channel attention module, and obtain the attention weight value of each channel through global average pooling and two layers of linear transformation layers with relu as activation function;
[0048] Step S3.2.6: Use
[0049]
[0050] Map the weights to the [0,1] interval and compare the channel weight values with X c The second output is obtained by multiplying the channel feature map of
[0051] Preferably, step S4 includes the following sub-steps:
[0052] Step S4.1: Perform two rounds of opening operations on the mask with a kernel size of (3,3) to remove point noise that does not meet the preset value;
[0053] Step S4.2: Extract the closed contour C of the binary mask k , and for each C k Calculate its pixel area A k , and adopt
[0054]
[0055] Converted to equivalent diameter d k ;
[0056] Step S4.3: Using the Rosin-Rammler function distribution as the slag particle size distribution curve;
[0057]
[0058] Where d represents the input of the function, i.e., the slag particle size, d′ is the characteristic particle size, which means that 63.2% of the particles are smaller than this value, and n is the uniformity coefficient.
[0059] Step S4.4: Set the equivalent diameter set {d k}Arrange in ascending order and fit the two major parameters by nonlinear least squares method.
[0060] According to the present invention, a full-face tunnel boring machine rock slag particle size distribution prediction system is provided, comprising:
[0061] Module M1: Collect images of rock slag on the conveyor belt during tunnel boring machine construction, mark the rock slag contours and obtain a binary mask;
[0062] Module M2: Preprocesses the slag image and binary mask to obtain the model input required for different tasks;
[0063] Module M3: Build a self-supervised learning model and pre-train it using an unlabeled dataset; build a semantic segmentation model based on the self-supervised learning model, load the pre-trained weights, and train it on the labeled dataset;
[0064] Module M4: Performs closed contour traversal and pixel statistics on the segmentation mask output by the semantic segmentation model, and fits the two key distribution parameters of the Rosin-Rammler curve based on the nonlinear least squares method;
[0065] Module M5: Evaluate the prediction effect of slag particle size distribution based on the fitting results.
[0066] Preferably, the module M1 includes:
[0067] An image acquisition device is installed above the conveyor belt to obtain the original rock slag image covering the rock block area; the obtained image dataset is divided into a pre-training set, a training set, and a test set; the rock slag contours of the training set and test set images are manually annotated to obtain the corresponding binary masks.
[0068] Preferably, the module M2 includes the following submodules:
[0069] Module M2.1: Preprocess the images in the pre-training set by randomly cropping them to 224*224 images from the original images and randomly adding one of four types of engineered noise for data augmentation, converting them into tensors.
[0070] Module M2.2: using
[0071]
[0072]
[0073] Simulate the motion blur noise caused by camera dynamic blur caused by conveyor belt jitter, where (i, j) represents the point coordinates in the kernel function K, L is the kernel size, θ is the random rotation angle, Line(θ,L) represents the set of all points on a line defined by parameters θ and L, and (x,y) represents the pixel coordinates of image I;
[0074] use
[0075] FI illum (x,y)=I(x,y)·[1-α·exp(-((x c -x) 2 +(y c -y) 2 )0]
[0076] Simulates the uneven illumination noise caused by brightness gradient caused by local lighting loss or equipment occlusion, where (x c ,y c ) is the image center coordinate, α is the attenuation intensity, and (x, y) represents the pixel coordinate of image I;
[0077] use
[0078] I haze (x,y)=I(x,y)·(1-β)+A·β
[0079] Simulate rock dust suspension and dust mist noise caused by water spraying, where A = 255 is the atmospheric light intensity, β is the fog concentration, and (x, y) represents the pixel coordinates of image I;
[0080] use
[0081]
[0082] Simulate image compression artifact noise, where DCT represents discrete cosine transform, I represents the uncompressed image matrix to be transmitted, which is usually 8*8, round is the rounding function, and the quantization matrix Q is a diagonal matrix that determines the compression strength of each frequency component and is related to the compression quality factor q;
[0083] Module M2.3: Perform the same random cropping on the training set images and masks, interpolate to obtain 224*224 images and convert them into tensors;
[0084] Module M2.4: Center-crop the test set images and masks to obtain 1440*1440 pixel images. Add the four types of noise mentioned above to simulate the harsh working conditions on site to generate four additional test sets. Interpolate to obtain 224*224 images and convert them into tensors.
[0085] Preferably, the module M3 includes the following submodules:
[0086] Module M3.1: Use the PyTorch framework to build a self-supervised learning model and perform multi-view noise enhancement contrast pre-training on a certain number of unlabeled image datasets to learn robust feature representations of rock debris;
[0087] Module M3.2: Use the Pytorch framework to build a semantic segmentation model based on noise-resistant self-supervision combined with frequency domain bias decomposition, load pre-trained weights and train it on a manually annotated dataset.
[0088] Preferably, the module M3.1 includes the following submodules:
[0089] Module M3.1.1: A contrastive self-supervised pre-training method based on the SimCLR framework with the core network architecture of ResNet34;
[0090] Module M3.1.2: Using a dual-branch network structure, the original slag image is enhanced with multi-class noise data to generate positive sample pairs. By maximizing the similarity of different enhanced views of the same image, the driven model automatically captures the inherent semantic features of rock slag images.
[0091] Preferably, the module M3.2 includes the following submodules:
[0092] Module M3.2.1: Use fast Fourier transform to convert the input to the frequency domain and then move the high-frequency region to the center of the spectrum;
[0093] Module M3.2.2: Generate a learnable Gaussian frequency domain mask using the following formula
[0094]
[0095] G high =1-G low
[0096]
[0097] where r uv Indicates the distance from the frequency domain coordinate (u, v) to the center, H, W represent the height and width of the image respectively, α c ,β w,γ,b are learnable parameters, which are used to control the center position of the Gaussian mask, control the width of the Gaussian mask, balance the weights of high and low frequency components, and limit the effective frequency range. S represents the Sigmoid function, G low , G high and G represent the low-frequency component selection mask, high-frequency component selection mask, and enhancement mask of the image, respectively;
[0098] Module M3.2.3: Adoption
[0099]
[0100]
[0101] Where ⊙ represents the matrix dot product; decompose the amplitude spectrum into high frequency M h With low frequency M l Component, input amplitude mapping network Enhanced features are obtained
[0102] Module M3.2.4: Use inverse Fourier transform to preserve the phase and inversely transform the high-frequency and low-frequency features to achieve spatial feature reconstruction to obtain the high-frequency reconstructed image X h and low-frequency reconstructed image X l ;
[0103] Module M3.2.5: X h and X l Splicing to get X c And input the channel attention module, and obtain the attention weight value of each channel through global average pooling and two layers of linear transformation layers with relu as activation function;
[0104] Module M3.2.6: Adoption
[0105]
[0106] Map the weights to the [0,1] interval and compare the channel weight values with X c The second output is obtained by multiplying the channel feature map of
[0107] Preferably, the module M4 includes the following submodules:
[0108] Module M4.1: Perform two rounds of opening operations on the mask with a kernel size of (3,3) to remove point noise that does not meet the preset value;
[0109] Module M4.2: Extracting closed contours C from binary masks k , and for each C k Calculate its pixel area A k , and adopt
[0110]
[0111] Converted to equivalent diameter d k ;
[0112] Module M4.3: Use the Rosin-Rammler function distribution as the slag particle size distribution curve;
[0113]
[0114] Where d represents the input of the function, i.e., the slag particle size, d′ is the characteristic particle size, which means that 63.2% of the particles are smaller than this value, and n is the uniformity coefficient.
[0115] Module M4.4: Equivalent diameter set {d k}Arrange in ascending order and fit the two major parameters by nonlinear least squares method.
[0116] Compared with the prior art, the present invention has the following beneficial effects:
[0117] 1. This paper proposes a comparative pre-training method for multi-view noise enhancement and four classic engineering noise simulation methods. Through different data enhancement methods, robust feature representation of rock slag can be learned from a large number of unlabeled rock slag images, which greatly reduces the workload of manual labeling and increases the robustness of the prediction model.
[0118] 2. The present invention proposes a frequency domain bias decomposition method, which uses fast Fourier transform to convert feature maps into the frequency domain, uses trainable Gaussian masks to adaptively separate high-frequency edge components from multi-channel spatial features, reconstructs high-frequency images and performs jump connections to share key features, thereby enhancing the interpretability of the prediction model and greatly improving the prediction accuracy.
[0119] 3. This paper proposes a TBM rock slag particle size distribution prediction method based on noise-resistant self-supervised learning combined with bias frequency decomposition. This method can accurately identify the current particle size distribution curve from rock slag images. Only two parameters (characteristic particle size d′ and uniformity coefficient n) are required to fully describe the entire particle size distribution, significantly reducing data storage and transmission requirements and directly reflecting key information such as rock mechanical properties and TBM rock breaking efficiency. This method helps crews identify unfavorable geology, detect tool wear, and make timely adjustments to the excavation state, thereby improving the automation and intelligence level of tunnel boring machines. BRIEF DESCRIPTION OF THE DRAWINGS
[0120] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0121] Figure 1This is a flowchart of the implementation of the method for predicting the particle size distribution of rock slag of a full-section tunnel boring machine based on noise-resistant self-supervised learning combined with frequency domain bias decomposition proposed in the present invention.
[0122] Figure 2 This is a structural diagram of the noise-resistant self-supervised learning model proposed in the present invention.
[0123] Figure 3 This is a structural diagram of the frequency domain offset decomposition layer proposed in the present invention.
[0124] Figure 4 This is a structural diagram of the noise-resistant self-supervised learning combined with frequency domain bias decomposition method proposed in the present invention.
[0125] Figure 5 It is a rock slag image in the test set, the corresponding real mask, the mask predicted by the full-section tunnel boring machine rock slag particle size distribution prediction method based on noise-resistant self-supervised learning combined with frequency domain bias decomposition proposed in the present invention, and the corresponding contour map.
[0126] Figure 6 This is a comparison chart of the predicted distribution and actual distribution of a rock slag image in the test set under five environments using the full-section tunnel boring machine rock slag particle size distribution prediction method based on noise-resistant self-supervised learning combined with frequency domain bias decomposition proposed in the present invention. DETAILED DESCRIPTION
[0127] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0128] Example 1
[0129] Reference Figure 1 As shown in FIG, a method for predicting the particle size distribution of rock slag from a full-face tunnel boring machine based on noise-resistant self-supervised learning combined with frequency domain bias decomposition includes:
[0130] Step S1: Collecting images of rock slag on a conveyor belt during tunnel boring machine construction, and manually marking the rock slag contours to obtain a binary mask;
[0131] Specifically, in step S1:
[0132] An image acquisition device was installed above the conveyor belt to obtain a 2560 x 1440 pixel raw rock slag image covering the rock mass area. The resulting image dataset was divided into pre-training, training, and test sets. The rock slag contours in the training and test images were carefully manually annotated to obtain the corresponding binary masks.
[0133] Step S2: Preprocess the image and mask to obtain the model input required for different tasks;
[0134] Specifically, in step S2, the following is adopted:
[0135] Step S2.1: Preprocess the unlabeled pre-training set images and randomly crop them to 224*224 pixels using the transforms.RandomResizeCrop function in torchvision;
[0136] Step S2.2: Use
[0137]
[0138] Simulate the motion blur noise caused by camera dynamic blur caused by conveyor belt jitter, where (i, j) represents the point coordinates in the kernel function K, L is the kernel size, θ is the random rotation angle, Line(θ,L) represents the set of all points on a line defined by parameters θ and L, and (x,y) represents the pixel coordinates of image I;
[0139] use
[0140] FI illum (x,y)=I(x,y)·[1-α·exp(-((x c -x) 2 +(y c -y) 2 ))]
[0141] Simulates the uneven illumination noise caused by brightness gradient caused by local lighting loss or equipment occlusion, where (x c ,y c ) is the image center coordinate, α is the attenuation intensity, and (x, y) represents the pixel coordinate of image I;
[0142] use
[0143] I haze (x,y)=I(x,y)·(1-β)+A·β
[0144] Simulate rock dust suspension and dust mist noise caused by water spraying, where A = 255 is the atmospheric light intensity, β is the fog concentration, and (x, y) represents the pixel coordinates of image I;
[0145] use
[0146]
[0147] Simulate image compression artifact noise, where DCT represents discrete cosine transform, I represents the uncompressed image matrix to be transmitted, which is usually 8*8, round is the rounding function, and the quantization matrix Q is a diagonal matrix that determines the compression strength of each frequency component and is related to the compression quality factor q;
[0148] Randomly add one of the four types of engineering noise mentioned above to the cropped image for data augmentation, and then use the transforms.ToTensor function in torchvision to convert it into a tensor;
[0149] Step S2.3: Perform the same random cropping on the training set images and masks, then use the transforms.Resize function in torchvision to get 224*224 images and use the transforms.ToTensor function in torchvision to convert them into tensors;
[0150] Step S2.4: Use the transforms.CenterCrop function in torchvision to center-crop the test set images and masks to obtain 1440*1440 pixel images. Then use the transforms.Resize function in torchvision to obtain 224*224 pixels and convert them into tensors using the transforms.ToTensor function in torchvision.
[0151] Step S3: Refer to Figure 2 As shown in the figure, a self-supervised learning model is built using the Pytorch framework, and multi-view noise enhancement contrast pre-training is performed on a large number of unlabeled image datasets to learn robust feature representations of rock slag.
[0152] Specifically, in step S3:
[0153] Step S3.1: Use the torch.nn package under the Pytorch framework to build a comparative self-supervised pre-training model based on the SimCLR framework with the core network architecture of ResNet34;
[0154] Step S3.2: Use a dual-branch network structure to generate positive sample pairs for the original rock slag image under multi-class noise data enhancement. use
[0155]
[0156] Calculate the temperature scaled cross entropy loss. Where s i,i is the similarity of the positive sample pair. It is the sum of similarities of all pairs of samples. The temperature coefficient τ = 0.1 controls the sharpness of the feature distribution.
[0157] Step S4: Refer to Figure 3 and Figure 4 As shown in the figure, a semantic segmentation model based on noise-resistant self-supervision combined with frequency domain bias decomposition is built using the Pytorch framework, pre-trained weights are loaded, and training is performed on a manually annotated dataset;
[0158] Specifically, the process of frequency domain offset decomposition layer in step S4 adopts:
[0159] Step S4.1: Use fast Fourier transform to convert the input to the frequency domain, and then move the high-frequency area to the center of the spectrum;
[0160] Step S4.2: Use
[0161]
[0162] G high =1-G low
[0163]
[0164] where r uv Indicates the distance from the frequency domain coordinate (u, v) to the center, H, W represent the height and width of the image respectively, α c ,β w ,γ,b are learnable parameters, which are used to control the center position of the Gaussian mask, control the width of the Gaussian mask, balance the weights of high and low frequency components, and limit the effective frequency range. S represents the Sigmoid function, G low , G high and G represent the low-frequency component selection mask, high-frequency component selection mask, and enhancement mask of the image, respectively;
[0165] Generating learnable Gaussian frequency domain masks
[0166] Step S4.3: Use
[0167]
[0168]
[0169] Where ⊙ represents the matrix dot product; decompose the amplitude spectrum into high frequency M h With low frequency M l Component, input amplitude mapping network (3-layer CNN) enhanced features are obtained
[0170] Step S4.4: Use inverse Fourier transform to preserve the phase and inversely transform the high-frequency features and low-frequency features to achieve spatial feature reconstruction to obtain the high-frequency reconstructed image X h and low-frequency reconstructed image X l ;
[0171] Step S4.5: Set X h and X l Splicing to get X c And input the channel attention module, and obtain the attention weight value of each channel through global average pooling and two layers of linear transformation layers with relu as activation function;
[0172] Step S4.6: Use
[0173]
[0174] Map the weights to the [0,1] interval and compare the channel weight values with X c The second output is obtained by multiplying the channel feature map of
[0175] Step S5: Perform closed contour traversal and pixel statistics on the segmentation mask output by the model, and use the nonlinear least squares method to fit the two key distribution parameters of the Rosin-Rammler curve;
[0176] Specifically, the distribution fitting in step S5 adopts:
[0177] Step S5.1: Perform two rounds of opening operations on the mask with a kernel size of (3,3) to remove some tiny point noise;
[0178] Step S5.2: Extract the closed contour C of the binary mask k , and for each C k Calculate its pixel area A k , and adopt
[0179]
[0180] Converted to equivalent diameter d k ;
[0181] Step S5.3: using the Rosin-Rammler function distribution as the slag particle size distribution curve;
[0182]
[0183] Where d represents the input of the function, i.e., the slag particle size, d′ is the characteristic particle size, which means that 63.2% of the particles are smaller than this value, and n is the uniformity coefficient.
[0184] Step S5.4: Set the equivalent diameter set {d k}Arrange in ascending order and fit the two major parameters by nonlinear least squares method.
[0185] Step S6: Evaluate the prediction effect of the slag particle size distribution based on the true distribution curve of the test set and the model prediction curve.
[0186] Specifically, in step S6:
[0187] All images in the test set are input into the model to obtain the predicted mask map of semantic segmentation. The two parameters of the Rosin-Rammler function fitted by the mask map are compared with the parameters of the manually labeled mask map.
[0188]
[0189] where MAPE dmax Represents the maximum particle size error, MAPE d′ and MAPE n represents the error in the distribution parameters, Indicates the maximum particle size value predicted by the model, Indicates the true maximum particle size value manually marked, d . ' pred , d′ true , n pred , n true Same thing.
[0190] Calculate the error between the predicted distribution and the true distribution to evaluate the model prediction effect.
[0191] Example 2
[0192] A full-face tunnel boring machine rock slag particle size distribution prediction system, comprising:
[0193] Module M1: Collect images of rock slag on the conveyor belt during tunnel boring machine construction and manually mark the rock slag contours to obtain a binary mask;
[0194] Specifically, in the module M1:
[0195] An image acquisition device was installed above the conveyor belt to obtain a 2560 x 1440 pixel raw rock slag image covering the rock mass area. The resulting image dataset was divided into pre-training, training, and test sets. The rock slag contours in the training and test images were carefully manually annotated to obtain the corresponding binary masks.
[0196] Module M2: Preprocess images and masks to obtain model inputs required for different tasks;
[0197] Specifically, in the module M2:
[0198] Module M2.1: Preprocess the unlabeled pre-training set images and randomly crop them to 224*224 pixels using the transforms.RandomResizeCrop function in torchvision;
[0199] Module M2.2: using
[0200]
[0201] Simulate the motion blur noise caused by camera dynamic blur caused by conveyor belt jitter, where (i, j) represents the point coordinates in the kernel function K, L is the kernel size, θ is the random rotation angle, Line(θ,L) represents the set of all points on a line defined by parameters θ and L, and (x,y) represents the pixel coordinates of image I;
[0202] use
[0203] FI illum (x,y)=I(x,y)·[1-α·exp(-((x c -x) 2 +(y c -y) 2 ))]
[0204] Simulates the uneven illumination noise caused by brightness gradient caused by local lighting loss or equipment occlusion, where (x c ,y c ) is the image center coordinate, α is the attenuation intensity, and (x, y) represents the pixel coordinate of image I;
[0205] use
[0206] I haze (x,y)=I(x,y)·(1-β)+A·β
[0207] Simulate rock dust suspension and dust mist noise caused by water spraying, where A = 255 is the atmospheric light intensity, β is the fog concentration, and (x, y) represents the pixel coordinates of image I;
[0208] use
[0209]
[0210] Simulate image compression artifact noise, where DCT represents discrete cosine transform, I represents the uncompressed image matrix to be transmitted, which is usually 8*8, round is the rounding function, and the quantization matrix Q is a diagonal matrix that determines the compression strength of each frequency component and is related to the compression quality factor q;
[0211] Randomly add one of the four types of engineering noise mentioned above to the cropped image for data augmentation, and then use the transforms.ToTensor function in torchvision to convert it into a tensor;
[0212] Module M2.3: Perform the same random cropping on the training set images and masks, then use the transforms.Resize function in torchvision to get 224*224 images and use the transforms.ToTensor function in torchvision to convert them into tensors;
[0213] Module M2.4: Use the transforms.CenterCrop function in torchvision to center-crop the test set images and masks to obtain 1440*1440 pixel images, then use the transforms.Resize function in torchvision to obtain 224*224 images and convert them into tensors using the transforms.ToTensor function in torchvision;
[0214] Module M3: Use the PyTorch framework to build a self-supervised learning model and perform multi-view noise enhancement comparative pre-training on a large number of unlabeled image datasets to learn robust feature representations of rock debris;
[0215] Specifically, in the module M3:
[0216] Module M3.1: Use the torch.nn package under the Pytorch framework to build a comparative self-supervised pre-training model based on the SimCLR framework with the core network architecture of ResNet34;
[0217] Module M3.2: Using a dual-branch network structure, the original slag image is enhanced with multi-class noise data to generate positive sample pairs. use
[0218]
[0219] where s i,i is the similarity of the positive sample pair. It is the sum of similarities of all pairs of samples. The temperature coefficient τ = 0.1 controls the sharpness of the feature distribution.
[0220] Calculate temperature-scaled cross entropy loss;
[0221] Module M4: Use the Pytorch framework to build a semantic segmentation model based on noise-resistant self-supervision combined with frequency domain bias decomposition, load pre-trained weights and train it on a manually annotated dataset;
[0222] Specifically, the process of the frequency domain offset decomposition layer in the module M4 adopts:
[0223] Module M4.1: Use fast Fourier transform to convert the input to the frequency domain and then move the high-frequency area to the center of the spectrum;
[0224] Module M4.2: Adoption
[0225]
[0226] G high =1-G low
[0227]
[0228] where r uv Indicates the distance from the frequency domain coordinate (u, v) to the center, H, W represent the height and width of the image respectively, α c ,β w ,γ,b are learnable parameters, which are used to control the center position of the Gaussian mask, control the width of the Gaussian mask, balance the weights of high and low frequency components, and limit the effective frequency range. S represents the Sigmoid function, G low , G high and G represent the low-frequency component selection mask, high-frequency component selection mask, and enhancement mask of the image, respectively;
[0229] Generating learnable Gaussian frequency domain masks
[0230] Module M4.3: Adoption
[0231]
[0232]
[0233] Where ⊙ represents the matrix dot product; decompose the amplitude spectrum into high frequency M h With low frequency M l Component, input amplitude mapping network (3-layer CNN) enhanced features are obtained
[0234] Module M4.4: Use inverse Fourier transform to preserve the phase and inversely transform the high-frequency and low-frequency features to achieve spatial feature reconstruction to obtain the high-frequency reconstructed image X h and low-frequency reconstructed image X l ;
[0235] Module M4.5: X h and X l Splicing to get X cAnd input the channel attention module, and obtain the attention weight value of each channel through global average pooling and two layers of linear transformation layers with relu as activation function;
[0236] Module M4.6: Adopt
[0237]
[0238] Map the weights to the [0,1] interval and compare the channel weight values with X c The second output is obtained by multiplying the channel feature map of
[0239] Module M5: Perform closed contour traversal and pixel statistics on the segmentation mask output by the model, and use the nonlinear least squares method to fit the two key distribution parameters of the Rosin-Rammler curve;
[0240] Specifically, the distribution fitting in the module M5 adopts:
[0241] Module M5.1: Perform two rounds of opening operations on the mask with a kernel size of (3,3) to remove some tiny point noise;
[0242] Module M5.2: Extracting closed contours C from binary masks k , and for each C k Calculate its pixel area A k , and adopt
[0243]
[0244] Converted to equivalent diameter d k ;
[0245] Module M5.3: Use the Rosin-Rammler function distribution as the slag particle size distribution curve;
[0246]
[0247] Where d represents the input of the function, i.e., the slag particle size, d′ is the characteristic particle size, which means that 63.2% of the particles are smaller than this value, and n is the uniformity coefficient.
[0248] Module M5.4: Equivalent diameter set {d k}Arrange in ascending order and fit the two major parameters by nonlinear least squares method.
[0249] Module M6: Evaluate the prediction effect of slag particle size distribution based on the true distribution curve of the test set and the model prediction curve.
[0250] Specifically, in the module M6:
[0251] All images in the test set are input into the model to obtain the predicted mask map of semantic segmentation. The two parameters of the Rosin-Rammler function fitted by the mask map are compared with the parameters of the manually labeled mask map.
[0252]
[0253] in Represents the maximum particle size error, MAPE d′ and MAPE n Represents the error in the distribution parameters.
[0254] Calculate the error between the predicted distribution and the true distribution to evaluate the model prediction effect.
[0255] from Figure 5 Figure 6 It can be seen that the proposed full-face tunnel boring machine slag particle size distribution prediction model based on noise-resistant self-supervised learning combined with frequency domain bias decomposition is very accurate. When the pre-training samples are 7000, the training samples are 900, and the test samples are 100, MAPE is 6.7% d′ MAPE is 11.8% n The proposed method for predicting the particle size distribution of rock slag from full-face tunnel boring machines based on noise-resistant self-supervised learning combined with frequency domain bias decomposition has a high recognition accuracy. MAPE is 8.39% d′ MAPE is 12.65% n In the uneven illumination noise test set, MAPE is 9.16% d′ MAPE is 14.21% n In the water mist and dust noise test set, MAPE is 7.67% d′ MAPE is 12.83% n is 18.51%. In the image compression noise test set, MAPE is 6.76% d′ MAPE is 11.80% n The results show that the proposed method for predicting the particle size distribution of rock slag in full-face tunnel boring machines based on noise-resistant self-supervised learning combined with frequency domain bias decomposition still has high recognition accuracy and strong robustness in harsh environments.
[0256] Those skilled in the art will appreciate that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, and units for implementing various functions can also be considered as both software modules implementing the method and structures within the hardware component.
[0257] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.
Claims
1. A method for predicting particle size distribution of rock slag from a full-face tunnel boring machine, characterized in that: include: Step S1: collecting an image of rock slag on a conveyor belt during tunnel boring machine construction, marking the rock slag contour to obtain a binary mask; Step S2: Preprocess the slag image and binary mask to obtain the model input required for different tasks; Step S3: Building a self-supervised learning model and pre-training it using an unlabeled dataset; building a semantic segmentation model based on the self-supervised learning model, loading the pre-trained weights and training it on the labeled dataset; Step S4: Perform closed contour traversal and pixel statistics on the segmentation mask output by the semantic segmentation model, and fit the two key distribution parameters of the Rosin-Rammler curve based on the nonlinear least squares method; Step S5: Evaluate the prediction effect of the slag particle size distribution based on the fitting results.
2. The method for predicting the particle size distribution of rock slag of a full-face tunnel boring machine according to claim 1, characterized in that: The step S1 comprises: An image acquisition device is installed above the conveyor belt to obtain the original rock slag image covering the rock block area; the obtained image dataset is divided into a pre-training set, a training set, and a test set; the rock slag contours of the training set and test set images are manually annotated to obtain the corresponding binary masks.
3. The method for predicting particle size distribution of rock slag of a full-face tunnel boring machine according to claim 2, characterized in that: The step S2 includes the following sub-steps: Step S2.1: Preprocess the images in the pre-training set by randomly cropping them to 224*224 images from the original images and randomly adding one of four types of engineered noise for data augmentation, and converting them into tensors; Step S2.2: Use Simulate the motion blur noise caused by camera dynamic blur caused by conveyor belt jitter, where (i, j) represents the point coordinates in the kernel function K, L is the kernel size, θ is the random rotation angle, Line(θ,L) represents the set of all points on a line defined by parameters θ and L, and (x,y) represents the pixel coordinates of image I; use FI illum (x,y)=I(x,y)·[1-α·exp(-((x c -x) 2 +(and c -and) 2 ))] Simulates the uneven illumination noise caused by brightness gradient caused by local lighting loss or equipment occlusion, where (x c ,y c ) is the image center coordinate, α is the attenuation intensity, and (x, y) represents the pixel coordinate of image I; use I haze (x,y)=I(x,y)·(1-β)+A·β Simulate rock dust suspension and dust mist noise caused by water spraying, where A = 255 is the atmospheric light intensity, β is the fog concentration, and (x, y) represents the pixel coordinates of image I; use Simulate image compression artifact noise, where DCT represents discrete cosine transform, I represents the uncompressed image matrix to be transmitted, round represents the integer function, and the quantization matrix Q is a diagonal matrix that determines the compression strength of each frequency component and is related to the compression quality factor q; Step S2.3: Perform the same random cropping on the training set images and masks, interpolate to obtain 224*224 images and convert them into tensors; Step S2.4: Center-crop the test set images and masks to obtain 1440 × 1440 pixel images. Add the four types of noise mentioned above to simulate the harsh working conditions on site to generate four additional test sets. Interpolate to obtain 224 × 224 images and convert them into tensors.
4. The method for predicting the particle size distribution of rock slag of a full-face tunnel boring machine according to claim 1, characterized in that: The step S3 includes the following sub-steps: Step S3.1: Use the PyTorch framework to build a self-supervised learning model and perform multi-view noise enhancement contrast pre-training on a certain number of unlabeled image datasets to learn a robust feature representation of rock debris; Step S3.2: Use the Pytorch framework to build a semantic segmentation model based on noise-resistant self-supervision combined with frequency domain bias decomposition, load pre-trained weights and train it on a manually annotated dataset.
5. The method for predicting particle size distribution of rock slag of a full-face tunnel boring machine according to claim 4, characterized in that: The step S3.1 includes the following sub-steps: Step S3.1.1: Use the SimCLR framework-based contrastive self-supervised pre-training method with the core network architecture of ResNet34; Step S3.1.2: Use a dual-branch network structure to generate positive sample pairs for the original rock slag image under multi-class noise data enhancement. By maximizing the similarity of different enhanced views of the same image, the driven model automatically captures the inherent semantic features of rock slag images.
6. The method for predicting particle size distribution of rock slag for a full-face tunnel boring machine according to claim 4, characterized in that: The step S3.2 includes the following sub-steps: Step S3.2.1: Use fast Fourier transform to convert the input to the frequency domain, and then move the high-frequency area to the center of the spectrum; Step S3.2.2: Generate a learnable Gaussian frequency domain mask using the following formula G high =1-G low where r uv Indicates the distance from the frequency domain coordinate (u, v) to the center, H, W represent the height and width of the image respectively, α c ,β w ,γ,b are learnable parameters, which are used to control the center position of the Gaussian mask, control the width of the Gaussian mask, balance the weights of high and low frequency components, and limit the effective frequency range. S represents the Sigmoid function, G low , G high and G represent the low-frequency component selection mask, high-frequency component selection mask, and enhancement mask of the image, respectively; Step S3.2.3: Use Where ⊙ represents the matrix dot product; decompose the amplitude spectrum into high frequency M h With low frequency M l Component, input amplitude mapping network Enhanced features are obtained Step S3.2.4: Use inverse Fourier transform to preserve the phase and inversely transform the high-frequency features and low-frequency features to achieve spatial feature reconstruction to obtain the high-frequency reconstructed image X j and low-frequency reconstructed image X l ; Step S3.2.5: Set X h and X l Splicing to get X c And input the channel attention module, and obtain the attention weight value of each channel through global average pooling and two layers of linear transformation layers with relu as activation function; Step S3.2.6: Use Map the weights to the [0,1] interval and compare the channel weight values with X c The second output is obtained by multiplying the channel feature map of 7. The method for predicting particle size distribution of rock slag for a full-face tunnel boring machine according to claim 1, characterized in that: The step S4 includes the following sub-steps: Step S4.1: Perform two rounds of opening operations on the mask with a kernel size of (3,3) to remove point noise that does not meet the preset value; Step S4.2: Extract the closed contour C of the binary mask k , and for each C k Calculate its pixel area A k , and adopt Converted to equivalent diameter d k ; Step S4.3: Using the Rosin-Rammler function distribution as the slag particle size distribution curve; Among them, d represents the input of the function, that is, the particle size of the slag, d ′ is the characteristic particle size, which means that 63.2% of the particles are smaller than this value, and n is the uniformity coefficient; Step S4.4: Set the equivalent diameter set {d k }Arrange in ascending order and fit the two major parameters by nonlinear least squares method.
8. A full-face tunnel boring machine slag particle size distribution prediction system, characterized in that: include: Module M1: Collect images of rock slag on the conveyor belt during tunnel boring machine construction, mark the rock slag contours and obtain a binary mask; Module M2: Preprocesses the slag image and binary mask to obtain the model input required for different tasks; Module M3: Build a self-supervised learning model and pre-train it using an unlabeled dataset; build a semantic segmentation model based on the self-supervised learning model, load the pre-trained weights, and train it on the labeled dataset; Module M4: Performs closed contour traversal and pixel statistics on the segmentation mask output by the semantic segmentation model, and fits the two key distribution parameters of the Rosin-Rammler curve based on the nonlinear least squares method; Module M5: Evaluate the prediction effect of slag particle size distribution based on the fitting results.
9. The full-face tunnel boring machine slag particle size distribution prediction system according to claim 8, characterized in that: The module M1 includes: An image acquisition device is installed above the conveyor belt to obtain the original rock slag image covering the rock block area; the obtained image dataset is divided into a pre-training set, a training set, and a test set; the rock slag contours of the training set and test set images are manually annotated to obtain the corresponding binary masks.
10. The full-face tunnel boring machine slag particle size distribution prediction system according to claim 9, characterized in that: The module M2 Includes the following submodules: Module M2.1: Preprocess the images in the pre-training set by randomly cropping them to 224*224 images from the original images and randomly adding one of four types of engineered noise for data augmentation, converting them into tensors. Module M2.2: using Simulate the motion blur noise caused by camera dynamic blur caused by conveyor belt jitter, where (i, j) represents the point coordinates in the kernel function K, L is the kernel size, θ is the random rotation angle, Line(θ,L) represents the set of all points on a line defined by parameters θ and L, and (x,y) represents the pixel coordinates of image I; use FI illum (x,y)=I(x,y)·[1-α·exp(-((x c -x) 2 +(and c -and) 2 ))] Simulates the uneven illumination noise caused by brightness gradient caused by local lighting loss or equipment occlusion, where (x c ,y c ) is the image center coordinate, α is the attenuation intensity, and (x, y) represents the pixel coordinate of image I; use I haze (x,y)=I(x,y)·(1-β)+A·β Simulate rock dust suspension and dust mist noise caused by water spraying, where A = 255 is the atmospheric light intensity, β is the fog concentration, and (x, y) represents the pixel coordinates of image I; use Simulate image compression artifact noise, where DCT represents discrete cosine transform, I represents the uncompressed image matrix to be transmitted, round represents the integer function, and the quantization matrix Q is a diagonal matrix that determines the compression strength of each frequency component and is related to the compression quality factor q; Module M2.3: Perform the same random cropping on the training set images and masks, interpolate to obtain 224*224 images and convert them into tensors; Module M2.4: Center-crop the test set images and masks to obtain 1440*1440 pixel images. Add the four types of noise mentioned above to simulate the harsh working conditions on site to generate four additional test sets. Interpolate to obtain 224*224 images and convert them into tensors.
Citation Information
Patent Citations
Real-time surrounding rock classification prediction method based on TBM rock slag image
CN116682010A
TBM rock slag classification method based on global attention mechanism and residual neural network
CN117876776A