Shield tunneling machine cutter state fault prediction system based on multi-modal data fusion
Through multimodal data fusion and improved network structure, the problem of insufficient adaptability of the tool status fault prediction system of the shield machine is solved, and more efficient fault prediction and adaptability is achieved, and the accuracy and real-timeness of the tool status monitoring of the shield machine is improved.
Patent Information
- Application Number
- CN202510664581.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing shield machine tool status fault prediction system lacks adaptability and cannot effectively deal with data differences in different geological conditions and shield machine models, resulting in poor generalization capabilities.
The multimodal data fusion method is adopted to achieve cross-domain feature extraction and fault prediction through low-rank matrix decomposition and compression, and improve the lightweight weight migration network and space-time cross-attention dynamic migration network, combining singular value decomposition, improve the lightweight weight migration network and adversarial training feature alignment algorithm.
It improves the accuracy and real-time prediction of tool state faults of the shield machine, enhances the generalization ability of the model, can adapt to different shield machine and working conditions, and reduces the computational complexity.
Smart Images

Figure CN120277615A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of shield machine fault prediction, and particularly to a shield machine tool state fault prediction system based on multi-modal data fusion. Background Art
[0002] The application of shield tunneling method in tunnel engineering is extensive. As the core equipment of shield tunneling method, the state of the cutter of the shield machine directly affects the construction progress, cost and safety. However, the cutter of the shield machine works under complex geological conditions and faces problems such as high pressure and high wear, resulting in frequent cutter failures.
[0003] At present, the existing shield machine tool state fault prediction systems are based on fixed feature extraction methods and model structures, lacking the adaptive ability to local features of data. Facing the data differences of different geological conditions and different shield machine models, the generalization ability of these models is poor and they cannot be effectively applied to various actual engineering scenarios. Therefore, a shield machine tool state fault prediction system based on multi-modal data fusion is proposed here. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art and achieve the above object, the present invention proposes the following technical solutions:
[0005] A shield machine tool state fault prediction system based on multi-modal data fusion, comprising:
[0006] A state acquisition module: collecting multi-modal state data, decomposing the multi-modal state data into time and space related features, and obtaining the optimal state data through low-rank matrix decomposition compression;
[0007] Among them, the low-rank matrix decomposition compression is implemented based on a low-rank matrix approximation algorithm improved by singular value decomposition;
[0008] An anomaly simulation module: collecting fault state data, constructing an improved lightweight weight transfer network for cross-domain feature extraction to obtain cross-domain common feature vectors, and obtaining common features through a feature alignment algorithm based on adversarial training based on the cross-domain common feature vectors and the fault state data;
[0009] Among them, the improved lightweight weight transfer network improves the traditional convolution kernel, introduces an additional branch network in the convolutional layer, takes the input feature map of the current convolutional layer as the input, outputs control parameters through convolution and fully connected operations, and adjusts the weights of the convolution kernel according to the values of the control parameters;
[0010] A fault prediction module: performing fault prediction through a spatio-temporal cross-attention dynamic migration network based on the obtained optimal state data and common features;
[0011] Among them, the spatio-temporal cross-attention dynamic migration network includes an improved dynamic sparsification strategy and a segmented causal convolution trigger mechanism.
[0012] The process of obtaining time-related features is as follows:
[0013] Collect multi-source sensor data. For the key time-related features in the multi-source sensor data, obtain the cutter head rotation speed signal through an encoder, and use the fast Fourier transform to convert the rotation speed signal in the time domain to the frequency domain to obtain time-related features;
[0014] The process of obtaining space-related features is as follows:
[0015] Adopt the finite element analysis method to perform mechanical modeling on the shield machine cutter and cutter head to obtain the stress distribution data of the cutter under different working conditions. Use a high-definition industrial camera to take periodic photos of the cutter surface to obtain the cutter surface image. Adopt an image segmentation algorithm to obtain the area of the wear area and the shape features of the wear area. Based on the stress distribution data, the area of the wear area, and the shape features of the wear area, obtain space-related features.
[0016] The implementation process of the singular value decomposition-improved low-rank matrix approximation algorithm is as follows:
[0017] Integrate the extracted time-related features and space-related features to construct a feature matrix. Among them, the feature matrix includes a left singular matrix, a diagonal matrix, and a right singular matrix;
[0018] According to the energy proportion of the singular values in the feature matrix, retain the first k larger singular values to construct an approximation matrix;
[0019] Set an energy proportion threshold, and accumulate a fixed number of points from large to small according to the energy proportion of the singular values in the feature matrix until the accumulated sum exceeds the energy proportion threshold, and determine that the number of retained singular values is the optimal number of singular values
[0020] Based on the optimal number of singular values Select columns from the left singular matrix, select the first diagonal elements from the diagonal matrix, and select the first rows from the right singular matrix to obtain the optimal approximation matrix;
[0021] The data points corresponding to the optimal approximation matrix are the optimal state data.
[0022] The design process of the branch network structure in the improved lightweight weight migration network is as follows:
[0023] The first layer of the branch network is a convolutional layer. After passing through this convolutional layer, a feature map is obtained, followed by a global average pooling layer, and the global average pooling layer compresses the feature map;
[0024] Connect a fully connected layer to the compression result of the global average pooling layer to obtain an intermediate vector, and then pass through another fully connected layer to output control parameters, and adjust the weights of the convolutional kernels based on the control parameters.
[0025] The implementation process of adjusting the weights of the convolutional kernels according to the control parameters is as follows:
[0026] Based on the weight matrix of the traditional convolutional kernel, divide the convolutional kernel weight matrix into multiple sub-regions through an interpolation algorithm, and calculate the interpolation coefficients for each sub-region according to the value of the control parameter;
[0027] Adjust the weights of each sub-region according to the interpolation coefficients to obtain an adjusted weight matrix;
[0028] When performing convolution operations, use the adjusted weight matrix to perform convolution calculations on the input feature map to achieve the adjustment of the convolutional kernel weights.
[0029] The process of obtaining the common features is as follows:
[0030] The improved lightweight weight transfer network consists of two sub-networks with the same structure but non-shared weights;
[0031] Input the fault state data into one of the sub-networks of the constructed improved lightweight weight transfer network. After the forward propagation of the network, extract the first feature vector of the image, perform normalization processing on the first feature vector, input the normalized first feature vector into another sub-network, extract the second feature vector, and splice the first feature vector and the second feature vector to obtain a cross-domain common feature vector;
[0032] Design a feature alignment algorithm based on adversarial training, introduce a discriminator network, with the input being the fault state data and the cross-domain common feature vector, and the output being the probability of judging whether the two come from the same distribution;
[0033] Introduce an adversarial loss term, and update the weights of the improved lightweight weight transfer network through the backpropagation algorithm of the adversarial loss term, making it difficult for the discriminator network to distinguish between the fault state data and the cross-domain common feature vector, and obtaining the common features.
[0034] The spatio-temporal cross-attention dynamic transfer network consists of an input layer, an improved dynamic sparsification strategy, a segmented causal convolution trigger mechanism, a fully connected layer, and a classifier.
[0035] The implementation process of the improved dynamic sparsification strategy is as follows:
[0036] The output of the improved dynamic sparsification strategy is a sparsified attention weight matrix;
[0037] Determine the top 3 key spatial nodes closely related to tool failures based on common features, and achieve sparsification by dynamically adjusting the mask matrix through a predefined mask matrix;
[0038] The process of dynamically adjusting the mask matrix is as follows:
[0039] Based on historical data and a machine learning regression model, initially obtain the complexity of the working conditions between the optimal state data and common features, and use the optimal state data and common features as input features;
[0040] Use the working condition complexity obtained by pre-artificial calibration as a label to train the regression model, enabling the model to learn the mapping relationship between the input features and the working condition complexity, and output the working condition complexity index;
[0041] Based on the working condition complexity index, construct a mapping function. Set different basic mask matrices according to the tool working stage, and adjust the basic mask matrix template through the mapping function to change the mask values in the area around the key nodes to obtain a sparsified attention weight matrix.
[0042] The implementation process of the segmented causal convolution trigger mechanism is as follows:
[0043] Real-time monitor the variance of the optimal state data to judge the change of the tool running state. Preset the variance mutation threshold, and obtain the change of the tool running state based on the variance mutation threshold;
[0044] When the variance of the vibration signal mutates, enable the deep dilated convolution to output the state output sequence;
[0045] Among them, the deep dilated convolution is achieved by increasing the interval of the convolution kernel;
[0046] When the variance of the vibration signal does not mutate, use one-dimensional convolution for shallow calculation to output the state output sequence.
[0047] The implementation process of the fault prediction is as follows:
[0048] Input the sparsified attention weight matrix and the state output sequence output by the segmented causal convolution into the fully connected layer at the same time. The fully connected layer inputs the sparsified attention weight matrix and the state output sequence output by the segmented causal convolution for abstract mapping, and through linear transformation and activation function operations, convert the input feature vectors into high-level feature representations;
[0049] Input the high-level feature representation into the classifier. Divide the tool failure types into G categories. The classifier calculates the probability of each category, sets the fault threshold, and among the various categories output by the classifier, the situation where the category probability is less than or equal to the fault threshold is in a fault state.
[0050] The present invention has the following beneficial effects:
[0051] In the present invention, firstly, by collecting multi-modal state data of the shield machine cutter, such as time and space related features such as cutter head rotation speed, cutter stress distribution, wear area characteristics, etc., and fusing and processing them, the operating state of the cutter can be comprehensively reflected, avoiding the limitations brought by single data;
[0052] Secondly, by using the improved lightweight weight transfer network and adaptive receptive field convolutional kernel, the receptive field size of the convolutional kernel can be dynamically adjusted according to the local features of the input data, and the effective features of different regions can be extracted more accurately. At the same time, the feature alignment algorithm based on adversarial training realizes the effective fusion of cross-domain features, enhances the generalization ability of the model, and enables it to better adapt to different shield machines and working conditions;
[0053] Finally, by improving the dynamic sparsification strategy, the sparsification is achieved by determining the key spatial nodes and dynamically adjusting the mask matrix, reducing the computational complexity and avoiding the interference of irrelevant information. The segmented causal convolution trigger mechanism intelligently selects deep dilated convolution or shallow calculation according to the variance mutation of the cutter vibration signal, which can capture long sequence features when the cutter is abnormal and reduce the computational amount when it is normal, improving the real-time performance and accuracy of fault prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a system block diagram of a shield machine cutter state fault prediction system with multi-modal data fusion proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0056] Embodiment 1
[0057] As Figure 1 shown, a shield machine cutter state fault prediction system with multi-modal data fusion proposed by the present invention includes:
[0058] State acquisition module: Collect multi-modal state data, collect time and space related features of the shield machine multi-modal state data, and obtain the optimal state data through low-rank matrix decomposition compression;
[0059] Using high-precision sensors to monitor the operating status of the shield machine's cutting tools and collect multi-source sensor data. Obtain the cutter head rotation speed signal R(t) through the encoder installed on the cutter head drive shaft, and use the Fast Fourier Transform (FFT) to convert the rotation speed signal in the time domain to the frequency domain to obtain the time-related feature [R].
[0060] Specifically, the cutter head rotation speed signal R(t) is a data that changes with time t, which reflects the rotation speed of the shield machine's cutter head at different times. Its value will change dynamically over time, showing obvious time series characteristics. For example, during the tunneling process of the shield machine, as factors such as geological conditions change and the load changes, the cutter head rotation speed will be adjusted in real time. This changing trend over time is the best manifestation of the time-related feature.
[0061] Adopt the finite element analysis method to conduct mechanical modeling on the shield machine's cutting tools and cutter head. According to the geometric shape and material properties (elastic modulus E, Poisson's ratio v) of the cutting tools, directly obtain the stress distribution data S1 (concentrated area and stress magnitude) of the cutting tools under different working conditions. At the same time, use the high-definition industrial camera installed near the cutter head to take periodic photos of the cutting tool surface, obtain the cutting tool surface image, and adopt the image segmentation algorithm to accurately segment the worn area and the non-worn area in the cutting tool image, directly obtaining the area S2 of the worn area and the shape feature S3 of the worn area. These features can accurately describe the spatial distribution of the cutting tool wear, forming the space-related feature [S].
[0062] Adopt the improved low-rank matrix approximation algorithm based on Singular Value Decomposition (SVD) to perform singular value decomposition on the feature matrix X to obtain the optimal state data [X]. The implementation process is as follows:
[0063] Integrate the extracted time-related feature [R] and space-related feature [S] to construct the feature matrix X. Among them, the feature matrix includes the left singular matrix, the diagonal matrix, and the right singular matrix.
[0064] Assume that there are m time-related features and n space-related features, and the sampling time points are N. The dimension of the feature matrix X is N×(m + n).
[0065] Feature matrix X = UΣV T , where U is the left singular matrix, Σ is the diagonal matrix, its diagonal elements are singular values, V is the right singular matrix, and T represents the transpose.
[0066] According to the energy proportion of the singular values in the feature matrix X, retain the first k larger singular values (k << min(N, m + n)) to construct the approximate matrix Specifically, calculate the energy proportion of the singular values:
[0067]
[0068] Among them, σ j is the j-th singular value after the singular value decomposition of the feature matrix X, represents the sum of the squares of the first min(N, m + n) singular values;
[0069] Set an energy ratio threshold θ (0.95), and accumulate a fixed point number 0.1 from large to small for the energy ratio of the singular values in the feature matrix until the accumulated sum exceeds θ. At this time, the number of singular values to be retained is determined as the optimal number of singular values
[0070] Determine to retain After the number of singular values, select the first columns from the left singular matrix U to form a new matrix Select the first diagonal elements from the diagonal matrix Σ to form a diagonal matrix Select the first rows from the right singular matrix V to form the optimal approximation matrix
[0071] Specifically, this optimal approximation matrix is a low-rank approximation of the original feature matrix X. Since only the optimal number of singular values larger singular values and their corresponding left and right singular vector parts are retained, the redundant information related to the smaller singular values is removed, reducing the rank of the matrix;
[0072] This new matrix is the optimal matrix obtained after low-rank matrix decomposition compression. The corresponding data points are the optimal state data [X]. While retaining the main feature information of the original feature matrix X (guaranteed by the energy ratio threshold θ), it retains the best part of the data, reduces the overall dimension, and facilitates the subsequent shield machine tool state fault prediction and analysis tasks.
[0073] Abnormal simulation module: Collect fault state data, construct an improved lightweight weight transfer network for cross-domain feature extraction to obtain cross-domain common feature vectors, and obtain common features based on the cross-domain common feature vectors and fault state data through a feature alignment algorithm based on adversarial training;
[0074] Collect historical fault data from different shield machines and other related construction machinery (roadheaders). These data include equipment operation parameters (such as rotational speed, torque, thrust) at the time of fault occurrence, sensor data (vibration, temperature, pressure, etc.), and fault types (tool wear, crack, fracture). Preprocess the collected fault data using statistical methods to detect and remove outliers to obtain fault state data H;
[0075] The basic lightweight weight transfer network consists of two sub-networks with the same structure but non-shared weights. Each sub-network adopts a deep convolutional neural network structure. The first convolutional layer of the sub-network uses a convolutional kernel of size 3×3, with 16 kernels, a stride of 1, a padding of 1, and the ReLU function as the activation function. After passing through the convolutional-pooling layer combination of each sub-network, the deep features of the input data (fault status data H) are gradually extracted. Finally, the feature map is flattened and mapped to a low-dimensional feature space through a fully connected layer, and the output feature vector has a dimension of d H ;
[0076] To improve the efficiency and accuracy of feature extraction, the traditional convolutional kernel is improved, and an adaptive receptive field convolutional kernel is adopted to dynamically adjust the receptive field size of the convolutional kernel according to the local features of the input data (fault status data H);
[0077] Specifically, an additional branch network is introduced in the convolutional layer. This branch network takes the input feature map of the current convolutional layer as the input, and through convolutional and fully connected operations, outputs a control parameter α. According to the value of α, the weights of the convolutional kernel are adjusted so that the convolutional kernel can capture more appropriate-scale features in different regions;
[0078] The design process of the branch network structure is as follows:
[0079] The additional branch network introduced in the convolutional layer of the basic lightweight weight transfer network has the input data (fault status data H) of the current convolutional layer as its input. The first layer of the branch network is a convolutional layer that uses a convolutional kernel of size 3×3 and a number of C1. After passing through this convolutional layer, the feature map F1 is obtained. Then, a global average pooling layer is used to compress the feature map F1 into a vector v with a dimension of 1×1×C1;
[0080] The purpose of this step is to aggregate the spatial information for more efficient processing later;
[0081] Based on the compression result of the global average pooling layer (the vector v with a dimension of 1×1×C1), a fully connected layer is connected, with the number of neurons being C2 (for example, C2 = 16) and the activation function being the ReLU function to obtain the intermediate vector v1. Then, through another fully connected layer, a scalar value is output, and this scalar value is the control parameter α. Through such a network structure, the branch network can extract effective information from the input feature map to generate the control parameter;
[0082] The implementation process of adjusting the weights of the convolutional kernel according to the value of α is as follows:
[0083] Based on the weight matrix W of the traditional convolution kernel, the weights of the convolution kernel are adjusted according to the control parameter α. Based on the interpolation algorithm, the convolution kernel weight matrix W is divided into multiple sub-regions (divided into small blocks of z×z, where z is the size of the convolution kernel). For each sub-region W i (i represents the index of the sub-region), an interpolation coefficient β i (α) is calculated, and the calculation method is:
[0084]
[0085] where and are two preset different weight coefficients, corresponding to the weights under different receptive field scales respectively;
[0086] Then, the weights of each sub-region are adjusted according to the interpolation coefficient to obtain the adjusted weight matrix When performing the convolution operation, the adjusted weight matrix is used to perform convolution calculation on the input feature map. In this way, in different regions of the input feature map, due to different values of α, the weights of the convolution kernel will be adjusted accordingly, so that the convolution kernel can capture more appropriate-scale features. For example:
[0087] In the region where the feature changes violently (the change amplitude exceeds the set threshold), α makes the receptive field of the convolution kernel larger to capture more extensive context information; while in the region where the feature is relatively smooth (the change amplitude is lower than or equal to the set threshold), α will make the receptive field smaller to focus on local detail features;
[0088] For the fault state data H, it is input into a sub-network of the constructed improved lightweight weight transfer network (the improved lightweight weight transfer network is also composed of two sub-networks with the same structure but non-shared weights). After the forward propagation of the network, the first feature vector f1 of the image is extracted, the first feature vector f1 is normalized, and its numerical range is mapped to the interval [0,1]. The normalized first feature vector f1 is input into another sub-network to extract the second feature vector f2, and the first feature vector f1 and the second feature vector f2 are concatenated to obtain the cross-domain common feature vector [f];
[0089] A feature alignment algorithm based on adversarial training is designed. A discriminator network D is introduced, whose input is the fault state data H and the cross-domain common feature vector [f]. The output is the probability of judging whether the two are from the same distribution. At the same time, an adversarial loss term L is introduced. The weights of the lightweight weight transfer network are updated and improved through the back propagation algorithm of the adversarial loss term L, making it difficult for the discriminator network to distinguish between the fault state data H and the cross-domain common feature vector [f], thereby achieving effective alignment and fusion of cross-domain features and obtaining more universal and transferable common features.
[0090] Specifically, due to the complexity and diversity of shield machine tool fault data, traditional convolution kernels have limitations in capturing features. The adaptive receptive field convolution kernel dynamically adjusts the receptive field size according to the local characteristics of the input data, expands the receptive field in areas with drastic feature changes to capture more contextual information, and shrinks the receptive field in areas with smooth features to focus on local details, which can more accurately extract effective features from different areas.
[0091] Fault prediction module: Based on the obtained optimal state data and common features, fault prediction is performed through the spatiotemporal cross-attention dynamic migration network;
[0092] The spatiotemporal cross-attention dynamic migration network consists of an input layer, an improved dynamic sparsification strategy, a segmented causal convolution trigger mechanism, a fully connected layer, and a classifier;
[0093] The feature fusion module fuses the best state data and common features obtained previously, and performs cross-attention calculations in time and space dimensions by improving the dynamic sparsification strategy and the segmented causal convolution trigger mechanism to capture the best state data [X] and common features. The relationship between them in time and space;
[0094] The implementation process of the improved dynamic sparsification strategy is as follows:
[0095] The traditional attention calculation method involves the calculation of all spatial nodes, which has high computational complexity and contains redundant information. In order to reduce the computational complexity and avoid information leakage, an improved dynamic sparsification strategy is adopted;
[0096] Based on common features The first three key spatial nodes closely related to tool failure are determined (e.g. nodes corresponding to areas with severe tool wear and stress concentration), and the mask matrix is dynamically adjusted to achieve sparsification through a predefined mask matrix M during the spatiotemporal cross-attention process;
[0097] Specifically, the predefined mask matrix M is dynamically adjusted to achieve sparsification according to the current working stage of the tool (excavation, tool change, shutdown) and the complexity of the working conditions (such as soil hardness changes);
[0098] The process of dynamically adjusting the mask matrix is as follows:
[0099] Based on historical data and combined with a machine learning regression model, initially obtain the optimal state data [X] and common features Regarding the complexity of the working conditions between them, use the optimal state data [X] and common features as input features;
[0100] Use the pre - calibrated working condition complexity obtained manually as a label to train the regression model, enabling the model to learn the mapping relationship between the optimal state data, common features, and the working condition complexity, and output a working condition complexity index I;
[0101] Specifically, the working condition complexity obtained through pre - manual calibration or statistical analysis based on actual fault conditions, that is, the complexity of the artificial judgment of the working condition through information such as soil type, hardness, water content, etc., is classified into low, medium, and high levels;
[0102] Meanwhile, set different basic mask matrices M according to the tool working stage template , and then, adjust the basic mask matrix template through a mapping function θ(I) to change the mask values in the area around the key nodes to obtain a sparsified attention weight matrix;
[0103] For example, set a change threshold φ. When the working condition complexity index I ≥ φ (high), expand the range of mask values in the area around the key nodes, so that the attention calculation considers more surrounding information to a certain extent. When the working condition is simple I ≥ φ (low), maintain the characteristics of the original mask matrix to highlight the key node information, thereby realizing the adaptive generation of the mask matrix and more flexibly balancing the computational complexity and information capture ability;
[0104] Specifically, the mapping function θ(I) is a polynomial function constructed based on the working condition complexity index I. During the construction process, study the operating state and fault conditions of the tool under different working condition complexity indices I, and determine the parameters in the mapping function θ(I) through the optimization algorithm gradient descent, expressed as: θ(I) = aI 2 + bI + γ. By continuously adjusting the values of a, b, and γ, make the mask matrix adjusted according to this function perform optimally in the fault prediction task, thereby determining the specific form and parameter values of the mapping function;
[0105] The process of obtaining the sparsified attention weight matrix is as follows:
[0106] The dimension of the mask matrix M is the same as the spatial node dimension in the attention calculation. The position corresponding to the key spatial node is 1, and the rest are 0. When calculating the attention weights, perform an element - by - element multiplication operation on the attention calculation result and the mask matrix M, that is, A sparse= A × M, where A is the original attention weight matrix, and A sparse is the attention weight matrix after sparsification;
[0107] Specifically, the attention weight matrix after sparsification is obtained by improving the dynamic sparsification strategy. In the subsequent feature weighted summation process, only the information of key spatial nodes is retained to participate in the calculation, effectively reducing the computational complexity and avoiding the interference of irrelevant spatial node information on fault prediction;
[0108] The implementation process of the segmented causal convolution trigger mechanism is as follows:
[0109] Design a segmented causal convolution trigger mechanism to monitor the variance σ(t) of the optimal state data [X] in real time and judge the change of the tool running state;
[0110] When a sudden change in the variance of the vibration signal is detected, that is, ((σ(t) - σ(t - Δt) > τ), where Δt is the time interval and τ is the variance mutation threshold), it is considered that the tool has an abnormal situation, and at this time, the deep dilated convolution is enabled;
[0111] Specifically, the deep dilated convolution can expand the receptive field of the convolution kernel without increasing the number of parameters by increasing the interval of the convolution kernel, so as to capture the feature information on a longer time series;
[0112] Let the convolution kernel size of the deep dilated convolution be The dilation rate is ω. For the input sequence, its output formula is:
[0113]
[0114] where, is the convolution kernel weight of the deep dilated convolution, is the index of the input optimal state data [X], and y(t) is the state data sequence output by the segmented causal convolution trigger mechanism;
[0115] When there is no sudden change in the variance of the vibration signal, directly use the (one-dimensional convolution) shallow layer calculation to reduce the dimension of the optimal state data [X] to obtain the state output sequence y(t);
[0116] Based on the sparsified attention weight matrix A sparse and the state output sequence y(t) output by the segmented causal convolution, input them into the fully connected layer and the classifier, calculate the probability of the tool fault type, and output the fault prediction result. The specific process is as follows:
[0117] The sparsified attention weight matrix A sparseThe state output sequence y(t) of the segmented causal convolution output and the sparsified attention weight matrix A of the input data of the fully connected layer are simultaneously input into the fully connected layer. sparse An abstract mapping is performed on the state output sequence y(t) of the segmented causal convolution output and the input data of the fully connected layer (the sparsified attention weight matrix A). Through linear transformation and activation function operations, the input feature vectors are converted into high-level feature representations.
[0118] The high-level feature representations processed by the fully connected layer are input into the classifier. For the tool failure types, there are G categories. The classifier calculates the probability p of each category. G , based on these probability values, it is determined whether different categories are in a failure state. A failure threshold (0.5) is set. Among the categories output by the classifier, the situation where the category probability p G is less than or equal to the failure threshold is in a failure state.
[0119] For example:
[0120] Suppose the shield machine tool failure types are divided into 3 categories: tool wear, tool crack, and tool fracture, that is, G = 3. By obtaining the sparsified attention weight matrix A sparse (after flattening, it is a 100-dimensional vector) and the state output sequence y((t) of the segmented causal convolution output (suppose it is a 50-dimensional vector), after splicing, a 150-dimensional input feature vector is obtained. After calculation by the fully connected layer and applying the ReLU activation function, a 200-dimensional high-level feature representation is obtained. After inputting it into the softmax classifier, the probability of tool wear p{1} = 0.2, the probability of tool crack p{2} = 0.3, and the probability of tool fracture p{3} = 0.5 are calculated. Compared with the failure threshold (0.5), it is found that the tool is in a failure state in terms of wear, crack, and fracture, and a reminder is sent to the staff.
[0121] In the application, several formulas involved are calculated by taking their numerical values after dimensionless. The establishment of the formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. Some coefficients or weights in the formula are set by those skilled in the art according to the actual situation, so no more details will be elaborated here.
[0122] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution.
[0123] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A shield machine tool state fault prediction system for multimodal data fusion, characterized in that, Including: State acquisition module: Collect multi-modal state data, decompose the multi-modal state data into time and space related features, and obtain the optimal state data through compression by low-rank matrix factorization; Among them, the low-rank matrix factorization compression is implemented based on a low-rank matrix approximation algorithm improved by singular value decomposition; Abnormal simulation module: Collect fault state data, construct an improved lightweight weight transfer network for cross-domain feature extraction to obtain cross-domain common feature vectors, and obtain common features through a feature alignment algorithm based on adversarial training based on the cross-domain common feature vectors and fault state data; Among them, the improved lightweight weight transfer network improves the traditional convolution kernel, introduces an additional branch network in the convolutional layer, the additional branch network takes the input feature map of the current convolutional layer as the input, outputs control parameters through convolution and fully connected operations, and adjusts the weights of the convolution kernel according to the values of the control parameters; Fault prediction module: Based on the obtained optimal state data and common features, perform fault prediction through a spatio-temporal cross-attention dynamic transfer network; Among them, the spatio-temporal cross-attention dynamic transfer network includes an improved dynamic sparsification strategy and a segmented causal convolution trigger mechanism.
2. The shield machine cutter state fault prediction system for multimodal data fusion according to claim 1, wherein The process of obtaining the time-related features is as follows: Collect multi-source sensor data, for the key time-related features in the multi-source sensor data, obtain the cutter head speed signal through an encoder, use the fast Fourier transform to convert the speed signal in the time domain to the frequency domain to obtain the spectrum of the speed signal, analyze the peak frequency in the spectrum, and determine the main frequency components of the cutter speed fluctuation to obtain time-related features; The process of obtaining the space-related features is as follows: Adopt the finite element analysis method to perform mechanical modeling on the shield machine cutter and cutter head to obtain the stress distribution data of the cutter under different working conditions, use a high-definition industrial camera to take periodic pictures of the cutter surface to obtain the cutter surface image, adopt an image segmentation algorithm to obtain the area and shape features of the worn area, and obtain space-related features based on the stress distribution data, the area of the worn area and the shape features of the worn area.
3. The shield machine tool state fault prediction system for multimodal data fusion according to claim 2, characterized in that, The implementation process of the low-rank matrix approximation algorithm improved by singular value decomposition is as follows: Integrate the extracted time-related features and space-related features to construct a feature matrix, where the feature matrix includes a left singular matrix, a diagonal matrix, and a right singular matrix; According to the energy ratio of the singular values in the feature matrix, retain the first k larger singular values to construct an approximate matrix; Set an energy ratio threshold, and accumulate a fixed number of singular values in the feature matrix from the largest to the smallest energy ratio until the cumulative sum exceeds the energy ratio threshold, and determine that the number of retained singular values is the optimal number of singular values Based on the optimal number of singular values Select from the left singular matrix columns, select the first diagonal elements from the diagonal matrix, and select the first rows from the right singular matrix to obtain the optimal approximation matrix; The data points corresponding to the best approximate matrix are the best state data.
4. A shield machine tool state fault prediction system for multimodal data fusion according to claim 3, characterized in that, The process of designing the branch network structure in the improved lightweight weight transfer network is as follows: The first layer of the branch network is a convolutional layer. After passing through this convolutional layer, a feature map is obtained, followed by a global average pooling layer, and the global average pooling layer compresses the feature map; Connect a fully connected layer based on the compression result of the global average pooling layer to obtain an intermediate vector, and then pass through a fully connected layer to output control parameters, and adjust the weights of the convolution kernel based on the control parameters.
5. A shield machine tool state fault prediction system for multimodal data fusion according to claim 4, characterized in that, The implementation process of adjusting the weights of the convolution kernel according to the control parameters is as follows: Based on the weight matrix of the traditional convolution kernel, through the interpolation algorithm, the convolution kernel weight matrix is divided into multiple sub-regions, and the interpolation coefficient is calculated for each sub-region according to the value of the control parameter; Adjust the weights of each sub-region according to the interpolation coefficient to obtain the adjusted weight matrix; When performing the convolution operation, use the adjusted weight matrix to perform convolution calculation on the input feature map to achieve the adjustment of the convolution kernel weights.
6. The shield machine tool state fault prediction system for multimodal data fusion according to claim 5, characterized in that, The process of obtaining the common features is as follows: The improved lightweight weight transfer network consists of two sub-networks with the same structure but non-shared weights; Input the fault state data into one sub-network of the constructed improved lightweight weight transfer network. After the forward propagation of the network, extract the first feature vector of the image, perform normalization processing on the first feature vector, and input the normalized first feature vector into another sub-network to extract the second feature vector. Concatenate the first feature vector and the second feature vector to obtain the cross-domain common feature vector; Design a feature alignment algorithm based on adversarial training, introduce a discriminator network, with the input being the fault state data and the cross-domain common feature vector, and the output being the probability of judging whether the two come from the same distribution; Introduce an adversarial loss term, and update the weights of the improved lightweight weight transfer network through the backpropagation algorithm of the adversarial loss term, making it difficult for the discriminator network to distinguish between the fault state data and the cross-domain common feature vector, and obtain the common features.
7. A shield machine cutter tool state fault prediction system for multimodal data fusion according to claim 6, characterized in that, The spatio-temporal cross-attention dynamic transfer network consists of an input layer, an improved dynamic sparsification strategy, a segmented causal convolution trigger mechanism, a fully connected layer, and a classifier.
8. A shield machine tool state fault prediction system for multimodal data fusion according to claim 7, characterized in that, The implementation process of the improved dynamic sparsification strategy is as follows: The output of the improved dynamic sparsification strategy is the sparsified attention weight matrix; Based on the common features, determine the top 3 key spatial nodes closely related to the tool fault, and dynamically adjust the mask matrix through a predefined mask matrix to achieve sparsification; The process of dynamically adjusting the mask matrix is as follows: Combined with historical data and a machine learning regression model, initially obtain the complexity of the working conditions between the optimal state data and the common features, and use the optimal state data and the common features as input features; Use the working condition complexity obtained by pre-artificial calibration as the label, and train the regression model with the label to enable the model to learn the mapping relationship between the input features and the working condition complexity, and output the working condition complexity index; Construct a mapping function based on the working condition complexity index, set different basic mask matrices according to the tool working stage, adjust the basic mask matrix template through the mapping function, and change the mask values in the area around the key nodes to obtain the sparsified attention weight matrix.
9. The shield machine tool state fault prediction system for multimodal data fusion according to claim 8, wherein The implementation process of the segmented causal convolution trigger mechanism is as follows: Real-time monitor the variance of the optimal state data, judge the change of the tool running state, preset the variance mutation threshold, and obtain the change of the tool running state based on the variance mutation threshold; When the variance of the vibration signal mutates, enable the deep dilated convolution to output the state output sequence; Among them, the deep dilated convolution is achieved by increasing the interval of the convolution kernel; When the variance of the vibration signal does not mutate, use one-dimensional convolution shallow calculation to output the state output sequence.
10. A shield machine tool state fault prediction system for multimodal data fusion according to claim 9, characterized in that, The implementation process of the fault prediction is as follows: The sparsified attention weight matrix and the state output sequence output by the segmented causal convolution are simultaneously input into the fully connected layer. The fully connected layer inputs the sparsified attention weight matrix and the state output sequence output by the segmented causal convolution for abstract mapping, and through linear transformation and activation function operations, the input feature vectors are converted into high-level feature representations; The high-level feature representations are input into the classifier. The tool failure types are divided into G categories. The classifier calculates the probabilities of each category, sets a failure threshold. Among the various categories output by the classifier, the cases where the category probability is less than or equal to the failure threshold are in the failure state.
Citation Information
Cited By
Flash furnace fault prediction method based on graph fusion and multi-stage learning
CN120508967A
Intelligent operation monitoring method and device of PEM fuel cell
CN121097136A
Large language model progressive field fine tuning and knowledge fusion method oriented to shield engineering
CN121303257A
A method for progressive domain fine-tuning and knowledge fusion of large language models for tunnel boring machines.
CN121303257B