Structural component fatigue life prediction method based on CNN-BiGRU-MHSA fusion model
By using the CNN-BiGRU-MHSA fusion model, which combines fracture mechanics and multi-source data, a comprehensive damage accumulation evolution equation is constructed. This solves the problem of insufficient accuracy and reliability in fatigue life prediction in existing technologies, and achieves high-precision fatigue life prediction for structural components.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LANZHOU JIAOTONG UNIV
- Filing Date
- 2026-03-25
- Publication Date
- 2026-04-24
AI Technical Summary
Existing methods for predicting the fatigue life of structural components fail to effectively combine fracture mechanics physical mechanisms with multi-source heterogeneous data. Single data-driven models lack physical constraints when facing variable loads and notch effects, resulting in limited generalization ability. Traditional damage evolution models fail to effectively coordinate multiple damage components such as tensile exhaustion, low-cycle fatigue, and crack propagation, leading to insufficient prediction accuracy and reliability.
A CNN-BiGRU-MHSA fusion model is adopted. By acquiring multi-source heterogeneous time-series signals and load spectrum data, and combining geometric boundary parameters to calculate interaction parameters, the nominal stress is transformed into the equivalent notch structural stress. A comprehensive damage accumulation evolution equation is constructed. Features are processed using a one-dimensional convolutional neural network, a bidirectional gated recurrent unit, and a multi-head self-attention mechanism to generate fatigue life prediction values and optimize network parameters.
It improves the model's ability to analyze multi-source features and its predictive generalization, reduces computational bias, enhances prediction accuracy and robustness, quantifies failure mechanisms at different stages, and solves the gradient dissipation problem in long-term series computation.
Smart Images

Figure CN121920242A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of structural health monitoring and reliability assessment technology, specifically a method for predicting the fatigue life of structural components based on a CNN-BiGRU-MHSA fusion model. Background Technology
[0002] In fields such as marine engineering, aerospace, and rail transportation, critical structural components are subjected to the combined effects of vibration, impact, and cyclic loading over extended periods during service, making them highly susceptible to low-cycle fatigue damage. This damage typically manifests as the initiation and propagation of microcracks. If these cracks are not identified and repaired in a timely manner, they will gradually worsen over time, eventually leading to structural fracture and failure, or even catastrophic safety accidents. Therefore, full-life-cycle condition monitoring and accurate fatigue life prediction of structural components are crucial for ensuring the safe and stable operation of industrial equipment.
[0003] Existing fatigue life prediction methods are mainly divided into fracture mechanics methods based on physical mechanisms and data-driven machine learning methods. Regarding physical mechanisms, traditional methods typically rely on finite element numerical simulations or stress-life curves to assess fatigue conditions. These methods require deriving stress intensity factors and crack propagation laws based on empirical or semi-empirical models (such as the Paris formula). However, environmental factors in actual working conditions are complex and variable, and physical models often rely on numerous manually set parameters, resulting in high computational costs and difficulty in fully considering the effects of multi-physics coupling. In particular, traditional analytical models struggle to accurately characterize the nonlinear characteristics of crack initiation and stress field distribution under complex boundary conditions, leading to significant theoretical deviations in prediction results.
[0004] With the development of artificial intelligence technology, data-driven methods based on deep learning are increasingly being applied to fatigue life prediction. Researchers utilize Long Short-Term Memory (LSTM) networks or Gated Recurrent Units (GRUs) to process time-series data, or employ Convolutional Neural Networks (CNNs) to extract spatial features of signals. While these individual deep learning models alleviate the limitations of traditional methods to some extent, they still have significant shortcomings in practical applications. For example, recurrent neural networks are prone to gradient vanishing or overfitting due to excessive parameters when processing long-sequence data, and they struggle to capture local spatial features; while convolutional neural networks, although adept at extracting local features, lack the ability to perceive long-distance temporal dependencies, making it difficult to describe the evolution of fatigue damage over time.
[0005] Furthermore, most existing hybrid neural network models fall into the category of purely data-driven approaches, lacking interpretability at the physical level. When faced with non-stationary scenarios such as scarce high-quality full-life-cycle data, shifting operating conditions, or strong noise interference, purely data-driven models often exhibit poor generalization ability and fail to accurately reflect the true physical state of damage within the structure. How to effectively integrate the spatiotemporal characteristics of multi-source heterogeneous monitoring data and embed the physical mechanisms of fracture mechanics into a deep learning framework to balance prediction accuracy, robustness, and physical interpretability is a pressing technical challenge in the field of structural fatigue life prediction. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a method for predicting the fatigue life of structural components based on a CNN-BiGRU-MHSA fusion model. The technical problem it solves is that existing methods for predicting the fatigue life of structural components fail to effectively combine fracture mechanics physical mechanisms with multi-source heterogeneous data. Single data-driven models lack physical constraints when facing variable loads and notch effects, resulting in limited generalization ability. At the same time, traditional damage evolution models fail to effectively coordinate multiple damage components such as tensile exhaustion, low-cycle fatigue, and crack propagation, resulting in insufficient accuracy and reliability of remaining life prediction under relevant working conditions.
[0007] To address the above problems, the present invention provides the following technical solution: This invention provides a method for predicting the fatigue life of structural components based on a CNN-BiGRU-MHSA fusion model, comprising the following steps: Acquire multi-source heterogeneous time-series signals and load spectrum data; calculate interaction parameters based on geometric boundary parameters and relative crack length, transform nominal stress into equivalent notched structural stress, and synthesize effective stress intensity factor, using equivalent notched structural stress and effective stress intensity factor as crack propagation mechanism features; couple tensile damage component, low-cycle fatigue damage component, and crack propagation damage component based on effective stress intensity factor to construct a comprehensive damage accumulation evolution equation, and solve for the normalized remaining life as the fatigue life label; splice multi-source heterogeneous time-series signals and crack propagation mechanism features to construct a standard input tensor; use a one-dimensional convolutional neural network to process the standard input tensor to extract local spatial feature vectors; use a bidirectional gated recurrent unit to process the local spatial feature vectors to extract hidden state features; use a multi-head self-attention mechanism to process the hidden state features to generate weighted fusion features; use a fully connected regression layer to process the weighted fusion features to output fatigue life prediction values, and optimize network parameters based on fatigue life labels.
[0008] Furthermore, in the process of acquiring multi-source heterogeneous time-series signals and load spectrum data, strain sequences are acquired through strain sensors, vibration acceleration sequences are acquired through accelerometers, and load spectrum data is simultaneously acquired through force sensors. Using the variation characteristics of the load spectrum data as a benchmark, time axis alignment is performed on the strain and vibration acceleration sequences, and interpolation resampling is performed on channel data at different sampling rates to ensure that the number of data points in all channels remains consistent within a unit time window.
[0009] Furthermore, the specific method for calculating the interaction parameters based on geometric boundary parameters and relative crack length is as follows: Calculate the dimensionless geometric correction function in the tensile and bending modes respectively, based on the geometric boundary parameters and instantaneous relative crack length; combine the stress distribution in the load spectrum data and the dimensionless geometric correction function to calculate the interaction parameters of the stress field of the notched structure; using these interaction parameters and the ligament width determined based on the geometric boundary parameters and instantaneous relative crack length, transform the nominal stress under the non-uniform stress field into the equivalent notched structure stress.
[0010] Furthermore, the specific process of constructing the comprehensive damage accumulation evolution equation is as follows: using the Paris crack propagation model and the effective stress intensity factor, the crack propagation life from the initial length to the fracture failure length is calculated; the tensile damage component characterizing the exhaustion of large strain ductility, the low-cycle fatigue damage component characterizing the consumption of cyclic plastic strain, and the crack propagation damage component based on the crack propagation life are added together and integrated to form the comprehensive damage accumulation evolution equation; based on the preset damage critical threshold, the equation is solved using an iterative method to obtain the theoretical total fatigue life, and the normalized remaining life is calculated based on the theoretical total fatigue life and the accumulated number of load cycles.
[0011] Furthermore, the steps for constructing the standard input tensor include: concatenating multi-source heterogeneous time-series signals and crack propagation mechanism features along the feature dimension to obtain a multi-dimensional feature matrix; independently performing max-min normalization on each feature dimension of the multi-dimensional feature matrix to map it to the [0,1] interval; setting the time window length and sliding step size, performing time-series slicing on the normalized data to obtain a sequence sample matrix, and using a many-to-one mapping strategy to determine the regression target of each matrix; stacking the sequence sample matrices to generate the standard input tensor and then dividing it into training and test sets.
[0012] Furthermore, the feature processing mechanisms of each network layer in the model are as follows: The one-dimensional convolutional neural network performs multi-channel parallel convolution operations on the standard input tensor using multiple one-dimensional convolutional kernels. A Same padding strategy is employed to maintain consistent dimensions. After nonlinear transformation using the ReLU activation function, a max-pooling layer is used to downsample and output a local spatial feature vector. A bidirectional gated recurrent unit calculates the reset and update gates based on the local spatial feature vectors, performing forward computation in parallel to capture historical evolution information of the sequence, and backward computation to capture future dependency information of the sequence. The forward and backward outputs are integrated to generate hidden state features. A multi-head self-attention mechanism uses a linear transformation matrix to map the hidden state features into query vectors, key vectors, and value vectors. The scaled dot product of the query vector and key vector is calculated and normalized using the Softmax function to obtain global feature attention weights. These weights are then used to weight the value vectors before concatenation and linear transformation to generate weighted fusion features. A fully connected regression layer performs dimensionality reduction mapping on the weighted fusion features to compress and integrate global semantic information. A linear transformation is performed using a linear weight matrix and bias terms to obtain a scalar calculation result as the fatigue life prediction value.
[0013] Furthermore, the steps to optimize the network parameters are as follows: calculate the mean squared error loss between the predicted fatigue life value and the fatigue life label in the dominant objective function; add an L2 weight decay regularization term to the dominant objective function and introduce a Dropout random deactivation mechanism before the fully connected regression layer; adjust the learning rate using a piecewise constant descent method; and use an adaptive moment estimation optimization algorithm to perform parameter update operations until the objective function converges.
[0014] This invention provides a method for predicting the fatigue life of structural components based on a CNN-BiGRU-MHSA fusion model. It has the following beneficial effects: 1. This invention transforms nominal stress into equivalent notch structural stress and synthesizes an effective stress intensity factor, which is then concatenated with multi-source heterogeneous time-series signals to construct a standard input tensor. This feature processing method introduces fracture mechanics mechanisms into the data-driven model, providing physical boundary conditions for the neural network. This reduces the computational bias caused by a single data-driven model under varying load conditions and notch geometry, and improves the model's analytical capability and predictive generalization for multi-source features.
[0015] 2. This invention couples tensile damage components, low-cycle fatigue damage components, and crack propagation damage components to construct a comprehensive damage accumulation evolution equation, and solves for the normalized remaining life as a fatigue life label. This calculation process coordinates the failure mechanisms of structural components at different stages. Compared with a single linear accumulation theory, it can quantify the synergistic effect of large strain ductility depletion, cyclic plastic strain consumption, and crack propagation, providing a physically consistent supervision label for subsequent model training.
[0016] 3. This invention employs a cascaded network consisting of a one-dimensional convolutional neural network, bidirectional gated recurrent units, and a multi-head self-attention mechanism to process standard input tensors. The convolutional layers perform multi-channel local dimensionality reduction and noise reduction, the recurrent units establish temporal forward and reverse correlations to handle long-term dependencies in fatigue accumulation, and the self-attention mechanism calculates global feature attention weights, assigning corresponding weights to data channels and time nodes that contribute significantly to damage evolution. This network architecture reduces feature redundancy in processing multi-source heterogeneous data and solves the gradient dissipation problem in long-term sequence computation. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the fatigue life prediction method for structural components based on the CNN-BiGRU-MHSA fusion model of the present invention. Figure 2 This is a schematic diagram of the structure of the hybrid neural network (CNN-BiGRU-MHSA) of the present invention; Figure 3 This is a linear regression analysis graph showing the predicted fatigue life values and actual fatigue life values in the test set of this invention. Figure 4 A bar chart comparing the performance of different network models of the present invention in terms of coefficient of determination and residual prediction error. Figure 5 This is a bar chart comparing the performance of different network models of the present invention on error-related evaluation metrics. Detailed Implementation
[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] See attached document Figure 1 This invention provides a fatigue life prediction method based on CNN-BiGRU-MHSA learning for different structural components, comprising the following steps: S100 collects multi-source heterogeneous time-series signals and load spectrum data of the structural components under test in service or test conditions, and calculates crack propagation mechanism characteristics and fatigue life labels based on fracture mechanics theory. S200 performs max-min normalization and sample serialization on multi-source heterogeneous time-series signals to eliminate differences in data dimensions and generate a standard input tensor, and divides the training set and test set. S300 constructs a convolutional neural network module, which uses a one-dimensional convolutional kernel to slide and compute on the time axis to extract local spatial feature vectors from the input tensor. S400 constructs a bidirectional gated loop unit module, which receives local spatial feature vectors and extracts hidden state features containing bidirectional temporal dependencies through forward and reverse gating mechanisms. S500 constructs a multi-head self-attention mechanism module, maps hidden state features to multiple subspaces, calculates global feature attention weights, and generates weighted fusion features. The S600 constructs a fully connected regression layer, receives weighted fusion features for high-order semantic integration, outputs the fatigue life prediction value of structural components, and optimizes the network parameters based on backpropagation of the loss function.
[0020] A fatigue life prediction method based on CNN-BiGRU-MHSA learning establishes an end-to-end deep learning prediction architecture, aiming to solve the problems of difficulty in obtaining parameters for traditional physical models under complex working conditions and the limited feature extraction capabilities of single machine learning models. The overall architecture achieves a direct mapping from raw monitoring data to fatigue life through concatenated convolutional neural networks, bidirectional gated recurrent units, and multi-head self-attention mechanisms.
[0021] In step S100, data acquisition encompasses acceleration signals, strain signals, and load spectra for different structural components such as aluminum alloy sheets, steel wires, and truss members. To ensure the physical interpretability of the data labels, a fracture mechanics mechanism model is introduced. Using physical quantities such as stress changes in the notched structure, interaction parameters, and stress intensity factors, the theoretical fatigue life and damage evolution process are calculated, serving as benchmark labels or auxiliary features for supervised learning. This ensures that model training is not only based on statistical data patterns but also conforms to the physical laws of material fatigue fracture.
[0022] In step S200, the preprocessing procedure standardizes the heterogeneous data collected from different sensors. Max-min normalization maps the data to a unified interval, eliminating the impact of amplitude differences on network weight updates. Simultaneously, sliding window slicing is applied to continuous time-series data to construct sequence samples containing historical information, and the training set data is randomly shuffled to improve the model's generalization ability.
[0023] In step S300, the convolutional neural network, as the front end of feature extraction, is mainly responsible for capturing local fluctuation patterns in the time series signal. For short-term abrupt changes and periodic features in vibration or stress signals, one-dimensional convolutional layers are used for filtering and feature abstraction. Nonlinear factors are introduced through the ReLU activation function, and max-pooling layers are used to downsample the feature map, reducing the data dimensionality while preserving key features, thus providing a compact spatial feature representation for subsequent time series modeling.
[0024] In step S400, the bidirectional gated recurrent unit module is responsible for handling the temporal dependencies of the convolutional layer output features. Fatigue damage is a cumulative process; the current state depends not only on the current input but also on the historical load history. The BiGRU structure captures the historical evolution information and future dependency information of the sequence (during offline training) through two independent information streams, forward and backward. The gating mechanism dynamically adjusts the forgetting and retention of information, effectively solving the gradient vanishing problem in long sequence training and outputting a hidden state sequence containing complete contextual information.
[0025] In step S500, the multi-head self-attention mechanism module further processes the hidden state of the BiGRU output. Given that load cycles at different stages of fatigue failure contribute differently to damage, the MHSA mechanism allows the model to automatically focus on key time steps or feature dimensions globally. By projecting features onto multiple subspaces and computing attention weights in parallel, the model can identify and enhance feature components that contribute significantly to lifetime prediction, suppress noise interference, and generate enhanced feature representations with a global perspective.
[0026] In step S600, the fully connected layer flattens and linearly combines the multidimensional weighted features output by MHSA, and obtains continuous lifetime prediction values through regression output layer. During the training phase, the error between the predicted value and the physical label generated in step S100 is calculated, and the weight parameters of each layer of CNN, BiGRU and MHSA are jointly optimized using the backpropagation algorithm until the model converges.
[0027] In step S100, the aim is to construct a digital representation that reflects the physical state of a structural component throughout its entire process from integrity to failure. The fatigue damage evolution of a structural component is a multi-scale physical process. The initiation and propagation of local microcracks lead to distortion of the local strain field of the material, while the stiffness degradation caused by crack accumulation is reflected in the overall vibration response of the structure. Therefore, this embodiment employs joint monitoring of local strain signals, overall acceleration signals, and external driving load signals to capture damage characteristics across different dimensions. Specifically, it includes the following sub-steps: S110, determine the monitoring object and key monitoring area of the structural component to be tested. In this embodiment, the structural component to be tested includes, but is not limited to, aerospace aluminum alloy thin plates, high-strength steel wires, and space truss members. Specific monitoring locations are set for different structural forms. Specifically, for aerospace aluminum alloy thin plates, the monitoring area is set within 0 to 50 mm along the expected crack propagation direction at the root of the prefabricated notch to capture the stress field at the crack tip; for high-strength steel wires, the monitoring area is set at the transition position between the anchor clamping area and the free section, which is a high-incidence area of stress concentration and fretting wear; for truss members, the monitoring area is selected at the stress concentration node welds and the mid-span position of the member.
[0028] S120, Deploy a multi-type sensor network to acquire heterogeneous physical quantities. This step obtains multi-dimensional physical information through different types of sensors. First, resistive strain gauges are attached to the key monitoring areas along the principal stress direction to acquire local micro-deformation signals of the structure, denoted as strain sequences. Secondly, piezoelectric accelerometers are rigidly fixed at locations with large modal displacements of the structure to acquire vibration signals reflecting the structure's dynamic characteristics, denoted as acceleration sequences. Next, by connecting a force sensor in series to the end of the loading actuator of the fatigue testing machine, or by using the force feedback simulation output interface of the testing machine, the cyclic load spectrum or random load history applied to the structure is obtained, and recorded as the load sequence. To ensure that the signal can capture the high-frequency transient characteristics caused by fatigue damage, the sampling frequency is set to more than 10 times the first-order natural frequency of the structure, typically ranging from 1kHz to 50kHz, and the analog-to-digital conversion resolution is set to no less than 16 bits to ensure the analytical accuracy of minute strain changes.
[0029] S130 performs multi-channel synchronous acquisition and timing alignment. A multi-channel dynamic signal acquisition instrument is used to sample the aforementioned sensor channels in parallel. Given the differences in response characteristics among sensors of different physical quantities, a unified external hardware clock trigger mode is used during the acquisition process to ensure zero phase difference in the time axis of the data from each channel. For minor time deviations caused by differences in transmission line length or sensor response delays, the strain and acceleration signals are time-axis shifted and calibrated using the peak moment of the load signal as a reference. Furthermore, if the original sampling rates of each channel are inconsistent, the data from the lower sampling rate channels is resampled using polynomial interpolation based on the highest sampling rate, ensuring that the number of data points in all channels remains consistent within a unit time window.
[0030] S140 defines the original input data matrix and data structure. The original input matrix is constructed from the acquired and aligned multi-source heterogeneous signals. Let the time length of a single sampling include At each time step, the system has a total of [number] connections. One strain sensor channel, One accelerometer sensor channel and For each load sensor channel, the total number of feature dimensions of the input data is... Defined as The original input data is defined as a two-dimensional real matrix. Its mathematical expression is At any given time state vector It is formed by cascading the instantaneous amplitudes of each sensor channel, i.e. In the formula, Indicates the first Each strain sensor at time The strain amplitude; Indicates the first An accelerometer at time The acceleration amplitude; Indicates the first Each load channel at time The load amplitude. This two-dimensional real matrix. This serves as the underlying input for subsequent deep neural network models, containing the structure in Complete spatiotemporal physical information within each time step.
[0031] This step utilizes the equivalent transformation principle in fracture mechanics to unify the stress state under different geometric configurations and loading conditions into an energy release rate index that drives crack propagation. This transformation eliminates the interference of structural size effects on the stress field and extracts physical characteristics directly related to the nature of material fatigue damage. The specific calculation process includes the following sub-steps: S210 calculates a dimensionless geometric correction function related to crack length. The stress concentration at the crack tip dynamically changes with crack length propagation and is constrained by the geometric boundaries of the structure. The relative crack length is defined. Instantaneous crack length With ligament width The ratio, i.e. Its value range is (0,1). For the tensile stress mode of the notched structure, the first dimensionless function is calculated. : ; For the bending stress mode of the notched structure, calculate the second dimensionless function. : ; In the above formula, Pi is a constant, and the trigonometric function parameters are in radians. These two functions serve as geometric correction coefficients for the tensile stress component and the bending stress component, respectively, and are used in the subsequent synthesis of the stress intensity factor.
[0032] S220, calculate the interaction parameters of the stress field in the notched structure. To decouple the nonlinear effects of notch geometry, load distribution, and crack propagation path, the interaction parameters are calculated. The calculation formula is as follows: ; In the formula, The integral result of the stress distribution along the notch ligament is taken as the difference between the maximum and minimum stresses in the load spectrum; and These are the components of local membrane stress and bending stress, respectively, representing the average stress value and the asymmetry of stress distribution on the crack surface; and The relative crack length calculated in step S210 The changing function value; This represents the differential component relative to the crack length. (Integration lower limit) The upper limit of integration corresponds to the ratio of the minimum initial crack length to the crack width that the non-destructive testing equipment can identify. This corresponds to the ratio of critical crack length to width determined by the structural fracture toughness criterion. Dimensionless data related to the fatigue properties of materials are defined as fatigue characteristic parameters.
[0033] S230, calculate the stress change in the notched structure. Based on the calculated interaction parameters, the nominal stress under the non-uniform stress field is transformed into the equivalent stress of the notched structure. The calculation formula is as follows: ; In the formula, This represents the change in nominal stress. The ligament width of the structural component under test on the crack propagation surface; Dimensionless data related to material fatigue properties are defined as fatigue characteristic parameters, reflecting the sensitivity of crack propagation rate to stress amplitude. This parameter is related to the microstructure of the material's grain structure and is usually used as a known constant input. The symbol is for multiplication, indicating multiplication. This formula introduces the ligament width... and interaction parameters The power-law correction achieves the normalization of fatigue data for structures of different thicknesses.
[0034] S240 calculates the stress intensity factor and the effective stress intensity factor. The stress intensity factor is the only parameter characterizing the strength of the elastic stress field at the crack tip. It combines the contributions of both tensile and bending stress modes. The calculation formula is: ; Based on this, the effective stress intensity factor is calculated. : ; In the formula, The maximum stress intensity factor is calculated by substituting the maximum peak stress from the load spectrum acquired in step S100 into the above... The calculation process (i.e., at this time) Obtained by taking the maximum stress value; The stress intensity factor corresponding to the initial crack is determined by the initial crack length. The effective stress intensity factor is obtained by substituting the corresponding stress level into the formula. By eliminating the influence of the initial damage threshold, it directly reflects the mechanical driving force that drives the crack to further expand from the initial state, and is a key physical feature input for subsequent neural network lifetime prediction.
[0035] A mapping relationship is established between physical mechanisms and data-driven approaches. By coupling fracture mechanics and damage mechanics models, the collected discrete strain and load data are transformed into continuous structural health status indicators. This process does not rely on expensive full-life failure tests, but rather derives the theoretical life based on the material's inherent physical properties, thus providing low-cost, high-precision supervision labels for deep learning models. The specific calculation process includes the following sub-steps: S250, the theoretical crack propagation life is calculated based on the Paris crack propagation model in the material damage model. First, in order to describe the dynamic behavior of crack propagation, this embodiment uses the Paris model to define the crack propagation rate. With effective stress intensity factor The power-law relationship between them can be expressed mathematically as follows: ; In the formula, Dimensionless data related to the fatigue properties of materials are defined as fatigue characteristic parameters; The Paris constant of the material is used, and the effective stress intensity factor obtained in step S240 is used. Based on the crack propagation characteristics of the material, the crack length from the initial crack is calculated. Crack length extending to fracture failure The number of load cycles experienced. This process is achieved by integrating the crack propagation rate formula, whose mathematical expression is: ; In the formula, The crack fatigue life value is solely contributed by the crack propagation stage; The length of the crack; This represents the infinitesimal component of the crack length. The initial crack length is determined by the lower limit of the non-destructive testing equipment or the inherent defect size of the material. The crack length at fracture failure is determined by the fracture toughness criterion. Dimensionless data related to material fatigue properties are defined as fatigue characteristic parameters. This integration process employs a numerical integration method, calculating and accumulating the crack increment for each load cycle until the crack length reaches [a certain value]. .
[0036] S260, Constructing the sub-terminal damage mechanism model in the material damage model. To quantify the degree of damage evolution of the structure under complex loads, this embodiment introduces a multi-mechanism coupled model consisting of tensile damage, low-cycle fatigue damage, and crack propagation damage. First, the tensile damage component is defined. This component characterizes the ductile exhaustion under large strain: ; Secondly, define the low-cycle fatigue damage component. This component describes the life loss due to cyclic plastic strain based on the Coffin-Manson relation: ; Finally, define the crack propagation damage component. This component characterizes the cumulative damage caused by crack propagation in the plastic zone: ; In the above formula, This represents the maximum plastic strain under cyclic loading. The monotonic fracture strain of the material is determined by a standard uniaxial tensile test. Dimensionless data related to the fatigue properties of materials are defined as fatigue characteristic parameters; To stabilize the plastic strain range of the hysteresis loop width, the acquired strain sequence was analyzed. Obtained by rainflow counting; This represents the crack fatigue life value, i.e., the current number of load cycles. Total fatigue life; The low-cycle fatigue hardening index of the material is determined by low-cycle fatigue testing. is the Paris constant of the material.
[0037] S270, Establish the comprehensive damage accumulation evolution equation and solve for the total lifespan. Based on the linear cumulative damage theory, when the structure experiences fracture failure, its comprehensive cumulative damage degree reaches a critical value of 1. Couple the three damage components in step S260 to establish an implicit equation regarding the total lifespan: ; In the formula, we will uniformly denote the variable representing lifespan as... (i.e., in the formula) equals in the critical state ), as the unknown quantity to be solved; The low-cycle fatigue hardening index of the material; To stabilize the plastic strain range of the hysteresis loop width, this equation is a nonlinear equation, and in this embodiment, the Newton-Raphson iterative method is used for numerical solution. (Setting...) The initial estimate is approximated iteratively until the residual between the result calculated on the left side of the equation and 1 is less than a preset convergence threshold (e.g., 10). -4 Thus, the theoretical total fatigue life under this working condition can be obtained. .
[0038] S280 generates a normalized remaining lifetime (RUL) tag sequence. The total fatigue life is then obtained. Then, for each time step in the training dataset Define its corresponding remaining lifetime label The calculation formula is: ; In the formula, Deadline Cumulative load cycles completed. Remaining lifetime label. It is a dimensionless value ranging from [0,1]: at the initial moment At the time of failure This sequence will serve as the ground truth for the subsequent regression layer of the deep neural network (S600), used to calculate the loss function and guide the update of network parameters.
[0039] See attached document Figure 2 This step is the core of feature engineering for multi-source heterogeneous data. Because the original strain, vibration, and load data, as well as the calculated fracture mechanics parameters, have completely different physical dimensions and numerical orders of magnitude, directly inputting them into a deep neural network can lead to gradient oscillations or convergence difficulties during model weight updates. Furthermore, fatigue damage is a historically dependent cumulative process, requiring the reconstruction of long-term continuous signals into short sequence fragments containing local evolution patterns to fit the input structure of time-series networks. The specific processing includes the following sub-steps: S310, perform channel-level independent normalization processing. This is applied to the original input matrix constructed in step S140. The physical feature parameter sequences calculated in steps S220 to S240 are cascaded along the feature dimension to construct a multi-dimensional feature matrix containing features of multi-source heterogeneous signals and crack propagation mechanisms. For each feature channel of this matrix... (in The value ranges from 1 to the total number of feature dimensions. Search for its maximum value over the entire time domain. and minimum value The maximum-minimum normalization algorithm is used to linearly map the data of each channel to the closed interval [0,1]. The calculation formula is as follows: ; In the formula, It is an index variable for the time dimension, i.e., time moment, and its value range covers the entire sampling process; For the first Each feature channel at the sampling time The value; These are the standardized values; To prevent division by zero errors, the smallest constant is typically taken to be 1 × 10⁻⁶. -8 Specifically, for the notched structural stress calculated in S230... and the effective stress intensity factor calculated in S240 Although it has a clear physical meaning, in order to ensure the consistency of the neural network input distribution, the same method is used to perform dimensionless processing.
[0040] S320, Constructing Overlapping Sliding Time Window Sequences. In order to capture the local temporal dependencies in the fatigue damage evolution process and to segment long-period monitoring data into fixed-length sequence samples, a sliding window technique is used to slice the standardized data matrix.
[0041] First, determine the length of the time window. This length needs to cover the number of sampling points involved in at least 3 to 5 complete load cycles of the structure to include complete hysteresis loop information. Its calculation is based on... ,in The number of iterations (take an integer, such as 5). Sampling frequency, This represents the average loading frequency. In practical implementation, Usually set to The form (such as 512, 1024, or 2048) is used to optimize computer memory alignment and computational efficiency. Next, the sliding step size is set. This parameter determines the overlap rate between samples. It is typically set to... (For example, take) This allows for 50% information overlap between adjacent samples, thus achieving data augmentation and smoothing boundary effects. For the first... A sample segment is extracted from the normalized matrix at time [time]. The beginning of the continuous Data from each time step is used to construct a sequence sample matrix. : ; In the formula, the starting time , For a dimension A two-dimensional real matrix, sample fragment The value range is determined by the total duration and the window length. and step length A joint decision.
[0042] S330, the alignment mapping between time-series samples and lifetime labels. In a supervised learning framework, each sequence sample matrix... A specific lifespan label must be assigned. This is because fatigue damage is a cumulative process, and the sequence sample matrix... Containing state information over a period of time, this embodiment employs a many-to-one mapping strategy, that is, using historical data to predict the state at the current moment. Based on the remaining lifetime tag sequence generated in step S280... Determine the first The regression target for each sample segment : ; The mapping rule uses the remaining lifetime label corresponding to the last time step within the input window as the regression target for that sample. This means that neural networks need to be based on the past. Based on historical information at each time step, the current remaining lifespan can be inferred.
[0043] S340, Random partitioning and tensor encapsulation of the dataset. To evaluate the model's generalization ability, the constructed set of sample pairs is... The dataset is divided into training, validation, and test sets according to a preset ratio. Stratified random sampling is used to ensure that samples in each dataset cover all stages of the lifecycle (i.e., from...). From the intact stage to (The failure phase). The specific partitioning ratio is set as follows: the training set accounts for 70% of the total sample size and is used for gradient descent updates of model parameters; the test set accounts for 30% of the total sample size and is used for final performance evaluation. After partitioning, each dataset is converted into a high-dimensional tensor format supported by the computing platform and encapsulated into batch data containing the batch size dimension to facilitate batch matrix operations on parallel computing units.
[0044] Because the number of full-life-cycle fatigue failure samples obtained in actual engineering projects is scarce and the cost is high, deep neural networks are prone to overfitting when trained on small sample datasets, meaning they perform well on the training set but have poor generalization ability under unknown working conditions. Furthermore, sensor signals in industrial environments inevitably contain environmental noise such as electromagnetic interference. Therefore, this embodiment employs a strategy combining data augmentation and multiple regularization to improve the robustness of the model. The specific implementation process includes the following sub-steps: S350 performs data augmentation based on random noise injection. To simulate measurement noise from the sensor under different operating conditions and to broaden the diversity of the training data, the input sequence sample matrix is augmented during each iteration of the training phase. Perform dynamic noise injection. Generate the enhanced sample matrix. The calculation formula is: ; In the formula, Represents the sequence sample matrix Dimension (i.e.) For the same random noise matrix, the matrix elements follow a standard normal distribution with a mean of 0 and a variance of 1; This is the noise intensity coefficient, used to control the amplitude of the injected noise. In this embodiment, The value of is set to 1% to 5% of the standard deviation of the original signal. This operation allows the samples processed by the neural network in different training epochs to have small numerical differences, thereby prompting the model to learn the underlying trends of the signal rather than memorizing specific noise patterns.
[0045] S360, configure Dropout random deactivation mechanism. To disrupt the complex co-adaptation relationship between neurons and prevent the model from over-relying on certain specific feature channels, a Dropout layer is introduced at the output of subsequent neural network layers. This embodiment uses an inverted Dropout strategy. Assume the output vector of a neuron in a certain layer is... The output vector after Dropout processing The calculation is as follows: ; In the formula, For a with For the same mask vector, its elements follow a Bernoulli distribution, i.e., with probability... The value is 1, with probability The value is 0; The preset inactivation probability is usually set between 0.2 and 0.5; Hadamard product (element-by-element multiplication); denominator This is used to scale activation values during the training phase to maintain consistency with expected output, so that the model can be used directly during the testing phase without adjusting the weights.
[0046] S370 introduces L2 weight decay regularization. To limit the numerical amplitude of model parameters and prevent the model from forcibly fitting outliers or noise points in the training data by increasing weight values, an L2 regularization term is added to the model's objective loss function. The corrected total loss function is shown below. Defined as: ; In the formula, The model predicted values and the regression target values defined in S330 The fundamental error loss between them (e.g., mean square error, MSE); It is the set of all trainable weight parameters in the network; For set A specific weight vector in the process; The summation symbol; The square of the L2 norm of the weight vector; The regularization factor (WeightDecay) is usually set to 10. -5 Up to 10 -3 This term applies a penalty gradient to the weights during backpropagation, proportional to their values, causing the weights to tend towards smaller values, thus smoothing the decision boundary of the model.
[0047] S380, Deploy the early stopping monitoring strategy. To automatically determine the optimal number of iterations during training and avoid overtraining, this embodiment sets up an early stopping mechanism based on validation set performance. After each training iteration, the current performance of the model is evaluated using the validation set partitioned in S340, and the validation set loss is calculated. Set the historical lowest validation loss variable. and tolerance counter If the current round Then update And save the current model's optimal parameters, while resetting... If it is 0; ,but Add 1. When Exceeding the preset patience threshold (For example, when set to 10 to 20 rounds), the training process is forcibly terminated, and the corresponding load is applied. The model parameters are used as the final training model.
[0048] This embodiment uses a one-dimensional convolutional neural network (1D-CNN) to process the enhanced sample matrix generated in step S350. Hierarchical spatial feature extraction is performed. This process is not simply data dimensionality reduction, but rather establishes a mapping relationship between the original vibration signal and the microcrack propagation mode inside the material through sliding calculations of the convolutional kernel in the time dimension. Convolutional neural networks, utilizing their weight-sharing and local connectivity characteristics, can extract translation-invariant local waveform features from noisy, non-stationary signals, such as abrupt stress concentration points or response modes at specific frequencies. The specific processing includes the following sub-steps: S410 performs multi-channel parallel convolution operations. The enhanced sample matrix... The input is fed into a convolutional layer, where a set of learnable convolutional kernels is used to extract features along the time series direction. In this embodiment, the number of convolutional kernels is set to 256, and the kernel size is... The value is set to 7. To maintain the same length of the output feature map as the input in the time dimension, a Same padding strategy is adopted, which involves padding both sides of the time axis boundary of the input data with zero values so that the output size after the convolution operation remains unchanged. The calculation expression for the output information of the convolutional layer is as follows: ; In the formula, To output the convolutional feature map at the index position The value; This is the index of the position where the convolutional kernel slides across the input temporal data; The size of the convolution kernel; For the first The weights of the convolutional kernel at each position; This refers to the enhanced sample matrix generated in step S350. , That is, the matrix is at the th... The element value at each position; The sliding stride of the convolution kernel is set in this embodiment. This is the bias term for the convolution kernel. The output of this step is... This forms an initial feature map that reflects the characteristics of different frequency bands.
[0049] S420, Nonlinear Activation Mapping. To enable the network to process nonlinear fatigue damage evolution relationships, especially for nonlinear elastoplastic behavior in hysteresis loops, the output convolutional feature map of step S410 is modified at the index position. value A modified linear unit (ReLU) is introduced for nonlinear transformation. This process forces the less than 0 values in the convolution output to 0, while leaving the greater than 0 values unchanged. Its mathematical expression is: This operation allows the network to select activation features that contribute significantly to lifetime prediction, filter out weak negative responses caused by background noise, and introduce sparsity to reduce the risk of model overfitting.
[0050] S430, Temporal Feature Dimensionality Reduction and Quality Filtering. To reduce computational cost and enhance the robustness of features to small temporal shifts, the activated feature map... Perform max pooling. This embodiment sets the size of the pooling window. The step size is 3. The step size is 2. That is, the step size on the feature map is... A sliding window is used to select the maximum value within the window as the representative feature of that region. The calculation expression is as follows: ; In the formula, For the pooling calculation, the first The set of maximum values at each position; This refers to the size of the pooling window; This is the position index inside the pooling layer window; This represents the numerical value of the activated feature map at the corresponding position. This step compresses the temporal dimension of the feature sequence, preserving the strongest activation response within the local region, thus forming a higher-order spatial feature sequence.
[0051] S440, Fully Connected Feature Integration and Normalization. To map the extracted local spatial features to a unified semantic space and normalize the feature strengths for use as input to subsequent temporal networks, the set of maximum values from the pooling layer outputs is used. The input is fed into a fully connected layer and processed using the softmax function. The calculation formula is: ; In the formula, It is the first step calculated by the softmax function. The probability vector of each output neuron; This refers to the set of maximum values output by step S430. Element; From the first The input neuron to the first The weight parameters of each output neuron; For the first The bias of each neuron. It should be noted that the fully connected layer in this embodiment is not used for final classification, but rather as a feature transformation layer, mapping the features extracted by convolution into a normalized feature vector with probability distribution characteristics. This vector will then be input into the BiGRU network to capture long-term temporal dependencies.
[0052] In fatigue life prediction, structural damage is a non-Markovian process with a memory effect. The current damage state depends not only on the current load input but also on historical accumulated damage and future load change trends. To effectively capture this long-short time-series dependency and address the gradient vanishing problem that traditional recurrent neural networks (RNNs) easily encounter when processing long sequences, this embodiment uses a BiGRU network to perform time-series modeling on the feature sequence output in step S440. This network structure consists of a forward GRU layer and a backward GRU layer, achieving complete extraction of fatigue damage features throughout the entire life cycle through bidirectional information flow interaction. The specific processing includes the following sub-steps: S510 calculates the gating state of the GRU unit. For any GRU computation unit in the BiGRU network, it receives the input features at the current time. (i.e., the normalized feature vector output by the preceding fully connected layer at time...) The value of the hidden state at the previous time step and the hidden state at the previous time step. In this embodiment, the number of hidden layer nodes in the GRU layer is set to 32 to control computational cost while ensuring model capacity. First, the reset gate is calculated using the Sigmoid activation function. and the update gate The reset gate controls whether to ignore the state information from the previous moment to eliminate interference from irrelevant noise (such as transient shocks); the update gate controls how much of the state information from the previous moment is carried over to the current state to maintain the cumulative effect of fatigue damage. The calculation formula is as follows: ; ; In the formula, This indicates the time index, i.e., the moment. For a moment The reset gate vector; For a moment Update the gate vector; This is the Sigmoid activation function, used to map the output to the (0,1) interval; The weight parameter matrix for resetting the gate; To update the weight parameter matrix of the gate; This indicates that the previous state will be hidden. With the current input Perform the splicing operation; This represents the dot product operation (Hadamard Product) of matrices, which is the multiplication of corresponding elements; , These represent the bias vectors corresponding to the update gate and reset gate, respectively, which are used to improve the network's fitting ability.
[0053] S520 calculates candidate states and updates hidden states. This is based on the reset gate. The control calculates the candidate hidden state at the current time. This candidate state reflects the current input. Under the influence of [the previous sentence], a new feature representation is formed by combining historical information filtered by the reset gate. Subsequently, the updated coefficient vector is used... Hidden state from the previous moment and the current candidate state Perform linear interpolation to obtain the hidden state at the current time step. In this embodiment, the update coefficients are... The value corresponds to the update gate vector calculated in step S510. The calculation formula is as follows: ; ; In the formula, Indicates the current time The candidate hidden state; It is the hyperbolic tangent activation function; The weight matrix required to calculate the candidate states; It is an identity matrix (or a vector of all ones), and its dimensions are the same as those of the hidden state. For a moment The updated coefficient vector is used to balance the ratio of memory to forgetting. When the value is close to 1, the model is more inclined to update to the new state. When the value is close to 0, the model tends to retain historical states. This represents the current hidden state calculated by a one-way GRU. This represents the bias vector corresponding to the candidate hidden state calculation; It represents the Hadamardi (or Hadama) stack.
[0054] S530, bidirectional temporal feature fusion. To overcome the limitation of unidirectional GRU in not being able to utilize future information, this embodiment performs GRU computation in parallel in two independent directions. The forward GRU layer runs along the positive direction of the time series (from...). arrive ) processes data to capture the cumulative evolution of fatigue damage; the inverse GRU layer moves in the opposite direction of the time series (from 1) Process the data, using load information from future time points to help determine the damage state at the current time, and capture the inverse causal dependency of crack propagation. Finally, hide the forward state. and reverse hidden state The fusion is performed using a weighted summation method. The fusion calculation formula is as follows: ; In the formula, The hidden state for positive GRU output; The hidden state for the reverse GRU output; The fusion weight matrix represents the positive state. The two weight matrices are learnable parameters, and the network automatically adjusts the contribution ratio of positive and negative information in the final feature through training. For fusion bias term; This is the fused bidirectional hidden state vector.
[0055] S540 generates temporal feature output. The fused bidirectional hidden state vector... Through the weight matrix of the output layer A linear transformation is performed, followed by activation using the Sigmoid function, to obtain the final output vector of the BiGRU network at the current time step. This output vector aggregates the contextual global dependencies of the input sequence, reflecting the relative position and damage level of the current time step throughout the entire fatigue life cycle. The calculation formula is as follows: ; In the formula, The output vector of the BiGRU layer; This is the weight matrix of the output layer; This is the Sigmoid activation function. The output vector... This will serve as input data for the subsequent multi-head self-attention mechanism (MHSA).
[0056] While the BiGRU network can extract forward and backward dependencies in time series, the contribution of load cycles at different times to stress accumulation varies in long-sequence fatigue data. To globally filter out the key features most discriminative for remaining life prediction and address the information dilution problem of long-distance dependencies, this embodiment introduces a multi-head self-attention mechanism after the BiGRU layer. This mechanism adaptively assigns higher weights to key time steps by projecting features into multiple different representation subspaces for parallel processing, thereby generating feature representations containing global contextual information. The specific processing includes the following sub-steps: S610, Multi-head Feature Subspace Linear Projection. The temporal feature sequence output by the preceding BiGRU network is used as the input to this step. To capture feature information from different dimensions, this embodiment sets the number of attention heads. It is 16. For the first... A person's attention The input features are mapped to query, key, and value spaces using independent weight matrices. The calculation formula is as follows: ; In the formula, Indicates the input sequence at the th The feature vectors of each sequence index (i.e., the feature vectors corresponding to the output of step S540) ,in (for sequence indexes); , , The first The corresponding query weight matrix, key weight matrix, and value weight matrix are parameters that can be learned during model training. , , These are the generated query vector, key vector, and value vector, respectively.
[0057] S620, scaling dot product attention weight calculation. Based on query vector. and key vector Attention scores are calculated to quantify the correlation between features at different locations in the sequence. To prevent the gradient vanishing in the softmax function due to excessively large dot product results, a scaling factor is introduced. The dot product result is adjusted. Then, the normalized attention weights are applied to the value vector. , obtained the The output features of each attention head. The calculation formula is as follows: ; In the formula, For the first The output vector of each attention head; Represents the transpose of the key vector; In this embodiment, the dimension of the key vector is... The value is the model's hidden layer dimension divided by the number of attention heads. Softmax is a normalized exponential function used to calculate the probability distribution of attention. This step is actually performed on the value vector of the entire sequence. A weighted sum is performed, and the larger the weight, the more important the corresponding input feature is to the current prediction task.
[0058] S630, multi-head feature concatenation and global fusion. To integrate feature information extracted from different subspaces, the output vectors of all attention heads are... to The vectors are concatenated and then fused using a linear transformation matrix to generate the final global context feature vector. The calculation formula is as follows: ; In the formula, This represents the final output feature of the multi-head self-attention mechanism; This represents a vector concatenation operation; This is the linear projection weight matrix of the output layer, used to map the concatenated high-dimensional features back to the target feature dimension. The feature vector output from this step aggregates the global spatiotemporal dependency information of multi-source heterogeneous data and will be used as the input to the subsequent fully connected layer for the final fatigue life numerical regression prediction.
[0059] After processing by the preceding multi-head self-attention mechanism (MHSA), the model has extracted high-order feature vectors containing global temporal dependencies and weighted by key factors. To map these abstract feature representations to the specific fatigue remaining life numerical space, this embodiment constructs a fully connected regression network after the attention layer. This network module aims to establish a mathematical mapping relationship between the high-dimensional feature space and the one-dimensional life prediction value through nonlinear transformation and dimensionality compression. The specific processing includes the following sub-steps: S640, Fully Connected Layer Feature Mapping and Nonlinear Activation. The global context feature vector output from step S630 is used as the input for this step. In this embodiment, the number of neurons in the fully connected layer is set to 1. The main function of this layer is to reduce the dimensionality of the spliced features output by the multi-head attention mechanism and extract the latent features most closely related to the evolution trend of fatigue damage. The specific operation process is as follows: the input global feature vector is multiplied by the learnable weight matrix of this layer, and the result is added to the bias vector to complete the linear projection of the feature space. In order to give the model the ability to handle nonlinear relationships, a modified linear unit (ReLU) activation function is introduced after the linear transformation. That is, each element in the feature vector after the linear transformation is filtered, negative values are set to zero, and positive values are kept unchanged. In addition, to prevent the model from overfitting on the training data, this embodiment introduces a Dropout strategy after the fully connected layer, with the dropout rate set to 0.1. That is, during training, some neuron connections are randomly and temporarily dropped with a 10% probability, forcing the network to learn more robust distributed features.
[0060] S650, Remaining Life Numerical Regression Output. The latent feature vector output in step S640 is input into the output layer. The output layer consists of a single neuron and aims to map the previously extracted deep latent features into a continuous scalar value. This step performs a linear regression operation, that is, multiplying the latent feature vector by the regression weight matrix of the output layer and adding the bias term of the output layer to directly obtain the normalized fatigue remaining life prediction value. This prediction result represents the health status index of the structure at the current moment. In practical applications, using the life extreme value parameters recorded in step S200, the dimensionless prediction value can be restored to the actual remaining cycle number of the structure through inverse normalization calculation.
[0061] To determine the weights and bias parameters of each component in the aforementioned CNN-BiGRU-MHSA hybrid neural network model—including convolutional kernels, GRU gating units, and fully connected layers—and enable it to accurately map fatigue life values from multi-source sensor data, end-to-end supervised learning based on a labeled training dataset is required. The essence of this process is to find the optimal solution in the parameter space that minimizes the difference between the model's predicted life and the actual remaining life of the structural component.
[0062] When constructing the target loss function, considering that fatigue remaining life prediction is a regression problem, and that in practical applications a greater penalty is needed for larger prediction biases to improve model robustness, this embodiment selects Mean Squared Error (MSE) as the optimization target for network training. The calculation logic of this loss function is as follows: for each sample in the training batch, calculate the square of the difference between its true normalized remaining life label and the model's predicted output value; then sum the squared differences of all samples in the batch and take the average. During the backpropagation phase of training, the model calculates the gradient of this average loss value relative to each learnable parameter in the network, thereby determining the direction and magnitude of parameter adjustment, forcing the model to focus on optimizing those samples with poor fit.
[0063] Regarding the parameter update strategy, considering the high dimensionality of fatigue monitoring data features, strong noise interference, and the complex parameter space of deep networks, this embodiment employs the Adaptive Moment Estimation (Adam) optimization algorithm. This algorithm combines the advantages of the momentum method and the RMSProp algorithm, enabling adaptive adjustment of the learning rate. Its core mechanism utilizes the first-order moment estimate of the gradient (i.e., the exponential moving average of the gradient, reflecting the average direction of the gradient) and the second-order moment estimate (i.e., the exponential moving average of the squared gradient, reflecting the degree of gradient fluctuation) to dynamically scale the update step size of each parameter. Specifically, for parameters with large gradient fluctuations or low update frequencies, the algorithm automatically adjusts its learning rate, thereby ensuring convergence speed while effectively avoiding the model getting trapped in local optima.
[0064] Furthermore, to comprehensively quantify the prediction accuracy and generalization ability of the trained model, this embodiment uses root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²). 2 The following are used as evaluation metrics: Root Mean Square Error (RMSE) is calculated by taking the arithmetic square root of the mean square of the deviations between predicted and actual values, and is used to measure the dispersion of prediction errors, being particularly sensitive to outliers; Mean Absolute Error (MAO) is calculated by taking the average of the absolute values of prediction deviations, directly reflecting the actual magnitude of the prediction error; and the Coefficient of Determination (COD) is calculated by subtracting the ratio of the sum of squared residuals to the sum of squared total squares from 1, characterizing the model's ability to explain the fatigue life decline trend. A COD value closer to 1 indicates a better fit to the data.
[0065] To fully verify the effectiveness and generalization ability of the aforementioned CNN-BiGRU-MHSA hybrid neural network model in the fatigue remaining life prediction task, it is necessary to evaluate the performance of the trained model using test set data independent of the training set. Since the model output layer yields normalized relative values within the [0,1] interval, directly using these values to calculate the error cannot objectively reflect the actual deviation in the number of fatigue cycles. Therefore, inverse normalization must be performed before calculating the metrics. Specifically, using the life extreme value parameters recorded during data preprocessing (i.e., the maximum and minimum life values in the dataset), a linear transformation is used to map the normalized predicted values and corresponding true labels of the model output back to the original physical space, restoring the true number of fatigue life cycles. All subsequent error quantification and performance evaluation are based on the restored true physical quantities to ensure that the evaluation results conform to the measurement standards of actual industrial scenarios.
[0066] For the quantitative evaluation of prediction accuracy, this embodiment selects the root mean square error (RMSE) as the core indicator. This indicator measures the dispersion of the deviation between the predicted value and the actual value. Its calculation logic is as follows: First, for each sample in the test set, calculate the square of the difference between its actual fatigue life value and the model's predicted life value; second, sum the squares of the differences of all test samples and calculate the average; finally, perform an arithmetic square root operation on this average. Because the root mean square error amplifies the error by square during the calculation process, it has a stronger penalty for larger deviations. In the field of fatigue life prediction, avoiding serious prediction inaccuracies (especially overestimation of life) is crucial. Therefore, RMSE can effectively verify the reliability of the model when dealing with extreme working conditions or sudden data changes.
[0067] Meanwhile, to reflect the actual average level of prediction error, this embodiment selects the mean absolute error (MAE) as an auxiliary evaluation index. This index has good robustness, is relatively less affected by individual outliers, and can intuitively reflect the physical magnitude of the prediction deviation. Its calculation logic is as follows: for each sample in the test set, calculate the absolute value of the difference between its actual fatigue life value and the model's predicted fatigue life value; then sum the absolute values of all samples and divide by the total number of samples to obtain the mean absolute error value. The smaller the MAE value, the closer the average distance between the model's predicted value and the actual value, and the higher the overall fitting accuracy of the model.
[0068] Furthermore, to evaluate the model's explanatory power for fatigue life decline trends, this embodiment introduces a coefficient of determination (R²). 2 The coefficient of determination (COP) is used as a consistency evaluation index. This index characterizes the proportion of the model that explains the changes in the dependent variable, reflecting the degree of agreement between the predicted curve and the actual lifetime decay curve. The specific calculation process is as follows: The residual sum of squares is obtained by calculating the sum of squares of the differences between the actual and predicted values of all test samples, and the total sum of squares is obtained by calculating the sum of squares of the differences between the actual and the mean of the actual values of all test samples. The coefficient of determination is obtained by subtracting 1 from the ratio of the residual sum of squares to the total sum of squares. The COP ranges from 0 to 1; the closer the value is to 1, the better the model fits the nonlinear mapping relationship between the input sensor characteristics and the output remaining lifetime.
[0069] See attached document Figure 3 To be continued Figure 5 In one embodiment, the present invention selects multiple sets of sample data from aluminum alloy 2024-T3 thin plate aerospace material, steel wire uniaxial fatigue test and truss Q235 material fatigue test as verification objects, and constructs a verification environment that combines digital simulation and actual measurement, aiming to verify the generalization performance and prediction accuracy of the structural component fatigue life prediction method based on the CNN-BiGRU-MHSA fusion model of the present invention.
[0070] To ensure that those skilled in the art can implement this, a specific computational experimental platform was built in this embodiment. The hardware environment for program execution is configured with a high-performance graphics processing unit, NVIDIA GeForce GTX 1650 SUPER, and the programming implementation and model training tools use Matlab 2023 R2. In the data preprocessing and network construction stages, considering the significant differences in the dimensions of multi-source sample data, and to suppress the risk of overfitting of deep learning models with small samples, this embodiment introduces a dual computational strategy of batch normalization and random deactivation. Specifically, firstly, max-min normalization is performed on the multi-source heterogeneous data to map data of different feature dimensions to a unified interval, thus standardizing the degree of influence of different feature dimensions on the prediction results; then, the processed data is divided into training and test sets, and the training set is input into the CNN-BiGRU-MHSA network structure for iterative training.
[0071] See attached document Figure 3 This diagram displays a two-dimensional scatter regression plot with the true values of the test set on the x-axis and the predicted values on the y-axis. The plot contains three key types of information: scatter points representing the predicted values of the model output, a reference line (Y=X) representing the ideal error-free state, and a fitted line generated based on the prediction results.
[0072] From the appendix Figure 3 The distribution trend shows that the predicted scatter points are highly concentrated on both sides of the reference line Y=X, and the fitted line almost coincides with the reference line. This indicates that the model can accurately establish the nonlinear mapping relationship between the feature activation intensity and physical properties during fatigue. (See attached figure) Figure 4 To be continued Figure 5 To quantify the prediction accuracy, this embodiment uses multiple evaluation metrics to compare the prediction errors of different network models. According to the statistical results in the attached figures, the determination coefficient of the Convolutional Neural Network-BiGRU-MHSA model of this invention reaches 0.97605; simultaneously, the root mean square error is 0.02232, the mean absolute error is 0.01676, the mean square error is 0.00050, the mean absolute percentage error (×10) is 0.19487, and the residual prediction error (×10) is 0.67626. Compared with other network structures in the chart, the model of this invention has the highest determination coefficient, and all error metrics are at the lowest level. The above experimental data show that the fusion model method proposed in this invention has reliable prediction accuracy and model stability when processing fatigue data of various material structures (such as aluminum alloy thin plates and steel wires).
Claims
1. A method for predicting the fatigue life of structural components based on a CNN-BiGRU-MHSA fusion model, characterized in that, Includes the following steps: Acquire multi-source heterogeneous time-series signals and load spectrum data; Interaction parameters are calculated based on geometric boundary parameters and relative crack length. The nominal stress is transformed into the equivalent notched structural stress, and the effective stress intensity factor is synthesized. The equivalent notched structural stress and the effective stress intensity factor are used as characteristics of the crack propagation mechanism. By coupling the tensile damage component, the low-cycle fatigue damage component, and the crack propagation damage component based on the effective stress intensity factor, a comprehensive damage accumulation evolution equation is constructed, and the normalized remaining life is obtained as the fatigue life label. By splicing the multi-source heterogeneous temporal signals with the crack propagation mechanism features, a standard input tensor is constructed. A one-dimensional convolutional neural network is used to process the standard input tensor to extract local spatial feature vectors. The local spatial feature vector is processed using a bidirectional gated loop unit to extract hidden state features; The hidden state features are processed using a multi-head self-attention mechanism to generate weighted fusion features; The weighted fusion features are processed using a fully connected regression layer to output fatigue life prediction values, and the network parameters are optimized based on the fatigue life labels.
2. The method for predicting the fatigue life of structural components based on the CNN-BiGRU-MHSA fusion model according to claim 1, characterized in that, The steps for acquiring multi-source heterogeneous time-series signals and load spectrum data include: The strain sequence is obtained by a strain sensor, the vibration acceleration sequence is obtained by an accelerometer, and the load spectrum data is obtained simultaneously by a force sensor. Using the variation characteristics of the load spectrum data as a benchmark, the strain sequence and the vibration acceleration sequence are aligned on the time axis, and the channel data with different sampling rates are interpolated and resampled to make the number of data points in all channels consistent within a unit time window.
3. The method for predicting the fatigue life of structural components based on the CNN-BiGRU-MHSA fusion model according to claim 1, characterized in that, The steps for calculating interaction parameters based on geometric boundary parameters and relative crack length, and converting nominal stress into equivalent notched structural stress, specifically include: Based on the geometric boundary parameters and the instantaneous relative crack length, the dimensionless geometric correction functions in the tensile mode and the bending mode are calculated respectively. The interaction parameters of the stress field of the notched structure are calculated by combining the stress distribution in the load spectrum data and the dimensionless geometric correction function. Using the interaction parameters and the ligament width determined based on the geometric boundary parameters and the instantaneous relative crack length, the nominal stress under the non-uniform stress field is transformed into the equivalent notch structural stress.
4. The method for predicting the fatigue life of structural components based on the CNN-BiGRU-MHSA fusion model according to claim 1, characterized in that, The specific steps for constructing a comprehensive damage accumulation evolution equation and solving it to obtain the normalized remaining lifetime include: Using the Paris crack propagation model and the effective stress intensity factor, the crack propagation lifetime from the initial length to the fracture failure length is calculated. The tensile damage component, which characterizes the ductility depletion under large strain, the low-cycle fatigue damage component, which characterizes the consumption of cyclic plastic strain, and the crack propagation damage component based on the crack propagation life are added together to form the comprehensive damage accumulation evolution equation. Based on a preset critical damage threshold, the comprehensive damage accumulation evolution equation is solved using an iterative method to obtain the theoretical total fatigue life. The normalized remaining life is calculated based on the theoretical total fatigue life and the accumulated number of load cycles.
5. The method for predicting the fatigue life of structural components based on the CNN-BiGRU-MHSA fusion model according to claim 1, characterized in that, The steps for constructing a standard input tensor by splicing the multi-source heterogeneous temporal signals with the crack propagation mechanism features include: The multi-source heterogeneous time-series signals are spliced with the crack propagation mechanism features to construct a multi-dimensional feature matrix; Each feature dimension of the multidimensional feature matrix is independently subjected to max-min normalization to map the data to the [0,1] interval; Set the time window length and sliding step size, perform time-series slicing on the normalized multidimensional feature matrix to obtain multiple sequence sample matrices, and use a many-to-one mapping strategy to determine the regression target corresponding to each sequence sample matrix. All sequence sample matrices are stacked to generate the standard input tensor, and the standard input tensor and the regression target are divided into the training set and the test set.
6. The method for predicting the fatigue life of structural components based on the CNN-BiGRU-MHSA fusion model according to claim 1, characterized in that, The steps for extracting local spatial feature vectors from the standard input tensor using a one-dimensional convolutional neural network include: The standard input tensor is subjected to multi-channel parallel convolution operations using multiple one-dimensional convolution kernels, and the same padding strategy is used to keep the output size consistent with the input size. The convolution output is nonlinearly transformed using the ReLU activation function; The output after nonlinear transformation is downsampled using a max pooling layer, and the downsampled result is used as the local spatial feature vector.
7. The method for predicting the fatigue life of structural components based on the CNN-BiGRU-MHSA fusion model according to claim 1, characterized in that, The steps for processing the local spatial feature vector using a bidirectional gated recurrent unit to extract hidden state features include: Based on the local spatial feature vector, the degree to which the reset gate control forgets the state information of the previous moment is calculated, and the degree to which the update gate control retains the state information of the previous moment is also calculated. The reset gate and the update gate are used to perform forward GRU computation and reverse GRU computation in parallel. Forward GRU computation is used to capture the historical evolution information of the sequence, and reverse GRU computation is used to capture the future dependency information of the sequence. The hidden state output by the forward GRU calculation is fused with the hidden state output by the reverse GRU calculation to generate the hidden state feature.
8. The method for predicting the fatigue life of structural components based on the CNN-BiGRU-MHSA fusion model according to claim 1, characterized in that, The steps for generating weighted fusion features by processing the hidden state features using a multi-head self-attention mechanism include: The hidden state features are mapped into query vectors, key vectors, and value vectors under multiple attention heads using a linear transformation matrix; Calculate the scaled dot product of the query vector and the key vector, and then normalize it using the Softmax function to obtain the global feature attention weights; The value vector is weighted using the global feature attention weights to obtain the output of each attention head, and the outputs of all attention heads are concatenated and linearly transformed to generate the weighted fusion feature.
9. The method for predicting the fatigue life of structural components based on the CNN-BiGRU-MHSA fusion model according to claim 1, characterized in that, The steps of processing the weighted fusion features using a fully connected regression layer to output fatigue life prediction values include: The weighted fusion features are subjected to dimensionality reduction mapping to compress and integrate the data features into global semantic information for the remaining lifetime; The global semantic information is subjected to linear transformation calculation using a linear weight matrix and bias terms to obtain scalar calculation results; The scalar calculation result is output as the fatigue life prediction value.
10. The method for predicting the fatigue life of structural components based on the CNN-BiGRU-MHSA fusion model according to claim 1, characterized in that, The steps for optimizing network parameters based on the fatigue life label include: Calculate the mean squared error loss between the predicted fatigue life value and the fatigue life label, and use the mean squared error loss as the dominant objective function. An L2 weight decay regularization term is added to the dominant objective function, and a Dropout random deactivation mechanism is introduced before the fully connected regression layer; The learning rate is dynamically adjusted using a piecewise constant descent method, and an adaptive moment estimation optimization algorithm is used to perform parameter update operations on the network parameters according to the dominant objective function until the dominant objective function converges.
Citation Information
Cited By
Gear fatigue life assessment method based on vibration and stress proxy backtracking model
CN122174406A