Method and system for predicting dynamic performance of piezoelectric vibration energy collector
By optimizing the hyperparameter combination using a hybrid deep learning model CNN-LSTM and NOM technology, the accuracy and efficiency issues of dynamic performance prediction for piezoelectric vibration energy harvesters were resolved, enabling accurate prediction under unknown excitation conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-04
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies struggle to accurately predict the dynamic performance of piezoelectric vibration energy harvesters, especially under nonlinear characteristics and unknown excitation conditions. Traditional methods suffer from insufficient modeling accuracy and high computational costs.
By employing a hybrid deep learning model CNN-LSTM combined with differentiable surrogate models and neural optimization machine (NOM) techniques, a nonlinear mapping relationship is constructed through training and optimizing hyperparameter combinations to achieve dynamic performance prediction of piezoelectric vibration energy harvesters.
It achieves accurate prediction of the dynamic performance of piezoelectric vibration energy harvester under unknown excitation conditions, avoiding the modeling complexity and high computational cost of traditional methods, and improving prediction efficiency and adaptability.
Smart Images

Figure CN121859971A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of energy harvesting technology and artificial intelligence, and in particular to a method and system for predicting the dynamic performance of a piezoelectric vibration energy harvester. Background Technology
[0002] Piezoelectric vibration energy harvesting technology can convert the mechanical vibration energy widely present in the environment into electrical energy, providing a green and sustainable solution for low-power electronic devices to achieve self-powering. It has demonstrated significant application value in fields such as the Internet of Things (IoT) and structural health monitoring. The dynamic output performance of piezoelectric vibration energy harvesters is a key indicator for evaluating their power supply capability. However, these performance indicators are influenced by a combination of external excitation conditions and internal structural parameters, and often exhibit significant nonlinear characteristics. This poses a serious challenge to performance prediction and reliability assessment in practical applications.
[0003] Currently, the prediction of the dynamic performance of such devices mainly relies on analytical methods based on physical models or computational methods based on numerical simulation. Physical modeling methods require establishing accurate mechanistic models and performing complex parameter identification, facing dual constraints of modeling accuracy and generalization ability when dealing with nonlinear dynamic behavior or complex structures. Furthermore, while numerical simulation methods can handle more complex operating conditions, their high computational cost and long simulation time make it difficult to meet the practical needs of rapid performance evaluation for a large number of unknown excitation conditions. These two traditional methods have significant shortcomings in prediction efficiency, adaptability, and practicality, thus making it difficult to accurately predict key dynamic performance indicators such as output voltage, power, and displacement of piezoelectric vibration energy harvesters. Summary of the Invention
[0004] This invention provides a method and system for predicting the dynamic performance of a piezoelectric vibration energy harvester, to solve the aforementioned problems existing in the prior art, namely, how to accurately predict the dynamic performance indicators of a piezoelectric vibration energy harvester in the prior art. This invention provides a method for predicting the dynamic performance of a piezoelectric vibration energy harvester, which includes: Obtain the original time series of multiple energy harvesters under vibration excitation; The constructed hybrid deep learning model CNN-LSTM is trained using the original time series data to determine the CNN-LSTM model after initial training. The acquired hyperparameter-performance data pairs are transformed into parameter space to determine the transformed data. Using the transformed data, a differentiable surrogate model is constructed to describe the nonlinear mapping relationship between hyperparameter combinations and the performance of CNN-LSTM models. The trained differentiable surrogate model is embedded into the NOM optimization framework, which includes a trainable input layer and a constraint layer. The hyperparameter combination is optimized by using the gradient descent algorithm to minimize the total loss function. Under the conditions of following boundary constraints and physical rationality constraints, the optimal hyperparameter combination is determined. The total loss function includes the prediction loss of the differentiable surrogate model and the constraint penalty term. The CNN-LSTM model after initial training is reconstructed based on the optimal hyperparameter combination and trained. The sequence features of the energy harvester under vibration excitation are input into the reconstructed and trained CNN-LSTM model to obtain the final prediction results of the dynamic performance of the piezoelectric vibration energy harvester.
[0005] Optionally, the hybrid deep learning model CNN-LSTM model includes a one-dimensional convolutional layer 1D-CNN, an LSTM hidden layer, and a fully connected layer; the one-dimensional convolutional layer 1D-CNN is used to extract local temporal features from the original time series; the LSTM hidden layer is used to obtain the nonlinear mapping relationship between the learning input stimulus and the output response in the local temporal features based on the gating mechanism, and to determine the intermediate features; the fully connected layer is used to perform a linear transformation on the intermediate features to determine the linearly transformed intermediate features.
[0006] Optionally, the extraction of local temporal features from the original time series specifically includes: in, The original time series, For convolution kernel weights, The kernel size is [size]. This represents the convolution operation. It is the ReLU activation function. It represents a local temporal feature.
[0007] Optionally, the boundary constraints and physical rationality constraints specifically include: For each hyperparameter, define a lower bound violation function. and upper bound violation function The boundary constraints are obtained using the following formula: The physical rationality constraints are obtained using the following formula: Where h is the hyperparameter vector, Represents the i-th hyperparameter. To predetermine the lower bound, As a preset upper bound, g low,i (h), g high,i (h) is the boundary violation function. These are the penalty coefficients for boundary constraints and physical rationality constraints, respectively. Let be the violation function of the j-th physical rationality constraint.
[0008] Optionally, the optimization of the hyperparameter combination using the gradient descent algorithm specifically includes: Where η is the learning rate, h t For the current t-th generation hyperparameter combination, This represents the gradient of the total loss function with respect to the hyperparameters.
[0009] Optionally, before reconstructing the CNN-LSTM model based on the optimal hyperparameter combination and training it, the optimal hyperparameter combination is discretized using a nearest neighbor matching strategy; wherein the type of the optimal hyperparameter combination includes discrete parameters, continuous parameters, and log-space parameters.
[0010] Optionally, the sequence features include the base acceleration sequence, the acquisition time sequence, and the external excitation frequency; the prediction results include the end mass vibration displacement sequence of the piezoelectric vibration energy harvester, the output voltage sequence, and the angle of the cantilever beam normal relative to the external excitation direction.
[0011] Optionally, the original time series is normalized, specifically including: The Min-Max scaling method is used to linearly transform the original time series to the [0,1] interval. The normalized data is determined using the following formula: in, The original time series, and These are the minimum and maximum values of the corresponding features in the training set data, respectively. This is the normalized data.
[0012] This invention provides a dynamic performance prediction system for a piezoelectric vibration energy harvester, realizing the aforementioned dynamic performance prediction method for a piezoelectric vibration energy harvester.
[0013] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a method for predicting the dynamic performance of a piezoelectric vibration energy harvester. This method uses CNN to extract local temporal features, which solves the problem of insufficient capture of nonlinear features by traditional physical models. It uses LSTM to capture long-term dynamic dependencies, such as the correlation between the resonance characteristics of the vibration energy harvester and the excitation frequency, avoiding the gradient vanishing problem and thus enhancing the ability to model complex temporal patterns. In addition, by introducing NOM technology, automatic hyperparameter optimization is achieved. Specifically, based on the data after parameter space transformation, a differentiable surrogate model is trained to describe the nonlinear mapping relationship between the hyperparameter combination and the performance of the CNN-LSTM model, to determine the optimal hyperparameter combination. This allows for accurate capture of the nonlinear dynamic characteristics of the piezoelectric vibration energy harvester using limited training data. Then, the initially trained CNN-LSTM model is reconstructed based on the optimal hyperparameter combination and trained to obtain accurate performance prediction results. This method can achieve accurate prediction of the dynamic performance of a piezoelectric vibration energy harvester under unknown excitation without relying on traditional physical modeling or time-consuming numerical simulation. Attached Figure Description
[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0015] Figure 1 A flowchart illustrating a method for predicting the dynamic performance of a piezoelectric vibration energy harvester, provided in an embodiment of the present invention; Figure 2 To construct a sample graph; Figure 3 This is a hierarchical diagram of the model; Figure 4 Optimize the flowchart for the NOM method; Figure 5 This is a diagram of the physical structure of a piezoelectric vibration energy harvester. Figure 6 The training loss iteration graph before optimization; Figure 7 A random sampling distribution diagram of hyperparameter combinations; Figure 8 The training loss curve for a differentiable surrogate model; Figure 9 Optimize the loss curve for each starting point; Figure 10 This is an iterative graph of the optimized training loss. Figure 11 The displacement of the end mass of the piezoelectric vibration energy harvester at 15.5 Hz is the vibration displacement. Figure 12The angle between the cantilever beam of the piezoelectric vibration energy harvester and the external excitation direction at 15.5Hz; Figure 13 The output voltage of the piezoelectric vibration energy harvester at 15.5Hz; Figure 14 The end mass vibration displacement of the piezoelectric vibration energy harvester at 18.3 Hz; Figure 15 The angle between the cantilever beam of the piezoelectric vibration energy harvester and the external excitation direction at 18.3Hz; Figure 16 The output voltage of the piezoelectric vibration energy harvester at 18.3Hz; Figure 17 The absolute error of the end mass vibration displacement of the piezoelectric vibration energy harvester at 15.5Hz; Figure 18 The absolute error of the angle between the cantilever beam of the piezoelectric vibration energy harvester and the external excitation direction at 15.5Hz; Figure 19 The absolute error of the output voltage is 15.5Hz. Figure 20 The absolute error of the end mass vibration displacement is 18.3Hz. Figure 21 The absolute error of the angle between the cantilever beam and the external excitation direction of the piezoelectric vibration energy harvester at 18.3Hz; Figure 22 The absolute error of the output voltage of the piezoelectric vibration energy harvester at 18.3Hz; Figure 23 RMSE, MAE, and 15.5Hz Evaluation indicators; Figure 24 RMSE, MAE, and 18.3Hz Evaluation indicators. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0017] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.
[0018] Example 1 Figure 1 This is a flowchart of a dynamic performance prediction method for a piezoelectric vibration energy harvester provided in an embodiment of the present invention, as shown below. Figure 1 As shown in the figure, this embodiment illustrates a method for predicting the dynamic performance of a piezoelectric vibration energy harvester, including: S1: Obtain the original time series of multiple energy harvesters under vibration excitation.
[0019] For example, a Min-Max normalization method can be used to independently calculate the normalization parameters (min and max) for each input feature based on the training set data. These parameters are the original time series (which may include base acceleration, acquisition time, and external excitation frequency) and the output target (displacement, voltage, angle). The validation and test sets are strictly transformed using the normalization parameters calculated from the training set to eliminate dimensional differences and avoid data leakage. Then, for the normalized time series data, a sliding window method is used to construct supervised learning samples. The time window length is set to N=5, meaning that the input of each sample is a sequence of input features from five consecutive time steps, and the output is the three performance index values for the next time step after that window. Finally, the processed samples can be encapsulated as PyTorch Dataset objects, and a DataLoader can be built to provide a standardized data interface for model training and evaluation.
[0020] For example, the normalization process uses the Min-Max scaling method to linearly transform the original data to the [0, 1] interval. The calculation formula is as follows:
[0021] in, The original data, and These are the minimum and maximum values of the corresponding features in the training set data, respectively. This is the normalized data.
[0022] For the normalized time series data, a sliding window method is used to construct a sample set, such as... Figure 2As shown, the input features of each sample consist of a time-step sequence of length N (N is the preset time window size), and the corresponding output target is the response value of the next time step after that time window. For time-series data of length L, the sample set constructed by the sliding window can be represented as:
[0023] Input sample: , Output target: , Where t is the starting time point of the window, t=1,2,…,(L−N). Each sample The output target contains feature data from N consecutive time steps. This is the target value for the next observation step.
[0024] S2: Train the constructed hybrid deep learning model CNN-LSTM using the original time series data to determine the initial trained CNN-LSTM model.
[0025] For example, a hybrid deep learning model, such as a CNN-LSTM model, may include one-dimensional convolutional layers (1D-CNN), LSTM hidden layers, fully connected layers, and an output layer, with the model hierarchy as follows: Figure 3 As shown, its forward propagation process is as follows: (1) Input layer: Receives a normalized tensor with dimensions (batch_size, TIME_STEPS, INPUT_DIM), where TIME_STEPS is the time window size and INPUT_DIM is the input feature dimension.
[0026] (2) One-dimensional convolutional layer (1D-CNN): A one-dimensional convolutional kernel slides along the time dimension to extract local temporal features. The convolution operation is defined as:
[0027] in, Given the input sequence, For convolution kernel weights, The kernel size is [size]. This represents the convolution operation. The ReLU activation function is used, and the output feature map has dimensions (batch_size, CONV_FILTERS, TIME_STEPS).
[0028] (3) LSTM Hidden Layers: The convolutional feature sequences are input into a multi-layer LSTM network. LSTM controls the flow of information through gating mechanisms (input gate, forget gate, output gate), and its unit update formula is as follows:
[0029] in, Forget gate activation vector, The input gate activation vector, The output gate activation vector, Candidate cell state, This represents the cell state at the current time step. The hidden state at the current time step; , , , Here is the weight matrix for the corresponding gate. , , , This is the bias vector for the corresponding gate; element-wise multiplication ( product), for Activation function.
[0030] (4) Fully Connected Layer: As a feature compression layer, this layer receives the hidden state output from the last time step of the LSTM and maps it to a dimension suitable for the prediction task through a linear transformation. The main function of this layer is to compress the high-dimensional LSTM hidden state features to a dimension suitable for the prediction task, realizing the transformation from abstract temporal features to specific performance metrics. Its mathematical expression is:
[0031] in, These are the output features of the fully connected layer. It is the weight matrix of the fully connected layer. It is the hidden state of the model at the last time step. It is the bias vector of the fully connected layer.
[0032] (5) Output layer: As the final prediction layer of the model, the output of the fully connected layer is directly used as the prediction result (intermediate features after linear transformation), without the need for additional nonlinear transformation. This layer directly maps the compressed features to the predicted values of physical quantities. The output dimension is determined according to the prediction task and directly corresponds to the various dynamic performance indicators of the energy harvester.
[0033] For example, supervised learning of a CNN-LSTM model can be performed using a training set. The core objective is to enable the model to accurately capture the complex nonlinear relationship between input stimulus and output response through a systematic training process. The training process includes three core steps: loss function definition, optimization strategy implementation, and iterative training loop.
[0034] (1) Definition of loss function: By defining a suitable loss function To quantify the difference between model predictions and actual observations, where This represents the model parameters. The loss function, by mathematically transforming the prediction error, provides clear guidance for model optimization, ensuring that the model can specifically improve prediction accuracy during the learning process. The general form of the loss function can be expressed as:
[0035] in, For the model's predicted output, For the true value, This is a function that measures the prediction error of a single sample.
[0036] (2) Implementation of optimization strategy: An advanced optimization algorithm is used to implement the parameter update strategy. The basic update rule can be expressed as follows: in, For learning rate, The algorithm designs the update direction based on the gradient of the current parameters. By adaptively adjusting the update step size of each parameter dimension, it navigates efficiently in the parameter space, accelerating the convergence process while maintaining training stability and effectively avoiding oscillations during optimization.
[0037] (3) Iterative Training Loop: Establish a complete iterative training loop mechanism. The training process proceeds in a cyclical manner, with each cycle including standard steps such as batch data processing, forward propagation calculation, loss evaluation, gradient backpropagation, and model parameter update. The forward propagation process can be represented as:
[0038] The loss is calculated as follows: in, For the average loss of a batch, For batch size, A function to measure the prediction error of a single sample. Let be the predicted value for the i-th sample. Let be the true value of the i-th sample. This batch processing method not only improves training efficiency but also enhances the model's generalization ability.
[0039] The entire training process optimizes the model parameters by minimizing the loss function, enabling the model to gradually learn the accurate mapping relationship from the input sequence to the output response.
[0040] S3: Perform parameter space transformation on the acquired hyperparameter-performance data pairs to determine the data after parameter space transformation; use the data after parameter space transformation to construct a differentiable surrogate model to describe the nonlinear mapping relationship between hyperparameter combinations and CNN-LSTM model performance.
[0041] For example, a three-layer fully connected neural network (multilayer perceptron, MLP) can be selected as the surrogate model, and its mathematical expression is as follows: in, For the first The activation function of the layer, and For the corresponding weights and biases, This is the activation function for the output layer.
[0042] Its specific structure is as follows: Input layer: 7 nodes, corresponding to a 7-dimensional hyperparameter vector Hidden layers: Two 128-node hidden layers using the ReLU activation function. Output layer: 1 node, using the Sigmoid activation function, outputs the prediction performance loss. During the training of the surrogate model, the number of training epochs (surrogate_epochs) was set to 1000, and the learning rate (surrogate_lr) was set to 0.0001. This structure is complex enough to capture the nonlinear relationship between hyperparameters and performance, but not so complex as to cause overfitting. The Sigmoid output layer ensures that the predicted values are within the range [0,1], facilitating subsequent optimization. The surrogate model is trained by minimizing the prediction error, learning the complex nonlinear relationship between hyperparameters and validation set loss, and establishing a differentiable mapping function. The training loss curve of the surrogate model is shown below. Figure 8 As shown, the loss rapidly decreased from 0.133 to below 0.01 within the first 200 epochs, indicating that the model quickly learned the basic mapping relationship. The final loss reached 0.004035, which is relatively small. An MSE of 0.0013 indicates that the predicted values deviate little from the true values, and the accuracy meets the requirements. A MAE of 0.0284 indicates that the mean absolute error is controlled within a reasonable range. An R² of 0.9440 indicates that the model explains 94.4% of the variance, demonstrating excellent goodness of fit. The model completed 1000 training iterations in 4.07 seconds, demonstrating high efficiency.
[0043] S4: Embed the trained differentiable surrogate model into the NOM optimization framework, which includes a trainable input layer and a constraint layer. Optimize the hyperparameter combination using the gradient descent algorithm to minimize the total loss function. Determine the optimal hyperparameter combination while adhering to boundary constraints and physical rationality constraints. The total loss function includes the prediction loss of the differentiable surrogate model and the constraint penalty term.
[0044] For example, this embodiment can employ Neural Optimization Machine (NOM) technology to construct a surrogate mapping relationship between hyperparameters and model performance, thereby achieving efficient optimal hyperparameter search and systematically optimizing the hyperparameters of the CNN-LSTM model to achieve optimal prediction performance. The NOM method optimization process is as follows: Figure 4 As shown.
[0045] For example, first, the key hyperparameter dimensions that affect the performance of the CNN-LSTM model are determined, and an n-dimensional hyperparameter search space is established. The search space encompasses network structure parameters. Training parameters and data processing parameters Three main categories:
[0046] The network structure parameters include hidden layer dimensions, number of convolutional filters, and network depth; training parameters include learning rate and batch size; and data processing parameters mainly involve the time window size of the input sequence. For each type of parameter, a reasonable range of values is determined, forming a multi-dimensional hyperparameter search space.
[0047] Based on the definition of the search space, a multi-dimensional constraint system is established to ensure the feasibility and rationality of the optimization process, specifically including: (1) Boundary constraints: Define the range constraints of each hyperparameter to ensure that the hyperparameter does not exceed the preset boundary during the optimization process: h i ∈[low i high i ], i=1,2,...,n;(2)Physical rationality constraints: Define physical rationality constraints between hyperparameters, such as the convolution kernel size must not exceed the length of the input time series: k_size≤time_steps.
[0048] For example, an initial hyperparameter combination sample set is generated within the hyperparameter search space using a random sampling strategy. For each hyperparameter combination Construct the corresponding CNN-LSTM model And training and validation were performed, among which, [x1,x2,...,x T ] represents the input time series data. Configure the model's hyperparameters and calculate the model performance loss using validation set data. As an evaluation result of this hyperparameter combination:
[0049] in, To determine the number of samples in the validation set, and These are the input features and the true labels, respectively. This evaluation process is entirely based on validation set data to ensure the objectivity and generalization ability of the performance evaluation. This stage obtains a sufficient number of samples to establish a preliminary mapping relationship between hyperparameters and model performance.
[0050] For example, the present invention can achieve unified processing of parameters at different scales by establishing a three-level parameter space transformation system, which may specifically include: The first level is the actual parameter space, which includes the range of values for the original parameters and covers parameters in both linear and logarithmic spaces. The second level is the surrogate model space, which uses a parameter transformer to uniformly convert the actual parameters into a linear space, where logarithmic parameters are transformed using a logarithmic transformation: The third level is the normalization space, which linearly maps the surrogate model space parameters to the [0,1] interval to ensure numerical stability; among them, the logarithmic space parameters are transformed using logarithmic transformation: Transform the logarithmic space parameters from the surrogate model space back to the actual parameter space: Denormalize the normalized space parameters to the proxy model space: The parameter space transformation ensures that all hyperparameters have a uniform data scale and differentiability during the optimization process, supporting the effective implementation of gradient optimization algorithms.
[0051] For example, the hyperparameter-validation set performance data obtained from the initial sampling can be used to... After parameter space transformation, a differentiable mapping function is constructed. As a surrogate model, the surrogate model takes a combination of hyperparameters as input and the performance loss of the predicted model on the validation set as output. By learning the complex nonlinear relationship between the hyperparameter space and performance, it establishes an efficient performance prediction mechanism.
[0052] The surrogate model is designed following the principle of function approximation, constructing a mapping relationship from input to output through multiple layers of nonlinear transformations. The model structure ensures sufficient expressive power to capture the complex relationship between hyperparameters and performance, while maintaining differentiability to provide a foundation for subsequent gradient optimization. The model is trained using supervised learning, with the optimization objective being to minimize the validation set loss and prediction error.
[0053] in, These are the parameters for the proxy model.
[0054] For example, a trained surrogate model is embedded into the NOM optimization framework, which comprises three core components: a trainable input layer, a frozen surrogate model, and a constraint layer. The trainable input layer serves as the starting point of the optimization process, carrying the initial values of hyperparameters and supporting multiple initialization strategies. The frozen surrogate model is a pre-trained surrogate neural network model whose weights and biases remain fixed during optimization, used to predict the model performance corresponding to the hyperparameter configuration. The constraint layer handles boundary constraints and physical rationality constraints, ensuring that the optimization results meet engineering requirements through a penalty function method. Through the collaborative work of these components, the NOM optimization framework achieves efficient gradient optimization of hyperparameters, avoiding the problem of traditional methods easily getting trapped in local optima.
[0055] For example, the trainable input layer, serving as a starting point for optimization and a container for trainable parameters, has the following functions: As an optimization variable carrying layer, it carries the hyperparameter variables to be optimized and serves as the direct optimization object of gradient descent; It supports multi-starting point initialization, allowing optimization to begin from different initial points, increasing the probability of finding the global optimum. A parameter space mapping is achieved by using trainable weights and biases to realize a nonlinear mapping from the input space to the design space. The input space is the direct operation space of the optimization algorithm, containing normalized parameter representations with parameter values in the range [0,1], which facilitates gradient calculation and numerical stability handling. The design space is the actual engineering application space, containing physically meaningful hyperparameter values with clear dimensions and value ranges, which can be directly used to construct CNN-LSTM models. The mathematical expression for the input layer is: in As the initial input vector, and These are the trainable weights and bias parameters.
[0056] For example, to overcome the limitation of traditional optimization methods that are prone to getting trapped in local optima, a hierarchical multi-starting-point optimization strategy is adopted: The first phase is a random exploration phase, where parameters are randomly generated within the hyperparameter space. Starting from a certain point, a relatively large learning rate is used for global exploration, with the goal of discovering multiple potential optimal solution regions; The second stage is the diversification guidance stage. Based on the existing sampled data, points with significant differences are selected as the starting point, and a medium learning rate is used to conduct a refined search of the region, with the goal of enhancing the diversity and coverage of the search. The third stage is the fine-grained search stage, which selects the best performing option. Starting from a certain point, we use a small learning rate to perform local fine-tuning optimization, with the goal of deep mining in the vicinity of a high-quality solution.
[0057] For example, the constraint layer, as a component specifically designed to handle various types of constraints, uniformly employs the constraint penalty function method to process all constraints: Boundary constraint handling, for each hyperparameter Define the lower bound violation function and upper bound violation function The boundary constraint penalty is: Physical rationality constraints define physical relationship constraints between hyperparameters, and the penalty for physical rationality constraints is: All constraints are incorporated into the optimization objective using a penalty function method. The total loss function is: in, To predetermine the lower bound, As a preset upper bound, , These are the penalty coefficients for various constraints. This refers to the amount of violation of physical rationality constraints.
[0058] For example, hyperparameters can be directly optimized using gradient descent to minimize the performance loss on the validation set. The optimization problem can be formulated as:
[0059] The gradient descent algorithm is used to iteratively optimize the hyperparameters under the guidance of the constraint layer: in, For learning rate, For the current t-th generation hyperparameter combination, Let be the gradient of the total loss function with respect to the hyperparameters. The goal of the entire optimization process is to find the optimal combination of hyperparameters on the validation set, ensuring the model's generalization performance. This process transforms the discrete hyperparameter selection problem into an optimization problem in a continuous space, utilizing the gradient information of the surrogate model to guide the search direction, achieving efficient optimal hyperparameter localization.
[0060] For example, the continuous optimal solution obtained through gradient optimization needs to undergo inverse parameter space transformation, which may include: First, performing inverse normalization to map the normalized parameters in the range [0,1] back to the surrogate model space: Then, a transformation from the proxy space to the actual space is performed, and the logarithmic space parameters are exponentially transformed: For example, in neural architecture search, the optimal solution of hyperparameters is obtained through gradient optimization. The hyperparameters reside in a continuous space, while in practical applications they often have discrete characteristics. Therefore, it is necessary to map the continuous optimal solution back to the original discrete hyperparameter space. To obtain practically usable hyperparameter configurations, this method employs a nearest neighbor matching strategy to determine the final discrete hyperparameter combinations:
[0061] in, This represents the preset set of discrete hyperparameter options. This represents the Euclidean distance metric. The strategy applies across all predefined discrete hyperparameter spaces. In the process, find the continuous optimal solution. The closest option will be the final result.
[0062] In practice, differentiated processing methods are adopted based on the definition type of hyperparameters in the search space. For hyperparameters defined as discrete options, discretization is performed using rounding. Specifically, firstly, continuous values are rounded:
[0063] Then, ensure that the rounded value is within a reasonable range of the original options: in, Indicates will Limited to the range Inside, This is a preset set of discrete options for the hyperparameter. This method preserves the integer properties of discrete parameters while avoiding being limited to a finite set of preset options, thus increasing the flexibility of hyperparameter search.
[0064] For hyperparameters (such as the learning rate) defined as continuous intervals, their continuity is preserved, and only boundary clipping is performed: For continuous parameters in a logarithmic space, the pruning operation is performed on a logarithmic scale to preserve its logarithmic properties.
[0065] For the optimal combination of discretized hyperparameters The final model will be trained and its performance fully evaluated to verify its effectiveness and superiority in real-world application scenarios.
[0066] For example, after the NOM neural optimization machine completes the hyperparameter search, it obtains the optimal combination of hyperparameters. This includes three main aspects: network structure features, training parameters, and data processing configuration. Based on these optimal parameters, the CNN-LSTM model is reconstructed and trained.
[0067] in and To determine the optimal hyperparameters Initialize the model weights and bias parameters. The model training objective is to minimize the loss function on the training set:
[0068] The training strategy uses a fixed number of rounds, and monitors the changes in the loss curve to ensure that the model converges to a stable state.
[0069] S5: Reconstruct the initially trained CNN-LSTM model based on the optimal hyperparameter combination, and train it. Input the sequence features of the energy harvester under vibration excitation into the reconstructed and trained CNN-LSTM model to obtain the final prediction result of the dynamic performance of the piezoelectric vibration energy harvester.
[0070] For example, the optimized model can be evaluated using a test set. Test metrics include, in addition to absolute error and loss function, the root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²). Where N is the sample size. For the sample true value, For the sample predicted value, is the average of the true values of all samples; RMSE represents the root mean square error, which measures the standard deviation between the predicted and true values. The smaller the value, the more accurate the prediction; MAE is the mean absolute error, which reflects the average level of the absolute value of the prediction error; R² is the coefficient of determination, which evaluates the proportion of the model that explains the variation of the target variable. The value ranges from [0,1]. The closer the value is to 1, the better the model fit.
[0071] Example 2 This embodiment uses an omnidirectional adaptive piezoelectric vibration energy harvester as the prediction object, and details the dynamic performance prediction process based on a CNN-LSTM model and a NOM neural optimization machine. The physical structure of the omnidirectional piezoelectric vibration energy harvester is as follows: Figure 5 As shown, its core components consist of four parts: a cantilever beam, an end mass, piezoelectric patches, and an adaptive rotation unit. The overall structure is a single cantilever beam. Piezoelectric patches are attached to the surface of the cantilever beam to convert mechanical energy into electrical energy. A stainless steel end mass is fixed at the free end of the beam, and its center of mass is eccentric to the axis of the rotation unit, providing the driving force for adaptive rotation. The rotation unit is connected to a support shaft via a deep groove ball bearing, and the support shaft is fixed to the base. This unit has no external power supply and passively releases the rotational freedom of the cantilever beam around its axis.
[0072] For example, the original dataset in this embodiment can be generated through numerical simulation to simulate the dynamic response of the omnidirectional piezoelectric vibration energy harvester at different external excitation frequencies. The specific data parameter configuration is as follows: (1) Training set: contains simulation data with 7 different parameter combinations. The external excitation frequency covers the range of 14Hz to 20Hz, and the initial angle between the cantilever beam normal and the external excitation direction is 90 degrees to ensure that the model can learn the dynamic characteristics under different working conditions.
[0073] (2) Validation set: contains a set of simulation data with a specific combination of parameters, used for hyperparameter tuning. The parameters are set as follows: external excitation frequency 16.7Hz.
[0074] (3) Test set: contains simulation data with two independent parameter combinations, used for final model performance evaluation. The parameters are set as follows: external excitation frequencies of 15.5 Hz and 18.3 Hz.
[0075] Each parameter combination has a sampling frequency of 1000Hz and a sampling time of 3s.
[0076] The sampling parameters of the training set are shown in Table 1: Table 1 Training set sampling parameters For example, a Min-Max normalization method can be used to independently calculate the normalization parameters (min and max) for each input feature (such as base acceleration, acquisition time, and external excitation frequency) and output target (such as displacement, voltage, and angle) based on the training set data. The validation and test sets are strictly transformed using the normalization parameters calculated from the training set to eliminate dimensional differences and avoid data leakage. Then, for the normalized time series data, a sliding window method is used to construct supervised learning samples. The time window length can be set to N=5, meaning that the input of each sample is a sequence of input features from 5 consecutive time steps, and the output is the three performance index values for the next time step after that window. Finally, the processed samples can be encapsulated as PyTorch Dataset objects, and a DataLoader can be constructed to provide a standardized data interface for model training and evaluation.
[0077] For example, a hybrid deep learning model (CNN-LSTM model) may include an input layer, a one-dimensional convolutional layer (1D-CNN), an LSTM hidden layer, a fully connected layer, and an output layer. Specific parameters can be configured as follows: (1) Input layer: Receives tensors with dimensions (16, 5, 4) after normalization in S1, corresponding to batch size (batch_size), time window length (time_steps), and input feature dimensions (time, base acceleration, initial angle, external excitation frequency).
[0078] (2) One-dimensional convolutional layer (1D-CNN): 128 one-dimensional convolutional kernels of size 5 slide along the time dimension. The ReLU activation function is used to enhance non-linear expressiveness; its mathematical expression is: This effectively alleviates the gradient vanishing problem. The output feature map has dimensions of (16, 128, 5).
[0079] (3) LSTM hidden layer: A three-layer LSTM structure is adopted, with 128 hidden units in each layer; the information flow is controlled by the input gate, forget gate and output gate to effectively capture long-term dependencies.
[0080] (4) Fully connected layer: The 128-dimensional output features of the last time step of the LSTM are mapped to the 3-dimensional physical quantity space through linear transformation, realizing the conversion and compression from abstract temporal features to specific performance indicators.
[0081] (5) Output layer: Receives the 3D physical quantities of the fully connected layer and outputs the results, corresponding to three output targets: end mass vibration displacement, cantilever beam angle and output voltage.
[0082] For example, a CNN-LSTM model can be trained using a training set, with the optimization objective being to minimize the mean squared error (MSE) between the predicted and true values. The training process includes the following steps:
[0083] Loss function: The mean squared error (MSE) is used as the loss function. in For the sample size, For the true value, These are predicted values.
[0084] Optimizer: The Adam optimizer is used, and its parameter update rules are as follows: in, For first-order moment estimation, It is a second-order moment estimate; This is the first-order moment estimate after bias correction. This is the second-order moment estimate after bias correction; , The exponential decay rate is estimated by moments; It is the numerical stability constant; This represents the parameter value in the t-th iteration. This represents the gradient of the loss function with respect to the parameters.
[0085] During training, the parameters can be set to 300 training epochs and 16 batch sizes. In each training epoch, the model processes each batch of data sequentially, performs forward propagation to calculate the predicted output and loss value, calculates the gradient through backpropagation, and updates the model parameters using the optimizer.
[0086] This embodiment employs NOM (Neural Optimization Machine) technology to systematically optimize the hyperparameters of the CNN-LSTM model to achieve optimal prediction performance. Generally, for the task of predicting the dynamic performance of energy harvesters, a search space containing seven key hyperparameters is defined. The range of each parameter is determined based on the following considerations:
[0087] (1) Network structure parameters: Hidden layer dimensions (16, 256): Covers the model capacity requirements from basic to complex. The upper limit of 256 is to balance GPU memory limitations and training efficiency.
[0088] Number of convolutional filters (16, 256): Matched to the hidden layer dimension, ensuring that feature extraction capability is commensurate with model complexity. Too many filters increase computational burden, while too few may limit feature extraction capability.
[0089] Network depth (1, 3): Limited to 3 layers or less to avoid gradient vanishing and training difficulties caused by excessively deep networks. For time series prediction tasks, excessively deep LSTM networks may actually reduce training stability.
[0090] Kernel size (3, 15): Covers short- to medium-term temporal dependencies, with a maximum of 15 to ensure it does not exceed the minimum time window length. Smaller kernels are suitable for capturing local features, while larger kernels can capture dependencies over a longer period.
[0091] (2) Training parameters: Learning rate (logarithmic space: 1e-4 to 1e-2): This range covers a reasonable spectrum for typical deep learning tasks, avoiding excessively large rates that cause oscillations or excessively small rates that lead to slow convergence. Using logarithmic space sampling better reflects the sensitive characteristics of the learning rate.
[0092] Batch size (16, 256): Choose within memory limits to balance training stability and convergence speed. Smaller batches may increase randomness and help escape local optima.
[0093] (3) Data processing parameters: Input sequence time window size (5, 15): Based on preliminary experiments, it was found that windows with more than 15 time steps offer limited performance improvement but significantly increase computational cost. Windows that are too short may fail to capture the complete dynamics, while windows that are too long introduce redundant information.
[0094] The hyperparameter space is defined as shown in Table 2: Table 2 Hyperparameter Space Definition For example, a random sampling strategy can be used to generate 200 initial hyperparameter combinations. Choosing 200 samples is based on a balance between computational resources and the training requirements of the surrogate model; too few samples make it difficult to build an accurate surrogate model, while too many significantly increase computational costs. For each hyperparameter combination:
[0095] (1) Dynamically reconstruct the data loader based on the current time window size time_steps and batch size batch_size to ensure that the data format matches the hyperparameters; (2) Construct the corresponding CNN-LSTM model and train it for 50 rounds; (3) Calculate the normalized MSE loss on the validation set as a performance metric; This process establishes a preliminary mapping relationship between hyperparameters and model performance, and obtains a sufficiently diverse range of training samples. While ensuring the model's expressive power, the limitations of actual device computing capabilities are fully considered. For example... Figure 6 After 50 training epochs, the loss was found to be stable. Therefore, the training epoch for a single evaluation was set to 50, which ensured sufficient model convergence while keeping the total optimization time within an acceptable range. 150 sets of samples provided ample learning data for the surrogate model. The hyperparameter combination was randomly sampled as follows: Figure 7 As shown. For example, based on 200 sets of hyperparameter-performance data pairs, parameter space transformation is performed: the hyperparameters in the actual parameter space are transformed to the surrogate model space, where logarithmic space parameters such as the learning rate are logarithmically transformed, while other parameters retain their original values; then the surrogate model space parameters are normalized to the [0,1] interval to ensure that the parameters are comparable at different scales.
[0096] For example, the trained surrogate model is embedded into the NOM optimization framework, which comprises three core components: a trainable input layer, a frozen surrogate model, and a constraint layer. The Adam optimizer is used for 200 steps of gradient descent optimization. The 200-step optimization is chosen to strike a balance between search sufficiency and computational efficiency: too few steps may not converge to the optimal solution, while too many steps increase unnecessary computational overhead.
[0097] For example, in the trainable input layer, a hierarchical multi-starting-point optimization strategy is adopted: first, six random starting points are generated for global exploration, with a learning rate of 0.001; then, four diverse starting points are selected based on sampled data for region refinement search, with a learning rate of 0.001; finally, three optimal starting points are selected for local fine-tuning optimization, with a learning rate of 0.0005. The learning rate was determined experimentally to ensure convergence stability without being too slow.
[0098] For example, in the constraint layer, a constraint penalty function method is uniformly used to handle all constraints. For boundary constraints, a boundary constraint is defined for each hyperparameter, with a penalty coefficient set to 1. For physical rationality constraints, the penalty coefficient is set to 1, and a convolution kernel size constraint is defined, specifically implemented as follows:
[0099] when When this happens, the corresponding penalty will be triggered.
[0100] For example, gradient descent algorithms (such as the Adam optimization algorithm) are used to iteratively optimize hyperparameters. The optimization objective is to minimize the total loss function consisting of the loss predicted by the surrogate model and the constraint penalty term. The optimization loss curves at each starting point are shown below. Figure 9As shown, random starting points (R1-R6) have a higher initial loss (around 0.0035) but exhibit the strongest exploration ability, converging to a lower loss in the later stages; diverse starting points (D1-D4) have better initial positions and a relatively smooth convergence path, demonstrating the advantage of being guided by sampled data; fine starting points (F1-F3) start from the region with the best performance and have the lowest initial loss. During the optimization process, some starting points have constraint violations (penalty value > 0), and the constraint penalties for all starting points rapidly decrease to near zero within 100 steps. The objective function is the main contributor to the total loss, and the constraint penalties are almost zero in the later stages of optimization, indicating that the physical rationality constraints are effectively satisfied. The three loss components of each starting point tend to be consistent in the later stages, suggesting that a global or strong local optimum may have been found.
[0101] For example, the continuous optimal solution obtained by gradient optimization can be discretized, specifically including: (1) Discrete parameters (hidden layer dimension, number of convolutional filters, etc.): rounded to the nearest integer and clipped to the range of preset options; (2) Continuous parameters: maintaining continuity and only performing boundary clipping; (3) Logarithmic space parameters (learning rate): performing clipping operation on the logarithmic scale; The selection of discretization strategy is based on practicality and interpretability considerations, ensuring that the final hyperparameter configuration can be directly used for model training, while maintaining consistency with the original search space. The optimal solution of hyperparameters and the discretization results are shown in Table 3:
[0102] Table 3. Optimal solutions and discretization results for hyperparameters. The following example illustrates the model training and performance verification process based on NOM-optimized parameters.
[0103] For example, after the NOM neural optimization machine completes the hyperparameter search, it obtains the optimal combination of hyperparameters. Specifically, the parameters include a hidden layer dimension of 151, a learning rate of 0.001466, 154 convolutional filters, a kernel size of 11, 2 network layers, a batch size of 176, and a time window size of 11. Based on these optimal parameters, the CNN-LSTM model is reconstructed and trained using supervised learning, aiming to minimize the mean squared error loss on the training set. Training is set to 300 epochs, and convergence is monitored by recording the loss value for each epoch. Figure 10 As shown.
[0104] For example, the performance of the NOM-optimized CNN-LSTM model can be fully validated using independent test sets (external excitation frequencies of 15.5Hz and 18.3Hz). Validation includes prediction curve comparison, absolute error analysis, and multi-index quantitative evaluation to comprehensively examine the model's prediction accuracy and generalization ability under unknown excitation conditions.
[0105] like Figures 11 to 16 As shown, the comparison between the model's predicted and actual values for the end-mass vibration displacement, the angle between the cantilever beam normal and the external excitation direction, and the output voltage of the omnidirectional adaptive piezoelectric vibration energy harvester at two test frequencies is presented. Figures 8 to 13 As can be seen, the model's predicted curves and the actual response curves closely match at excitation frequencies of 15.5Hz and 18.3Hz, indicating that the model can accurately capture the dynamic response characteristics of the omnidirectional adaptive piezoelectric vibration energy harvester under different frequency excitations. Even in regions of significant nonlinearity, such as peaks, valleys, and abrupt changes in the rate of change, the model's predicted values remain highly consistent with the actual values. This demonstrates that the CNN-LSTM hybrid model effectively combines the advantages of CNN in extracting local temporal patterns with LSTM in capturing long-term dynamic dependencies.
[0106] Figures 17 to 19 The time-series distribution of the absolute error of the model to each output target at an excitation frequency of 15.5 Hz is shown. Figure 17 The absolute error of the end mass vibration displacement is basically controlled within ±0.15cm, and the error curve fluctuates slightly around the zero line. Figure 18 The absolute error in the predicted cantilever beam angle is mainly distributed within the range of ±5°. Figure 19 The absolute error of the displayed output voltage is basically maintained within ±1.2V.
[0107] Figures 20 to 22 The error distribution at an excitation frequency of 18.3 Hz is shown. Figure 20 The predicted error of the end-mass vibration displacement converges to within ±0.12 cm. Figure 21 The display shows a significant improvement in the cantilever beam angle prediction error, with the fluctuation range narrowed to ±3°; Figure 22 The voltage prediction error has been reduced to within ±1.0V.
[0108] All error curves exhibit stable fluctuations without any trend of accumulation or divergence over time, and maintain similar fluctuation amplitudes and distribution characteristics at different excitation frequencies. This stable error pattern indicates that the model has a consistent ability to capture the dynamic characteristics of the piezoelectric vibration energy harvester, and will not produce significant performance fluctuations due to changes in excitation conditions, thus verifying the reliability of the model in practical applications.
[0109] To quantify model performance, the root mean square error (RMSE), mean absolute error (MAE), coefficient of determination (R²), and percentage of the maximum amplitude of the true signal were calculated at two test frequencies, respectively. Figure 23 , 24 As shown, the specific values are shown in Tables 4 and 5: Table 4 Performance evaluation of various output indicators at a test frequency of 15.5Hz Table 5 Performance evaluation of various output indicators at a test frequency of 18.3Hz The evaluation results show that: At a test frequency of 15.5Hz, the model exhibits extremely high prediction accuracy and good fit. The displacement prediction has an RMSE of 0.0635 and a MAE of 0.0462, with this error representing only 3.06% of the maximum amplitude of the true signal. Simultaneously, its R² value is as high as 0.9947, indicating that the model can explain 99.47% of the variation in this variable. The angle prediction has an RMSE of 2.5602 and a MAE of 2.1755, with a relative error accounting for 2.44% and an R² value of 0.9897, demonstrating the model's strong ability to capture nonlinear relationships. The voltage prediction achieves an R² of 0.9967 and an MAE of 0.5135, with a relative error accounting for 2.66%. These data collectively demonstrate that the model possesses excellent reliability and accuracy under specific operating conditions.
[0110] At the 18.3Hz test frequency, the model exhibited superior generalization ability, with several performance indicators outperforming those at 15.5Hz, demonstrating good universality. The absolute error of displacement prediction further decreased (RMSE=0.0539, MAE=0.0431), and the R² value stabilized at 0.9936. The R² value of angle prediction increased from 0.9897 to 0.9961, with a significant reduction in prediction error (RMSE from 2.5602 to 1.5833, MAE from 2.1755 to 1.1954), and its relative error percentage decreased significantly from 2.44% to 1.34%. The R² value of voltage prediction also increased to 0.9970, while its MAE decreased to 0.3637, and its relative error percentage decreased to 2.32%.
[0111] Evaluation results show that the proposed model achieves R² values higher than 0.989 for all output targets at both testing frequencies, with the relative proportion of the average prediction error being less than 3.8%. In particular, most indicators are superior at the 18.3Hz operating condition. This fully verifies that the model not only has excellent fit but also exhibits strong stability and generalization ability in the face of changing operating conditions, providing a reliable performance guarantee for its application in practical engineering.
[0112] To address the limitations of traditional methods for predicting the dynamic performance of piezoelectric vibration energy harvesters, which rely on complex physical modeling, face difficulties in parameter identification, and suffer from high computational costs and slow response to unknown excitation conditions in numerical simulation, this invention normalizes and applies a sliding window to the excitation-response time-series data to construct a structured sample set. It then extracts local temporal features using a CNN, captures long-term dynamic dependencies using an LSTM, and introduces NOM (Normally Oscillating Parameter) technology for automatic hyperparameter optimization. This allows for accurate capture of the nonlinear dynamic characteristics of piezoelectric vibration energy harvesters using limited training data. This method achieves rapid and accurate prediction of the dynamic performance of piezoelectric vibration energy harvesters under unknown excitations without relying on traditional physical modeling or time-consuming numerical simulations. Experimental results show that the NOM-optimized CNN-LSTM model significantly improves prediction accuracy across multiple key indicators such as displacement, angle, and voltage, demonstrating clear advantages in improving prediction efficiency and adaptability to nonlinear systems.
[0113] Example 3 The above describes a method for predicting the dynamic performance of a piezoelectric vibration energy harvester, provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding system for predicting the dynamic performance of a piezoelectric vibration energy harvester to implement the above-described method for predicting the dynamic performance of a piezoelectric vibration energy harvester.
[0114] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this invention.
Claims
1. A method for predicting the dynamic performance of a piezoelectric vibration energy harvester, characterized in that, include: Obtain the original time series of multiple energy harvesters under vibration excitation; The constructed hybrid deep learning model CNN-LSTM is trained using the original time series data to determine the CNN-LSTM model after initial training. The acquired hyperparameter-performance data pairs are transformed into parameter space to determine the transformed data. Using the transformed data, a differentiable surrogate model is constructed to describe the nonlinear mapping relationship between hyperparameter combinations and the performance of CNN-LSTM models. The trained differentiable surrogate model is embedded into the NOM optimization framework, which includes a trainable input layer and a constraint layer. The hyperparameter combination is optimized by using the gradient descent algorithm to minimize the total loss function. Under the conditions of following boundary constraints and physical rationality constraints, the optimal hyperparameter combination is determined. The total loss function includes the prediction loss of the differentiable surrogate model and the constraint penalty term. The CNN-LSTM model after initial training is reconstructed based on the optimal hyperparameter combination and trained. The sequence features of the energy harvester under vibration excitation are input into the reconstructed and trained CNN-LSTM model to obtain the final prediction results of the dynamic performance of the piezoelectric vibration energy harvester.
2. The dynamic performance prediction method for a piezoelectric vibration energy harvester as described in claim 1, characterized in that, The hybrid deep learning model CNN-LSTM includes a one-dimensional convolutional layer 1D-CNN, an LSTM hidden layer, and a fully connected layer. The one-dimensional convolutional layer 1D-CNN is used to extract local temporal features from the original time series. The LSTM hidden layer is used to obtain the nonlinear mapping relationship between the learning input stimulus and the output response in the local temporal features based on the gating mechanism, and to determine the intermediate features. The fully connected layer is used to perform a linear transformation on the intermediate features to determine the intermediate features after the linear transformation.
3. The dynamic performance prediction method for a piezoelectric vibration energy harvester as described in claim 2, characterized in that, The extraction of local temporal features from the original time series specifically includes: in, The original time series, For convolution kernel weights, The kernel size is [size]. This represents the convolution operation. It is the ReLU activation function. It represents a local temporal feature.
4. The dynamic performance prediction method for a piezoelectric vibration energy harvester as described in claim 1, characterized in that, The boundary constraints and physical rationality constraints specifically include: For each hyperparameter, define a lower bound violation function. and upper bound violation function The boundary constraints are obtained using the following formula: The physical rationality constraints are obtained using the following formula: Where h is the hyperparameter vector, Represents the i-th hyperparameter. To predetermine the lower bound, As a preset upper bound, g low,i (h), g high,i (h) is the boundary violation function. These are the penalty coefficients for boundary constraints and physical rationality constraints, respectively. Let be the violation function of the j-th physical rationality constraint.
5. The dynamic performance prediction method for a piezoelectric vibration energy harvester as described in claim 1, characterized in that, The optimization of hyperparameter combinations using the gradient descent algorithm specifically includes: Where η is the learning rate, h t For the current t-th generation hyperparameter combination, This represents the gradient of the total loss function with respect to the hyperparameters.
6. The dynamic performance prediction method for a piezoelectric vibration energy harvester as described in claim 1, characterized in that, Before reconstructing and training the CNN-LSTM model based on the optimal hyperparameter combination, the optimal hyperparameter combination is discretized using a nearest neighbor matching strategy; wherein, the type of the optimal hyperparameter combination includes discrete parameters, continuous parameters, and log-space parameters.
7. The dynamic performance prediction method for a piezoelectric vibration energy harvester as described in claim 1, characterized in that, The sequence features include the base acceleration sequence, the acquisition time sequence, and the external excitation frequency; the prediction results include the end mass vibration displacement sequence of the piezoelectric vibration energy harvester, the output voltage sequence, and the angle of the cantilever beam normal relative to the external excitation direction.
8. The method for predicting the dynamic performance of a piezoelectric vibration energy harvester as described in claim 1, characterized in that, The original time series is normalized, specifically including: The Min-Max scaling method is used to linearly transform the original time series to the [0,1] interval. The normalized data is determined using the following formula: in, The original time series, and These are the minimum and maximum values of the corresponding features in the training set data, respectively. This is the normalized data.
9. A dynamic performance prediction system for a piezoelectric vibration energy harvester, characterized in that, The method for predicting the dynamic performance of a piezoelectric vibration energy harvester as described in any one of claims 1-8 is implemented.