Electromechanical system reliability evaluation method based on deep reinforcement learning and importance sampling

By combining deep reinforcement learning with importance sampling, the challenges of multimodal data fusion and mechanism modeling for reliability assessment of complex electromechanical systems are solved, efficient and accurate reliability assessment is achieved, and the design and optimization of electromechanical systems are supported.

CN119885759BActive Publication Date: 2025-10-17UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411990192.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-10-17
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The reliability assessment of complex electromechanical systems faces challenges such as multimodal data fusion, difficulty in mechanism modeling, and complex high-dimensional nonlinear calculations. Existing methods are unable to efficiently and accurately handle the reliability assessment of complex systems.

Method used

Combining deep reinforcement learning and importance sampling methods, by establishing a functional block diagram of a complex electromechanical system, analyzing the failure modes of key components, and constructing a multi-physics field performance simulation analysis model, the deep learning model is used to fuse sensor data, combined with the important sampling strategy of adaptive MCMC, to optimize operating parameters to improve the accuracy and efficiency of reliability assessment.

Benefits of technology

It improves the efficiency and accuracy of reliability assessment of complex electromechanical systems, reduces computing resources and time requirements, and supports the design, maintenance, and optimization decisions of electromechanical systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119885759B_ABST
    Figure CN119885759B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on deep reinforcement learning and important sampling electromechanical system reliability evaluation method, first establish complex electromechanical system reliability block diagram, then analyze the sensor data of each level operating state of system, build electromechanical system sensor data and multi-physics field performance simulation analysis fusion dataset, then establish the reliability evaluation model based on deep reinforcement learning, based on the important sampling strategy of adaptive MCMC, the reliability of electromechanical system is evaluated, finally, the reliability sensitivity of electromechanical system operating parameters is analyzed, the operating parameters are optimized based on particle swarm intelligence algorithm, and the effectiveness of the optimization strategy is verified through environmental information.The method of the application combines the intelligent decision-making capability of deep reinforcement learning and the adaptive sampling strategy of important sampling, reduces the computing resources and time required for evaluation, while improving the accuracy and reliability of the evaluation results, to better support the design, maintenance and optimization decisions of electromechanical systems.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of reliability evaluation and optimization of complex mechatronic systems, and particularly relates to a mechatronic system reliability evaluation method based on deep reinforcement learning and important sampling. BACKGROUND

[0002] A complex mechatronic system is composed of mechanical structures, electrical modules, control modules and other components, and its failure modes involve multiple disciplines and interact with each other. On the one hand, the reliability evaluation of the complex mechatronic system is restricted by factors such as funds, environment and technology, and a large number of tests in a statistical sense cannot be carried out. On the other hand, due to the complexity of the mechatronic system, full-size refined calculation cannot be carried out under incomplete knowledge, and the precision of the physical model is not enough. The reliability evaluation of the complex mechatronic system faces difficulties such as multi-modal data fusion, difficulty in mechanism modeling, and complexity of high-dimensional nonlinear calculation. The reliability evaluation method based on the combination of the proxy model technology and the adaptive sampling strategy can improve the efficiency and precision of the reliability evaluation of the complex system by using advanced machine learning algorithms, mathematical fitting methods and reasonable sampling techniques. This method has become an important method for reliability evaluation of complex systems, but how to process multi-modal data and effectively fuse data and mechanism is still a difficulty and challenge faced by the reliability evaluation of complex mechatronic systems.

[0003] Deep reinforcement learning can better handle complex environments and decision-making problems by using the nonlinear mapping ability of deep neural networks. Secondly, the deep reinforcement learning algorithm can learn general patterns and strategies from a large amount of data by training with deep neural networks, thereby having good generalization ability. Compared with traditional reinforcement learning algorithms, deep reinforcement learning algorithms can automatically discover features and strategies through autonomous learning and optimization, and are more flexible and efficient in handling complex problems. However, deep reinforcement learning algorithms usually require a large number of training samples and computing resources to achieve good results. Important sampling technology can identify key areas that have a significant impact on reliability (such as areas close to the limit state of the structure) by analyzing the characteristics of complex systems, and concentrate sampling in areas that have a large impact on reliability and can reduce variance, thereby obtaining key samples that can reflect the reliability characteristics of the system, improving sampling efficiency and reducing reliability evaluation variance. Therefore, in order to overcome the shortcomings of existing methods and improve the efficiency and reliability of reliability evaluation, a reliability evaluation method based on deep reinforcement learning and important sampling is studied. SUMMARY

[0004] To solve the above technical problems, the present application provides a kind of based on deep reinforcement learning and important sampling electromechanical system reliability evaluation method, can efficiently, accurately process the reliability evaluation of complex electromechanical system uncertainty, by combining the intelligent decision-making ability of deep reinforcement learning and the probability optimization strategy of important sampling, reduce the required computing resources and time of evaluation, while improving the accuracy and reliability of evaluation results, to better support the design, maintenance and optimization decision of electromechanical system.

[0005] The technical scheme adopted by the present application is: a kind of based on deep reinforcement learning and important sampling electromechanical system reliability evaluation method, specific steps are as follows:

[0006] S1, establish complex electromechanical system function block diagram, analyze the failure mode of key components and its influence, and then establish system reliability block diagram;

[0007] First, complex electromechanical system is decomposed according to its function and structure level, identify each subsystem, module and key component, establish complex electromechanical system function block diagram;Again, based on FEMA method, the failure mode of key components of complex electromechanical system and its influence are analyzed, and finally, on the basis of completing electromechanical system function block diagram and failure mode of key components and its influence analysis, the function block diagram is converted into reliability block diagram;

[0008] S2, analyze system each level running state sensor data, and according to the failure mode of key components in step S1, establish key component multi-physics field performance simulation analysis model based on virtual prototype technology, build electromechanical system sensor data and multi-physics field performance simulation analysis fusion dataset;

[0009] S3, according to the electromechanical system sensor data and multi-physics field performance simulation analysis fusion dataset established in step S2, establish reliability evaluation model based on deep reinforcement learning;

[0010] S4, based on the reliability evaluation model established in step S3, establish important sampling strategy based on adaptive MCMC, evaluate the reliability of electromechanical system through reinforcement learning and environment interaction;

[0011] S5, based on step S4, analyze the reliability sensitivity of electromechanical system operating parameters, optimize operating parameters based on particle swarm intelligent algorithm, improve system reliability, and further verify the effectiveness of optimization strategy through environmental information;

[0012] Wherein the operating parameters include: motion trajectory, motor output torque, control parameter, joint output angle.

[0013] 2. The electromechanical system reliability evaluation method based on deep reinforcement learning and important sampling according to claim 1, wherein step S2 is as follows:

[0014] S21. Analyze the operating status sensor data of each level of the electromechanical system and define fault characteristic indicators;

[0015] The operating status sensor data at each level of the electromechanical system is multi-source heterogeneous information, including: motor temperature and current, motor output torque, end-effector motion image, vibration amplitude, acceleration, joint output angle and torque. First, the collected raw data is denoised to improve data quality. Second, key features are extracted from the denoised data, including: frequency domain features, time domain features, and time-frequency domain features.

[0016] Then, based on the operating status and potential failure modes of each level of the system, the fault characteristic indicator space is defined and associated with each component in the reliability block diagram;

[0017] The fault characteristic indicator space includes: motor temperature and current, motor output torque, end effector vibration amplitude, acceleration, joint output angle and torque;

[0018] S22. Based on step S21, according to the main failure modes of the complex electromechanical system, a multi-physics performance simulation analysis model of key components of the complex electromechanical system is established based on virtual prototyping technology to analyze the performance response of the key components of the system under multi-physics coupling;

[0019] S23. Construct a fusion data set of electromechanical system sensor data and multi-physics field performance simulation analysis.

[0020] 3. The electromechanical system reliability assessment method based on deep reinforcement learning and importance sampling according to claim 1, wherein step S3 is specifically as follows:

[0021] S31. Define the reliability assessment problem;

[0022] According to the main failure modes of complex electromechanical systems, the reliability assessment problem is defined as follows:

[0023] R k =P{g k (x1,x2,…,x n )≤0,k=1,2,…,l} (1)

[0024] Among them, R k represents the reliability of the kth performance response, P(·) represents the probability, g k (x1,x2,…,x n ) represents the kth performance limit state function, k=1,2,…,l; g k (x)≤0 means the performance response meets the requirements, g k(x) > 0 indicates performance failure, x = [x1, x2, …, x n ] represents a system input parameter vector;

[0025] S32, define the agent action in deep reinforcement learning and the reinforcement learning environment;

[0026] In deep reinforcement learning, the agent action is defined as optimizing the deep learning network parameters, a complex mechatronic system performance evaluation response prediction model based on deep learning is established, and the environment is a mechatronic system sensor data and performance response simulation data set at each level;

[0027] The multi-source heterogeneous sensor data is fused by using a deep learning model, and a deep learning model combining convolutional neural network CNN and long short-term memory LSTM network and fully connected neural network is constructed for image information and time series feature data in multi-source information, and a complex mechatronic system performance response evaluation and prediction model is established;

[0028] S33, define state-action pair value function, state value function and environment reward function, and evaluate the accuracy of the complex mechatronic system performance response evaluation and prediction model;

[0029] The state-action pair value function is defined based on Bellman equation, and the expression is as follows:

[0030] Q new (s t ,a t )=Q(s t ,a t )+α1[Aw(s t ,a t )+γminQ'(s t+1 ,a t+1 )-Q(s t ,a t )] (2)

[0031] Where s t represents the state at time t, a t represents the action at time t, Q(s t ,a t ) represents the state-action pair value function at time t, defined as formula (3), Aw(s t ,a t ) represents the reward function, defined as formula (4), Q new (s t ,a t ) represents the new state-action pair value function, γ represents the discount factor, Q'(s t+1 ,a t+1 ) represents the state-action pair value function at time t+1, and α1 represents the learning rate;

[0032] Q(s t ,a t )=||g tru -g pre || 2 (3)

[0033] wherein g pre represents the prediction result of the deep learning model, and g tru represents the real value of the environmental feedback or the simulation analysis result;

[0034]

[0035] wherein Δx represents the variation of the input parameter of the electromechanical system performance, represents the gradient of the state-action pair value function with respect to the input parameter;

[0036] The reinforcement learning and the policy gradient algorithm are fused to optimize the deep learning network parameters, and the evaluation network parameter update expression is as follows:

[0037]

[0038] wherein w represents the deep learning network weight, and α2 represents the learning rate.

[0039] 4. The electromechanical system reliability evaluation method based on deep reinforcement learning and importance sampling according to claim 1, characterized in that the step S4 is specifically as follows:

[0040] S41. According to the long-term return function, an importance sampling strategy based on an adaptive Markov chain Monte Carlo algorithm is established;

[0041] Exploration and learning are performed through the reinforcement learning, on the basis of which the cumulative reward is maximized, and the adaptive sampling strategy is optimized; G t represents the long-term return of the state-action pair, and the expression is as follows:

[0042]

[0043] wherein j represents the jth action-state pair, G(t+1) represents the return of the t+1 action-state pair, and γ represents the discount factor;

[0044] Given the current sample x(t), the transition of the next state x(t+1) follows the transition probability T(x(t+1)|x(t)), and satisfies the detailed balance condition, and the expression is as follows:

[0045] P(x(t))T(x(t+1)|x(t))=P(x(t+1))T(x(t)|x(t+1)) (7)

[0046] The transition probability is calculated by formula (8), and the expression is as follows:

[0047] T(x(t+1)|x(t))=P(s new ,r|s,a)=P[S(t+1)=s new ,Aw(t+1)=r|S(t)=s t ,A(t)=a t ] (8)

[0048] Wherein, r represents the reward of state-action pair, s new represents the new state at t+1, according to the definition of state-action pair value function, the expression of state-action pair value function is as follows:

[0049]

[0050] Wherein, q π (s,a) represents the state-action pair value function, E(σ) represents the expected function, V π (·) represents the state value function, and the expression is as follows:

[0051]

[0052] S42, based on the sampling strategy of step S41, the system reliability is evaluated through reinforcement learning and environment interaction;

[0053] Based on the reliability evaluation of sampling method, the indicator function is defined first, and the expression is as follows:

[0054]

[0055] Wherein, represents the kth failure mode, m represents the mth sampling point, I k,m represents the indicator function;

[0056] Then, the reliability of each failure mode is evaluated, and the reliability R k of the kth performance is as follows:

[0057]

[0058] Wherein, N represents the total sampling size.

[0059] 5. The mechanical and electrical system reliability evaluation method based on deep reinforcement learning and important sampling according to claim 1, wherein the step S5 is specifically as follows:

[0060] Based on the particle swarm optimization (PSO) intelligent algorithm, the running parameters of the system are optimized, and the expression is as follows:

[0061]

[0062] wherein x * represents the optimal set of operating parameters, R sys (x) represents the system reliability, R k (x), k = 1, 2, … l represents the reliability function of each failure mode of the electromechanical system.

[0063] Finally, according to the reliability optimization result, the operating parameters are adjusted, and the optimization result is verified based on the finite element method and sensor data.

[0064] The method of the present application first establishes a complex electromechanical system function block diagram, analyzes the failure modes of key components and their effects, then establishes a system reliability block diagram, analyzes the sensor data of each level of the system operating state, and according to the failure modes of key components, establishes a key component multi-physics performance simulation analysis model based on virtual prototype technology, constructs an electromechanical system sensor data and multi-physics performance simulation analysis fusion data set, then establishes a reliability evaluation model based on deep reinforcement learning, based on the adaptive MCMC importance sampling strategy, evaluates the reliability of the electromechanical system through reinforcement learning and environmental interaction, finally analyzes the reliability sensitivity of the operating parameters of the electromechanical system, optimizes the operating parameters based on the particle swarm intelligent algorithm, improves the system reliability, and verifies the effectiveness of the optimization strategy through environmental information. The method of the present application aims to improve the efficiency and accuracy of complex electromechanical system reliability evaluation, by combining the intelligent decision-making ability of deep reinforcement learning and the adaptive sampling strategy of importance sampling, reducing the required computing resources and time, while improving the accuracy and credibility of the evaluation results, to better support the design, maintenance and optimization decisions of the electromechanical system. BRIEF DESCRIPTION OF DRAWINGS

[0065] Figure 1 A flowchart of a method for electromechanical system reliability evaluation based on deep reinforcement learning and importance sampling according to the present application.

[0066] Figure 2 A structure diagram of an electromechanical system performance evaluation and prediction model based on deep learning in an embodiment of the present application.

[0067] Figure 3 A schematic diagram of a reliability evaluation framework based on deep reinforcement learning and importance sampling in an embodiment of the present application. DETAILED DESCRIPTION

[0068] The method of the present application will be further described below in conjunction with the drawings and embodiments.

[0069] As Figure 1 shown, the flowchart of a method for electromechanical system reliability evaluation based on deep reinforcement learning and importance sampling according to the present application, the specific steps are as follows:

[0070] S1, a complex electromechanical system function block diagram is established, failure modes of key components and their influences are analyzed, and then a system reliability block diagram is established;

[0071] First, the complex electromechanical system is decomposed according to its function and structure level, each subsystem, module and key component is identified, and a complex electromechanical system function block diagram is established; then, based on the FEMA (Failure Mode and Effects Analysis, FMEA) method, failure modes of key components of the complex electromechanical system and their influences are analyzed, and finally, on the basis of completing the electromechanical system function block diagram and the analysis of failure modes of key components and their influences, the function block diagram is converted into a reliability block diagram;

[0072] S2, analyze the sensor data of each level of the system, and based on the failure modes of the key components in step S1, establish a multi-physics field performance simulation analysis model of the key components based on virtual prototyping technology, and construct a fusion data set of sensor data and multi-physics field performance simulation analysis of the electromechanical system;

[0073] S3, based on the fusion data set of sensor data and multi-physics field performance simulation analysis of the electromechanical system established in step S2, a reliability evaluation model based on deep reinforcement learning is established;

[0074] S4, based on the reliability evaluation model established in step S3, an important sampling strategy based on adaptive MCMC is established, and the reliability of the electromechanical system is evaluated through reinforcement learning and environmental interaction;

[0075] S5, based on step S4, analyze the reliability sensitivity of the electromechanical system operating parameters, optimize the operating parameters based on the particle swarm intelligent algorithm, improve the system reliability, and further verify the effectiveness of the optimization strategy through environmental information;

[0076] The operating parameters include: motion trajectory, motor output torque, control parameter, joint output angle.

[0077] 2. The electromechanical system reliability evaluation method based on deep reinforcement learning and important sampling according to claim 1, wherein step S2 is as follows:

[0078] S21, analyze the sensor data of each level of the electromechanical system, and define the fault characteristic index;

[0079] The sensor data of the operation state of each level of the electromechanical system is multi-source heterogeneous information, including motor temperature and current, motor output torque, end effector motion image, vibration amplitude, acceleration, joint output angle and torque; first, the collected original data is denoised to improve the data quality; second, key features are extracted from the data after denoising, including frequency domain features, time domain features and time-frequency domain features;

[0080] Then, according to the operation state of each level of the system and the potential failure mode, a failure feature index space is defined, and is associated with each component in the reliability block diagram;

[0081] The failure feature index space includes motor temperature and current, motor output torque, end effector vibration amplitude, acceleration, joint output angle and torque.

[0082] S22, based on step S21, according to the main failure mode of the complex electromechanical system, a multi-physical field performance simulation analysis model of the key components of the complex electromechanical system is established based on virtual prototyping technology, and the performance response under the multi-physical field coupling of the key components of the system is analyzed.

[0083] S23, construct electromechanical system sensor data and multi-physical field performance simulation analysis fusion data set.

[0084] In this embodiment, the step S3 is specifically as follows:

[0085] S31, define the reliability evaluation problem;

[0086] According to the main failure mode of the complex electromechanical system, the reliability evaluation problem is defined as follows:

[0087] R k =P{g k (x1,x2,…,x n )≤0,k=1,2,…,l} (1)

[0088] Where R k represents the reliability of the kth performance response, P(·) represents the probability, g k (x1,x2,…,x n ) represents the kth performance limit state function, k=1,2,…,l. g k (x)≤0 indicates that the performance response meets the requirements, g k (x)>0 indicates performance failure, x=[x1,x2,…,x n ] represents the system input parameter vector.

[0089] S32, define the agent action and reinforcement learning environment in deep reinforcement learning;

[0090] In deep reinforcement learning, the action of the intelligent agent is defined as establishing a performance evaluation response prediction model for complex electromechanical systems based on deep learning, and the environment is the sensor data and performance response simulation data set of each level of the electromechanical system.

[0091] like Figure 2 As shown in the figure, a deep learning model is used to fuse multi-source heterogeneous sensor data. For the image information and time series feature data in the multi-source information, a deep learning model combining convolutional neural networks (CNN), long short-term memory (LSTM) networks and fully connected neural networks is constructed to establish a performance response evaluation and prediction model for complex electromechanical systems (a reliability evaluation model based on deep reinforcement learning).

[0092] The outputs of the CNN and LSTM are combined, leveraging both spatial and temporal features. Specifically, the RepeatVector layer is used to repeatedly extend the CNN output to the same number of time steps as the LSTM, and the expanded CNN output is used as the input to the LSTM. After the CNN and LSTM outputs are combined, a fully connected neural network (Dense layer) is added to further integrate the features and ultimately output the prediction results.

[0093] S33. Define the state-action pair value function, state value function and environment reward function, and evaluate the performance response evaluation and prediction model accuracy of complex electromechanical systems;

[0094] The state-action value function is defined based on the Bellman equation and is expressed as follows:

[0095] Q new (s t ,a t )=Q(s t ,a t )+α1[Aw(s t ,a t )+γminQ'(s t+1 ,a t+1 )-Q(s t ,a t )] (2)

[0096] Among them, s t represents the state at time t, a t represents the action at time t, Q(s t ,a t ) represents the value function of the state-action pair at time t, which is defined as shown in formula (3), Aw(s t ,a t ) represents the reward function, defined as formula (4), Qnew (s t ,a t ) represents the state-action pair new value function, γ represents the discount factor, Q'(s t+1 ,a t+1 ) represents the state-action pair value function at t+1 time, and α1 represents the learning rate.

[0097] Q(s t ,a t ) = ||g tru -g pre || 2 (3)

[0098] wherein g pre represents the deep learning model prediction result, and g tru represents the real value or simulation analysis result fed back by the environment.

[0099]

[0100] wherein Δx represents the change amount of the electromechanical system performance input parameter, represents the gradient of the state-action pair value function with respect to the input parameter.

[0101] The reinforcement learning and the policy gradient algorithm are fused to optimize the deep learning network parameters, and the network parameter update expression is as follows:

[0102]

[0103] wherein w represents the deep learning network weight, and α2 represents the learning rate.

[0104] In the embodiment, the step S4 is specifically as follows:

[0105] S41, according to the long-term return function, an important sampling strategy based on an adaptive Markov chain Monte Carlo algorithm is established;

[0106] The adaptive Markov chain Monte Carlo (Adaptive-MCMC) algorithm based on Bayesian optimization is a sampling method combining the Bayesian optimization idea and the MCMC sampling technology, and the core of the algorithm is to adaptively adjust the sampling strategy by using the historical sampling information to improve the efficiency of the sample and the exploration ability of the posterior distribution.

[0107] Exploration and learning are performed through the reinforcement learning, on the basis of which the cumulative reward is maximized, and the adaptive sampling strategy is optimized. G t represents the long-term return of the state-action pair, and the expression is as follows:

[0108]

[0109] where j denotes the jth action-state pair, G(t+1) denotes the t+1 action-state pair return, and γ denotes the discount factor.

[0110] Given the current sample x(t), the transition of the next state x(t+1) follows the transition probability T(x(t+1)|x(t)) and satisfies the detailed balance condition, which is expressed as follows:

[0111] P(x(t))T(x(t+1)|x(t))=P(x(t+1))T(x(t)|x(t+1)) (7)

[0112] The transition probability is calculated by equation (8), which is expressed as follows:

[0113] T(x(t+1)|x(t))=P(s new ,r|s,a)=P[S(t+1)=s new ,Aw(t+1)=r|S(t)=s t ,A(t)=a t ] (8)

[0114] where r denotes the reward for the state-action pair, s new denotes the new state at time t+1, and the state-action pair value function is defined according to the state-action pair value function, which is expressed as follows:

[0115]

[0116] where q π (s,a) denotes the state-action pair value function, E(·) denotes the expected function, and V π (·) denotes the state value function, which is expressed as follows:

[0117]

[0118] S42, based on the sampling strategy of step S41, the system reliability is evaluated through reinforcement learning and environment interaction;

[0119] Based on the reliability evaluation of the sampling method, the indicator function is first defined, which is expressed as follows:

[0120]

[0121] where, denotes the kth failure mode, m denotes the mth sampling point, and I k,m denotes the indicator function.

[0122] Then the reliability of each failure mode is evaluated, and the reliability Rk The expression is as follows:

[0123]

[0124] wherein N represents the total sample size.

[0125] In the embodiment, the step S5 is specifically as follows:

[0126] Based on the particle swarm optimization (PSO) intelligent algorithm, the operation parameters of the system are optimized to improve the system reliability, and the expression is as follows:

[0127]

[0128] wherein x * represents the optimal set of operation parameters, R sys (x) represents the system reliability, R k (x), k = 1, 2, … l represents the reliability function of each failure mode of the electromechanical system.

[0129] According to the reliability optimization result, the operation parameters are adjusted, and the optimization result is verified based on the finite element method and sensor data.

[0130] In order to verify the effectiveness of the method in the reliability evaluation and optimization of the complex electromechanical system, the embodiment is illustrated by an industrial robot. The system includes a mechanical arm, joint motor, reducer, controller, sensors (such as position sensor, force sensor) and power system and other key components, and the sensor data of the industrial robot at each level is collected, including: motor output torque, current, joint output angle and torque, end motion trajectory, vibration amplitude, acceleration, and part of the sensor data set is shown in Table 1. The failure mode of the industrial robot is defined as two failure modes that the repeatability error is greater than the allowable value and the end stability does not meet the requirements.

[0131] Table 1

[0132]

[0133] The uncertainty factors of the manufacturing error, load, material mechanical properties and dynamics model error of the industrial robot are characterized and quantified, a Simulink-based industrial robot dynamics simulation analysis model is established, the joint angle and the amplitude of the end effector are analyzed; at the same time, the elastic deformation of the mechanical arm is considered, and the elastic deformation of the mechanical arm is analyzed based on the finite element method. The simulation analysis data is corrected and confirmed according to the sensor data. On this basis, the sensor data and the simulation analysis data are fused to establish a reliability evaluation model based on deep reinforcement learning.

[0134] In summary, the method aims to improve the efficiency and accuracy of reliability evaluation of complex mechatronic systems, reduce the required computing resources and time for evaluation by combining the intelligent decision-making ability of deep reinforcement learning and the adaptability of important sampling, and improve the accuracy and reliability of the evaluation results to better support the design, maintenance and optimization decisions of mechatronic systems.

[0135] Those skilled in the art will appreciate that the embodiments described herein are presented for the purpose of aiding the reader in understanding the principles of the present application and should be understood as not limiting the scope of protection of the present application to such specific recitations and embodiments. Those skilled in the art can make various other specific modifications and combinations according to the technical inspiration disclosed in the present application without departing from the spirit of the present application, and these modifications and combinations are still within the scope of protection of the present application.

Claims

1. A method for reliability assessment of electromechanical systems based on deep reinforcement learning and importance sampling. The specific steps are as follows: S1. Establish a functional block diagram of a complex electromechanical system, analyze the failure modes and their impacts of key components, and then establish a system reliability block diagram; First, the complex electromechanical system is decomposed according to its functional and structural levels, and each subsystem, module and key components are identified to establish a functional block diagram of the complex electromechanical system. Then, the failure modes and their impacts of the key components of the complex electromechanical system are analyzed based on the FEMA method. Finally, based on the completion of the electromechanical system functional block diagram and the failure modes and their impacts analysis of the key components, the functional block diagram is converted into a reliability block diagram. S2. Analyze the operating status sensor data of each level of the system, and establish a multi-physics performance simulation analysis model for key components based on virtual prototyping technology based on the failure modes of key components in step S1, and construct a fusion data set of electromechanical system sensor data and multi-physics performance simulation analysis; S3. Establish a reliability assessment model based on deep reinforcement learning based on the electromechanical system sensor data and multi-physics field performance simulation analysis fusion data set established in step S2; S4. Based on the reliability assessment model established in step S3, an important sampling strategy based on adaptive MCMC is established to evaluate the reliability of the electromechanical system through reinforcement learning and environmental interaction; S5. Based on step S4, analyze the reliability sensitivity of the electromechanical system operating parameters, optimize the operating parameters based on the particle swarm intelligence algorithm, improve the system reliability, and further verify the effectiveness of the optimization strategy through environmental information; The operating parameters include: motion trajectory, motor output torque, control parameters, and joint output angle.

2. The electromechanical system reliability assessment method based on deep reinforcement learning and importance sampling according to claim 1, characterized in that: The step S2 is specifically as follows: S21. Analyze the operating status sensor data of each level of the electromechanical system and define fault characteristic indicators; The operating status sensor data at each level of the electromechanical system is multi-source heterogeneous information, including: motor temperature and current, motor output torque, end-effector motion image, vibration amplitude, acceleration, joint output angle and torque. First, the collected raw data is denoised to improve data quality. Second, key features are extracted from the denoised data, including: frequency domain features, time domain features, and time-frequency domain features. Then, based on the operating status and potential failure modes of each level of the system, the fault characteristic indicator space is defined and associated with each component in the reliability block diagram; The fault characteristic indicator space includes: motor temperature and current, motor output torque, end effector vibration amplitude, acceleration, joint output angle and torque; S22. Based on step S21, according to the main failure modes of the complex electromechanical system, a multi-physics performance simulation analysis model of key components of the complex electromechanical system is established based on virtual prototyping technology to analyze the performance response of the key components of the system under multi-physics coupling; S23. Construct a fusion dataset of electromechanical system sensor data and multi-physics field performance simulation analysis.

3. The electromechanical system reliability assessment method based on deep reinforcement learning and importance sampling according to claim 1, characterized in that: The step S3 is specifically as follows: S31. Define the reliability assessment problem; According to the main failure modes of complex electromechanical systems, the reliability assessment problem is defined as follows: R k =P{g k (x1,x2,…,x n )≤0,k=1,2,…,l} (1) Among them, R k represents the reliability of the kth performance response, P(·) represents the probability, g k (x1,x2, … ,x n ) represents the kth performance limit state function, k=1,2,…,l; g k (x)≤0 means the performance response meets the requirements, g k (x)>0 indicates performance failure, x=[x1,x2, … ,x n ] represents the system input parameter vector; S32. Define agent actions and reinforcement learning environments in deep reinforcement learning; In deep reinforcement learning, the agent's actions are defined as optimizing the parameters of the deep learning network. A deep learning-based prediction model for the performance evaluation and response of complex electromechanical systems is established. The environment is the sensor data and performance response simulation dataset of each layer of the electromechanical system. Using deep learning models to fuse multi-source heterogeneous sensor data, a deep learning model combining convolutional neural networks (CNNs), long short-term memory (LSTMs), and fully connected neural networks is constructed for image information and time series feature data in multi-source information. This model is used to establish performance response evaluation and prediction models for complex electromechanical systems. S33. Define the state-action pair value function, state value function and environment reward function, and evaluate the performance response evaluation and prediction model accuracy of complex electromechanical systems; The state-action value function is defined based on the Bellman equation and is expressed as follows: Q new (s t ,a t )=Q(s t ,a t )+α1[Aw(s t ,a t )+γminQ'(s t+1 ,a t+1 )-Q(s t ,a t )] (2) Among them, s t represents the state at time t, a t represents the action at time t, Q(s t ,a t ) represents the value function of the state-action pair at time t, which is defined as shown in formula (3), Aw(s t ,a t ) represents the reward function, defined as formula (4), Q new (s t ,a t ) represents the new value function of the state-action pair, γ represents the discount factor, Q'(s t+1 ,a t+1 ) represents the state-action pair value function at time t+1, and α1 represents the learning rate; Q(s t ,a t )=||g tru -g pre || 2 (3) Among them, g pre Represents the prediction result of the deep learning model, g tru Represents the actual value of environmental feedback or simulation analysis results; Where Δx represents the change in the input parameters of the electromechanical system performance, Represents the gradient of the state-action value function with respect to the input parameters; Fusion reinforcement learning and policy gradient algorithm optimize deep learning network parameters, the evaluation network parameter update expression is as follows: Among them, w represents the deep learning network weight and α2 represents the learning rate.

4. The electromechanical system reliability assessment method based on deep reinforcement learning and importance sampling according to claim 1, characterized in that: The step S4 is specifically as follows: S41. According to the long-term reward function, an important sampling strategy based on the adaptive Markov chain Monte Carlo algorithm is established; Explore and learn through reinforcement learning, maximize the cumulative reward on this basis, and optimize the adaptive sampling strategy; set G t represents the long-term reward of a state-action pair, which is expressed as follows: Where j represents the jth action-state pair, G(t+1) represents the reward of the t+1 action-state pair, and γ represents the discount factor; Given the current sample x(t), the transition to the next state x(t+1) follows the transition probability T(x(t+1)|x(t)) and satisfies the detailed balance condition, which is expressed as follows: P(x(t))T(x(t+1)|x(t))=P(x(t+1))T(x(t)|x(t+1)) (7) The transition probability is calculated by formula (8), which is as follows: T(x(t+1)|x(t))=P(s new ,r|s,a)=P[S(t+1)=s new ,Aw(t+1)=r|S(t)=s t ,A(t)=a t ](8) Among them, r represents the reward for the state-action pair, s new Represents the new state at time t+1. According to the definition of the state-action pair value function, the state-action pair value function expression is as follows: Among them, q π (s,a) represents the state-action pair value function, E(·) represents the expectation function, V π (·) represents the state value function, which is expressed as follows: S42. Based on the sampling strategy of step S41, evaluate the system reliability through reinforcement learning and environment interaction; For reliability assessment based on sampling method, we first define the indicator function, which is expressed as follows: in, represents the kth failure mode, m represents the mth sampling point, I k,m represents the indicator function; Then evaluate the reliability of each failure mode, the reliability of the kth performance R k The expression is as follows: Where N represents the total sample size.

5. The electromechanical system reliability assessment method based on deep reinforcement learning and importance sampling according to claim 1, characterized in that: The step S5 is specifically as follows: Based on the particle swarm PSO intelligent algorithm, the operating parameters of the system are optimized, and the expression is as follows: Among them, x * represents the optimal operating parameter set, R sys (x) represents the system reliability, R k (x), k = 1, 2, ... l represents the reliability function of each failure mode of the electromechanical system; Finally, according to the reliability optimization results, the operating parameters are adjusted, and the optimization results are verified based on the finite element method and sensor data.

Citation Information

Patent Citations

  • Ultra-dense network resource allocation method based on improved dual deep Q learning

    CN116634565A

  • A computer-implemented method and an apparatus for reinforcement learning

    WO2023225941A1