Robot polishing online flutter suppression method based on reinforcement learning
By adopting reinforcement learning-based methods during the robot polishing process, a time-enhanced strategy model is constructed, and impedance parameters are output to regulate the robot end effector, which solves the problems of insufficient adaptability and insufficient control accuracy in the existing technology, real-time suppression and efficient control of complex flutters are achieved.
Patent Information
- Application Number
- CN202510502735.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The prior art has problems of insufficient adaptability and insufficient control accuracy during the polishing process of robots, and cannot effectively suppress complex flutter phenomena.
A reinforcement learning-based method is adopted to collect vibration signals and force error signals, and multi-source feature vectors are constructed, and a convolutional neural network is used to fusion and dimensionality reduction with the multi-head attention module. Then, a time-enhanced strategy model is constructed through a deep reinforcement learning framework embedded in LSTM, and the impedance parameters are output to regulate the robot end effector to achieve flutter suppression.
It significantly improves the system's adaptability and control accuracy, can suppress complex flutter mechanisms in real time, and improves processing efficiency and finished product quality.
Smart Images

Figure CN120038761A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of robot vibration control and adjustment, and particularly to an online chatter suppression method for robot grinding based on reinforcement learning. Background Art
[0002] In related technologies, due to excellent characteristics such as light weight and high specific strength, thin-walled structures are widely used in fields such as aerospace, automotive manufacturing, electronic communication, and ship engineering. However, thin-walled structures often have characteristics such as large volume, complex structure, uneven wall thickness, and weak rigidity. These characteristics make the structure prone to vibration and deformation during the machining process, which in turn leads to a decline in surface quality and large machining errors, directly affecting the final performance.
[0003] Currently, the robot grinding technology based on force control has become an important part of advanced manufacturing due to its ability to adapt to complex surfaces, irregular shapes, and high-quality machining, such as using robots to grind thin-walled structures. However, during the robot grinding process, due to the dynamic contact between the tool and the workpiece and the fluctuations of machining parameters, chatter phenomena often occur. Chatter is an unstable self-excited vibration, which may lead to a decline in surface quality, increased tool wear, and increased energy consumption, thus seriously affecting the machining efficiency and the quality of finished products. Therefore, the control strategy of the robot needs to monitor and suppress vibration during the machining process.
[0004] In the prior art, the methods for achieving vibration suppression during the robot grinding and polishing process are relatively single and have various defects. For example, Chinese Patent (Patent No. CN20201155164.5) proposes a contact vibration suppression method for robot grinding and polishing. This prior art solution mainly collects the normal contact force of the workpiece in real time through a six-axis force sensor, decomposes and calculates the permutation entropy of the information, and then adjusts the robot according to the threshold of the index to achieve vibration suppression. However, there are complex changes in the working conditions during the robot grinding process. The prior art solution relies on the setting of thresholds, cannot adapt to the variable working conditions, lacks self-adaptability, and has insufficient control accuracy. Summary of the Invention
[0005] This application provides an online chatter suppression method for robot grinding based on reinforcement learning. Starting from the vibration signal and force error signal during the robot grinding process, a model is constructed to output a suppression strategy to control the end effector of the robot, realizing the real-time suppression and efficient control of the complex chatter mechanism during the process of the robot grinding thin-walled parts, significantly improving the self-adaptability and control accuracy of the system, and solving the problems of insufficient self-adaptability and insufficient control accuracy existing in the prior art.
[0006] In a first aspect, this application provides an online chatter suppression method for robot grinding based on reinforcement learning, including: Collecting discrete vibration signals during the robot grinding process, and determining a force error signal during the grinding process, wherein the force error signal is used to characterize a force control deviation during the grinding process; The discrete vibration signal and the force error signal are used as multi-source signals for preprocessing and normalized feature extraction fusion to obtain a normalized feature vector with high recognition; Convolutional neural network and multi-head attention module are used to perform feature fusion and dimensionality reduction on the normalized feature vector to obtain low-dimensional feature representation; Temporal encoding is performed on the low-dimensional feature representation to obtain a temporal state, and a temporal enhanced policy model is constructed through a deep reinforcement learning framework embedded with LSTM, wherein the policy model includes an Actor network and a Critic network of a deep deterministic policy gradient algorithm; Using the timing state as an input of an Actor network to obtain an output impedance parameter, and using the timing state and the impedance parameter as inputs of a Critic network to obtain an output quality evaluation value; Based on the quality evaluation value, the robot control parameters are updated by using a minimum loss function combined with entropy regularization optimization, and the robot control parameters are used as a chatter suppression strategy to regulate the robot end effector; Wherein, the robot control parameters include updated impedance parameters.
[0007] Optionally, collecting discrete vibration signals during the robot grinding process, and determining a force error signal during the grinding process, include: The discrete vibration signal collected during the robot grinding process is defined as , , and define the force error signal during robot grinding as ; in, Indicates time, It is the original discrete vibration signal, which has rich time-frequency information and nonlinear characteristics. is the force expected during robot grinding, is the actual measured force.
[0008] Optionally, the discrete vibration signal and the force error signal are used as multi-source signals for preprocessing and normalized feature extraction fusion to obtain a normalized feature vector with high recognition, including: Based on the selected wavelet basis function, the discrete vibration signal is subjected to multi-layer discrete wavelet decomposition and energy entropy quantification to obtain vibration signal features characterizing the flutter evolution trend, and original information of the force error signal is retained, and features of the force error signal are statistically analyzed by variance and / or mean difference to obtain force error signal features; The vibration signal feature and the force error signal feature are concatenated, and scale differences are eliminated by normalization to obtain a normalized feature vector.
[0009] Optionally, performing multi-layer discrete wavelet decomposition and energy entropy quantification on the discrete vibration signal based on a selected wavelet basis function to obtain vibration signal features characterizing the flutter evolution trend, and retaining original information of the force error signal, and obtaining force error signal features by statistically analyzing the features of the force error signal through variance and / or mean difference, including: according to , combined with the preset number of decomposition layers, the selected wavelet basis function is used to decompose the original vibration signal Perform discrete wavelet decomposition layer by layer to obtain low-frequency approximate coefficients and high-frequency detail coefficients after decomposition of each layer in wavelet transform; Combine the low-frequency approximate coefficients and high-frequency detail coefficients of each layer to get the final decomposition ; For the final decomposition ,according to ,right Each subband Calculating Energy ; Based on the energy calculated for each subband ,according to , calculate the energy proportion of each sub-band, and according to Define wavelet energy entropy , wavelet energy entropy is used to quantify the complexity and nonlinear characteristics of vibration signals, as well as important features that characterize the flutter evolution trend; Determine the vibration signal characteristics based on the wavelet energy entropy and the energy calculated for each sub-band; Retention force error signal Original information, and based on Calculate statistical features to obtain force error signal features; in, is the decomposition layer, The low-frequency coefficients obtained by layer decomposition are , the high frequency coefficient is , As the final decomposition in the wavelet decomposition, it means In the subband wavelet coefficients, is the total number of subbands, is the sampling point.
[0010] Optionally, the vibration signal feature and the force error signal feature are feature concatenated, and scale differences are eliminated by normalization to obtain a normalized feature vector, including: The vibration signal features extracted from the discrete vibration signal and , and the force error signal features extracted from the force error signal and are subjected to feature concatenation, and normalization processing for eliminating scale differences is performed according to to obtain a normalized feature vector; Among them, is the normalized feature vector.
[0011] Optionally, a convolutional neural network and a multi-head attention module are used to perform feature fusion and dimensionality reduction on the normalized feature vector to obtain a low-dimensional feature representation, including: The normalized feature vector is input into a convolutional layer, and an output feature map is obtained after a series of convolutional and pooling operations ; Based on the feature map output by the convolutional layer, the multi-head attention module is used to perform weighted fusion on the normalized feature vector, and a linear transformation is performed on the feature map according to to obtain a query , a key and a value ; With the query , the key and the value as inputs, according to , the attention output of each head is calculated through scaled dot-product attention; According to , the attention outputs of each head are concatenated, and the final multi-head attention output is obtained through a linear transformation; According to , the multi-head attention output is further reduced in dimension through a fully connected layer or a pooling layer to obtain a low-dimensional feature representation; Among them, are all learnable weight matrices, is the dimension of the key vector, is the number of heads, represents the attention output of the corresponding head.
[0012] Optionally, temporal encoding is performed on the low-dimensional feature representation to obtain a temporal state, and a temporal enhanced policy model is constructed through a deep reinforcement learning framework embedded with LSTM, including: An LSTM network is embedded in the reinforcement learning framework to perform temporal encoding on the preprocessed feature sequence, and according to , a temporal state containing historical dynamic information is generated; Construct an Actor-Critic framework based on the Deep Deterministic Policy Gradient algorithm to obtain a temporally enhanced policy model.
[0013] Optionally, taking the temporal state as the input of the Actor network to obtain the output impedance parameters, and using the temporal state and the impedance parameters as the input of the Critic network to obtain the output quality evaluation value, including: Taking the temporal state as the input within the Actor network, and according to processing the aging state through a fully connected layer and an LSTM layer to obtain the output action to be used as the regulation of the impedance parameters at the end of the robot; Taking the temporal state and the output action as the input of the Critic network, and according to outputting the corresponding value as the quality evaluation value for evaluating the quality of the current flutter suppression strategy; Among them, is the stiffness in the impedance parameters, is the damping in the impedance parameters, represents the parameters of the Actor network, are the parameters of the Critic network.
[0014] Optionally, based on the quality evaluation value, by minimizing the loss function and combining entropy regularization optimization, update the robot control parameters, and use the robot control parameters as the flutter suppression strategy to regulate the end effector of the robot, including: During the training of the policy model, introducing an entropy regularization term into the loss function of the Critic network; Adding entropy regularization optimization to the loss function of the Critic network to dynamically adjust the exploration weight and obtain the maximized value output by the Critic network; According to the obtained maximized value, using to update the Actor network, and through the updated Actor network and Critic network, obtain the updated robot control parameters; Among them, is the discount factor, is the target Critic network, is the output of the target Actor network. For the processing of temporal data, random exploration efficiency, and training convergence speed during the reinforcement learning process, through the weight Perform weighted sampling to enhance the model.
[0015] In summary, in the embodiments of the present application, starting from the vibration signal and force error signal during the robot grinding process, through discrete wavelet transform, energy entropy calculation, and statistical feature extraction, a multi-source feature vector with high recognition is constructed. After normalization, a convolutional neural network and a multi-head attention module are used for feature fusion and dimensionality reduction to obtain a low-dimensional feature representation. Then, a time-series enhanced policy model is constructed through the DDPG framework embedded with LSTM. In the policy model, the Actor network outputs impedance parameters based on the time-series state, and the Critic network realizes the stable update of the policy by minimizing the loss function and combining entropy regularization, and then controls the robot end effector. This embodiment realizes the real-time suppression and efficient control of the complex chatter mechanism during the robot grinding of thin-walled parts, significantly improves the self-adaptability and control accuracy of the system, and solves the problems of insufficient self-adaptability and insufficient control accuracy existing in the prior art. Description of the Drawings
[0016] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of a method for online chatter suppression of robot grinding based on reinforcement learning provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of the steps of a method for online chatter suppression of robot grinding based on reinforcement learning provided by an optional embodiment of the present application; Figure 3 It is a flowchart of online chatter suppression of robot grinding based on reinforcement learning provided by an optional example of the present application. Detailed Embodiments
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0020] For the convenience of understanding the embodiments of the present application, the following will further explain with reference to the accompanying drawings and specific embodiments. The embodiments do not limit the embodiments of the present application.
[0021] Figure 1 The following is a schematic flow chart of an online chatter suppression method for robot grinding based on reinforcement learning provided by the embodiments of the present application. As Figure 1 shown, the online chatter suppression method for robot grinding based on reinforcement learning provided by the embodiments of the present application may specifically include the following steps: Step 110, collect discrete vibration signals during the robot grinding process, and determine the force error signal during the grinding process.
[0022] Among them, the force error signal is used to characterize the force control deviation during the grinding process.
[0023] In this embodiment, discrete vibration signals and force error signals are mainly collected during the robot grinding process. The discrete vibration signals, also known as the original discrete vibration signals, have rich time-frequency information and nonlinear characteristics. The force error signal is mainly determined based on the force collected during the robot grinding process. The force error signal directly reflects the force control deviation during the grinding process and retains the real-time working condition information.
[0024] Step 120, perform preprocessing and normalization feature extraction and fusion on the discrete vibration signal and the force error signal as multi-source signals to obtain a normalized feature vector with high recognition.
[0025] In the related art, since the discrete vibration signal and the force error signal belong to multi-source signals, there is a problem of heterogeneity. The high-dimensional features extracted from the two signals will also have the problem of feature redundancy.
[0026] To solve the heterogeneity problem of multi-source signals and the problem of high-dimensional feature redundancy, in this embodiment, feature extraction and processing are respectively performed on the discrete vibration signal and the force error signal, and then feature fusion is performed. Specifically, in this embodiment, preprocessing is performed on the discrete vibration signal, including discrete wavelet transform and energy entropy calculation, to extract the features of the discrete vibration signal to quantify the complexity and nonlinear features of the discrete vibration signal, and statistical features are calculated for the force error signal. Then, the features extracted from the discrete vibration signal and the statistical features of the force error signal are fused by feature concatenation to obtain a normalized feature vector, providing a unified and highly recognizable input for subsequent feature compression and reinforcement models.
[0027] Step 130, use a convolutional neural network and a multi-head attention module to perform feature fusion and dimensionality reduction on the normalized feature vector to obtain a low-dimensional feature representation.
[0028] In this embodiment, to further reduce data redundancy and improve the ability to capture key features, a convolutional neural network (CNN) and a multi-head attention module are used to process the normalized feature vectors, including analyzing the feature maps of the normalized feature vectors, performing weighted fusion on the feature maps, and finally performing dimensionality reduction and compression on the normalized feature vectors to obtain a low-dimensional feature representation. This low-dimensional feature serves as an efficient input for the subsequent reinforcement learning model while retaining the features / characteristics of both signals.
[0029] Step 140: Perform temporal encoding on the low-dimensional feature representation to obtain a temporal state, and construct a temporally enhanced policy model through a deep reinforcement learning framework embedded with LSTM.
[0030] Among them, the policy model includes an Actor network and a Critic network of the deep deterministic policy gradient algorithm.
[0031] Step 150: Use the temporal state as the input of the Actor network to obtain the output impedance parameters, and use the temporal state and the impedance parameters as the input of the Critic network to obtain the output quality evaluation value.
[0032] Step 160: Based on the quality evaluation value, update the robot control parameters through a minimum loss function combined with entropy regularization optimization, and use the robot control parameters as a flutter suppression strategy to regulate the robot end effector.
[0033] Among them, the robot control parameters include the updated impedance parameters.
[0034] A unified description of Steps 140 - 160: In specific implementation, considering the strong temporal dependence of the flutter system, in this embodiment, an LSTM (Long Short-Term Memory) network is embedded in a deep reinforcement learning framework (such as the deep deterministic policy gradient DDPG) to construct a temporally enhanced policy model. This policy model includes an Actor-Critic framework based on the DDPG algorithm (composed of an Actor network and a Critic network). By performing temporal encoding on the low-dimensional feature representation to obtain a temporal state as the model input, after being processed by the Actor network, the temporal state outputs an action (represented as impedance parameters). Then, the current policy quality is evaluated by inputting it into the Critic network, and optimization exploration is carried out using the designed loss function to update the Actor network and the Critic network. Furthermore, the updated robot control parameters output by the policy model are obtained. The robot control parameters, as a flutter suppression strategy, mainly include the updated impedance parameters. The robot end effector is regulated using the updated impedance parameters.
[0035] It can be seen that the embodiments of the present application start from the vibration signal and force error signal in the robot grinding process. Through discrete wavelet transform, energy entropy calculation, and statistical feature extraction, a multi-source feature vector with high recognition is constructed. After normalization, a convolutional neural network and a multi-head attention module are used for feature fusion and dimensionality reduction to obtain a low-dimensional feature representation. Then, a time-series enhanced policy model is constructed through the DDPG framework embedded with LSTM. In the policy model, the Actor network outputs impedance parameters based on the time-series state, and the Critic network realizes the stable update of the policy by minimizing the loss function and combining entropy regularization, and then regulates the end effector of the robot. It can be seen that this embodiment realizes the real-time suppression and efficient control of the complex chatter mechanism in the process of the robot grinding thin-walled parts, significantly improves the self-adaptability and control accuracy of the system, and while solving the problems of insufficient self-adaptability and insufficient control accuracy existing in the prior art, it can also solve key problems such as multi-source data heterogeneity, time-series correlation, and insufficient adaptability of traditional methods.
[0036] Referring to Figure 2 , a schematic flow chart of steps of a method for online chatter suppression of a robot grinding based on reinforcement learning provided by an optional embodiment of the present application is shown. The method may specifically include the following steps: Step 210, collect discrete vibration signals during the robot grinding process, and determine the force error signal during the grinding process, where the force error signal is used to characterize the force control deviation during the grinding process.
[0037] In a specific implementation, referring to Figure 3 shown, to achieve the quantitative acquisition of discrete vibration signals and force error vibration signals during the robot grinding process, this embodiment defines the input signal through a formula. Specifically, the discrete vibration signal is mainly determined based on time and the original vibration signal, and the force error signal is mainly determined based on the desired force and the actually measured force during the robot grinding process.
[0038] In an optional embodiment, collecting discrete vibration signals during the robot grinding process and determining the force error signal during the grinding process may specifically include: defining the discrete vibration signal collected during the robot grinding process as , , and defining the force error signal during the robot grinding process as ; where represents time, is the original discrete vibration signal, which has rich time-frequency information and non-linear characteristics, is the desired force during the robot grinding process, is the actually measured force.
[0039] In this embodiment, by defining discrete vibration signals and force error signals, the acquisition of the two signals is realized, enabling the discrete vibration signals to have rich time-frequency information and non-linear characteristics; the force error signals can also retain real-time working condition information, and through the force error signals, the force control deviation during the grinding process can be intuitively determined in this embodiment to optimize the subsequent grinding control of the robot.
[0040] Step 220: Perform multi-level discrete wavelet decomposition and energy entropy quantization on the discrete vibration signals based on the selected wavelet basis function to obtain vibration signal features characterizing the evolution trend of chatter, and retain the original information of the force error signals. Statistically analyze the features of the force error signals through variance and / or mean difference to obtain force error signal features.
[0041] In specific implementation, the vibration signal features, as the features extracted from the discrete vibration signals, mainly include but are not limited to: sub-band energies corresponding to each wavelet coefficient, wavelet energy entropy, and other time-frequency statistics, etc.; the force error signal features, as the features extracted from the force error signals, mainly include but are not limited to: variance features and mean difference features.
[0042] Refer to Figure 3 As shown, for the feature extraction of discrete vibration signals: In this embodiment, the Daubechies wavelet basis function is selected to perform multi-level discrete wavelet decomposition on the discrete vibration signals, and the signals are represented as two types of coefficients, namely low-frequency approximation coefficients and high-frequency detail coefficients. Then, the two types of coefficients are represented in sub-bands, the energy of each sub-band is calculated, and then the wavelet energy entropy is determined. The wavelet energy entropy is used to quantify the important features characterizing the evolution trend of chatter of the vibration signals, that is, the discrete vibration signal features.
[0043] Refer to Figure 3 As shown, for the feature extraction of the force error signals, in this embodiment, the original information of the force error signals is used to calculate statistical features, such as through the mean value and / or variance, etc., so as to extract the signal features of the force error signals. Subsequently, the features corresponding to the two signals are further feature-cascaded, and the differences in scales of different signal features are eliminated through normalization processing to obtain normalized vector features.
[0044] Optionally, the above-mentioned performing multi-level discrete wavelet decomposition and energy entropy quantization on the discrete vibration signals based on the selected wavelet basis function to obtain vibration signal features characterizing the evolution trend of chatter, and retaining the original information of the force error signals, and statistically analyzing the features of the force error signals through variance and / or mean difference to obtain force error signal features may include: According to , combined with the preset decomposition level, use the selected wavelet basis function for the original vibration signals Perform layer-by-layer discrete wavelet decomposition to obtain the low-frequency approximation coefficients and high-frequency detail coefficients after each layer of decomposition in the wavelet transform; combine the low-frequency approximation coefficients and high-frequency detail coefficients of each layer of decomposition to obtain the final decomposition ; For the final decomposition , according to , for each sub-band in calculate the energy ; Based on the energy calculated for each sub-band , according to , calculate the proportion of the energy of each sub-band, and according to define the wavelet energy entropy . The wavelet energy entropy is used to quantify the complexity and non-linear characteristics of the vibration signal, and is an important feature characterizing the flutter evolution trend; determine the vibration signal characteristics according to the wavelet energy entropy and the energy calculated for each sub-band; retain the original information of the force error signal , and calculate the statistical characteristics according to to obtain the force error signal characteristics; where is the decomposition layer, and the low-frequency coefficient obtained from the decomposition of the th layer is , and the high-frequency coefficient is , is used as the final decomposition in the wavelet decomposition, representing the th wavelet coefficient in the th sub-band, is the total number of sub-bands, is the sampling point.
[0045] In a specific implementation, to ensure that the features extracted from the discrete vibration signal can quantify the complexity and non-linear characteristics of the vibration signal, this embodiment performs multi-layer wavelet decomposition on the wavelet transform of the discrete vibration signal to obtain the low-frequency and high-frequency coefficients of each layer as intermediate quantities of the wavelet transform, represents the signal as low-frequency and high-frequency coefficients, and finally obtains the final decomposition after multi-layer decomposition . Then, starting from the calculation of energy and energy entropy, important feature extraction is realized by using the above formula.
[0046] In addition, to ensure that the features extracted from the force error signal can retain the original information, this embodiment mainly uses the formula to calculate the statistical characteristics and realizes feature extraction by combining the mean and variance methods.
[0047] Step 230, perform feature concatenation on the vibration signal characteristics and the force error signal characteristics, and eliminate the scale difference through normalization to obtain a normalized feature vector.
[0048] In a specific implementation, referring to Figure 3As shown, in this embodiment, the features extracted from the discrete vibration signal (such as the energy corresponding to each sub-band, wavelet energy entropy, and other time-frequency statistics) are cascaded with the statistical features of the force error signal (such as variance, mean difference), etc., and through normalization processing to eliminate the scale difference, a normalized feature vector is obtained. This normalized feature vector provides a unified and highly recognizable input for subsequent feature compression and reinforcement learning models.
[0049] Optionally, cascading the vibration signal features with the force error signal features and eliminating the scale difference through normalization to obtain a normalized feature vector may include: the vibration signal features extracted from the discrete vibration signal and , are cascaded with the force error signal features extracted from the force error signal and , and normalization processing for eliminating the scale difference is performed according to to obtain a normalized feature vector; where is the normalized feature vector.
[0050] Specifically, this embodiment uses the formula to achieve the cascading of different features, eliminate the differences in scale between different features, and solve the problems of heterogeneity and high-dimensional feature redundancy of multi-source signals.
[0051] Step 240, using a convolutional neural network and a multi-head attention module to perform feature fusion and dimensionality reduction on the normalized feature vector to obtain a low-dimensional feature representation.
[0052] In a specific implementation, this embodiment uses the constructed convolutional neural network, combined with the multi-head attention module, takes the normalized feature vector as the input, performs feature fusion and dimensionality reduction on it, thereby further reducing data redundancy and improving the ability to capture key features, and obtains a low-dimensional feature representation.
[0053] In an alternative embodiment, using a convolutional neural network and a multi-head attention module to perform feature fusion and dimensionality reduction on the normalized feature vector to obtain a low-dimensional feature representation may specifically include: inputting the normalized feature vector into the convolutional layer, and obtaining an output feature map after a series of convolutional and pooling operations ; taking the feature map output by the convolutional layer as a reference, using the multi-head attention module to perform weighted fusion on the normalized feature vector, and performing a linear transformation on the feature map according to to obtain a query , a key and a value ; using the query , the key and the value as inputs, according to , calculate the attention output of each head through scaled dot - product attention; according to , splice the attention outputs of each head, and obtain the final multi - head attention output through a linear transformation; according to , further reduce the dimension of the multi - head attention output through a fully - connected layer or a pooling layer to obtain a low - dimensional feature representation; where are all learnable weight matrices, is the dimension of the key vector, is the number of heads, represents the attention output of the corresponding head.
[0054] Exemplarily, referring to Figure 3 , using the normalized feature vector as the input, perform convolution and pooling operations within the convolutional layer of the convolutional neural network. To improve the key feature capture ability, mainly use the formula to achieve the extraction of the feature map.
[0055] Then use the feature map as the input of the multi - head attention module, and perform weighted fusion within the multi - head attention module. The weighted fusion process includes but is not limited to: linear transformation. When performing linear transformation, successively use the formulas: , and to achieve the splicing of the attention output. Finally, use the formula to achieve feature dimension reduction. Using this formula for feature dimension reduction can retain the time - frequency energy distribution characteristics of the vibration signal and the statistical information of the force signal, and provide an efficient input for the subsequent reinforcement learning model.
[0056] Step 250, perform temporal encoding on the low - dimensional feature representation to obtain a temporal state, and construct a temporally enhanced policy model through a deep reinforcement learning framework embedded with LSTM.
[0057] Among them, the policy model includes an Actor network and a Critic network of the deep deterministic policy gradient algorithm.
[0058] In specific implementation, considering that the flutter system has strong temporal dependence, in this embodiment, an LSTM network is embedded in the reinforcement learning framework to perform temporal encoding on the pre - processed feature sequence to generate a state representation containing historical dynamic information. Then, on this basis, construct an Actor - Critic framework based on the deep deterministic policy gradient (DDPG) algorithm.
[0059] Optionally, the above-mentioned temporal encoding is performed on the low-dimensional feature representation to obtain a temporal state, and a temporal-enhanced policy model is constructed through a deep reinforcement learning framework embedded with LSTM, which may specifically include: embedding an LSTM network in the reinforcement learning framework, performing temporal encoding on the preprocessed feature sequence, and according to , generating a temporal state containing historical dynamic information ; constructing an Actor-Critic framework based on the deep deterministic policy gradient algorithm to obtain a temporal-enhanced policy model.
[0060] Thus, in this embodiment, an LSTM layer is introduced into the DDPG reinforcement learning framework to perform temporal encoding on the fused low-dimensional feature sequence, generate a state containing historical dynamic information, and improve the temporal perception ability of the policy model, so that the policy decision can not only reflect the current state but also fuse historical information, thereby overcoming the shortsightedness problem of traditional methods.
[0061] In a specific implementation, the temporal state obtained according to the formula not only reflects the current flutter state but also fuses historical information, thereby improving the temporal perception ability of the model.
[0062] Step 260, using the temporal state as the input of the Actor network to obtain the output impedance parameter, and using the temporal state and the impedance parameter as the input of the Critic network to obtain the output quality evaluation value.
[0063] In an optional embodiment, using the temporal state as the input of the Actor network to obtain the output impedance parameter, and using the temporal state and the impedance parameter as the input of the Critic network to obtain the output quality evaluation value may include: using the temporal state as the input in the Actor network, and performing fully connected layer and LSTM layer processing on the aging state according to to obtain the output action to be used as the regulation of the impedance parameter of the robot end effector; using the temporal state and the output action as the input of the Critic network, and outputting the corresponding value according to as the quality evaluation value for evaluating the quality of the current flutter suppression strategy; where is the stiffness in the impedance parameter, is the damping in the impedance parameter, represents the parameters of the Actor network, is the parameter of the Critic network.
[0064] In this embodiment, the regulation of the impedance parameters at the end of the robot mainly includes stiffness and damping , and the timing state is used as the input of the Actor network. After passing through several fully connected layers and LSTM layers, actions are output This is used as the input of the Critic network, and the formula is used to evaluate the quality of the current policy.
[0065] Step 270: Based on the quality evaluation value, update the robot control parameters through the minimum loss function combined with entropy regularization optimization, and use the robot control parameters as a flutter suppression strategy to regulate the end effector of the robot.
[0066] Among them, the robot control parameters include updated impedance parameters.
[0067] In an alternative embodiment, based on the quality evaluation value, update the robot control parameters through the minimum loss function combined with entropy regularization optimization, and use the robot control parameters as a flutter suppression strategy to regulate the end effector of the robot. Specifically, it may include: when training the policy model, introduce an entropy regularization term into the loss function of the Critic network ; Add entropy regularization optimization to the loss function of the Critic network , dynamically adjust the exploration weight, and obtain the value that maximizes the output of the Critic network ; According to the obtained maximized value, use to update the Actor network, and through the updated Actor network and Critic network, obtain updated robot control parameters; where is the discount factor, is the target Critic network, is the output of the target Actor network. For the processing of timing data, random exploration efficiency, and training convergence speed in the reinforcement learning process, weighted sampling is performed through the weight to enhance the model.
[0068] In this embodiment, with the entropy regularization term as the reference parameter, introduce entropy regularization optimization into the loss function of the Critic network , so as to maximize the value output by the Critic network.
[0069] In the related art, during the control process of existing robots for grinding thin-walled parts, they rely on passive mechanical structures to absorb the energy surge generated by chatter, and the frequency band for suppressing vibration is limited, resulting in poor applicability. In addition, most of the existing robot grinding process control methods rely on static threshold setting. When the chatter state reaches a certain fixed threshold, the processing parameters are adjusted, lacking the dynamic response ability to the chatter phenomenon.
[0070] To address the above technical problems, in this embodiment, a dynamic control strategy based on real-time vibration and force signal analysis is innovatively introduced, and a dynamic impedance parameter optimization framework based on the deep deterministic policy gradient algorithm is mainly proposed. Specifically, this framework realizes real-time chatter suppression and efficient and stable control of the grinding process through multi-source signal fusion, time-series enhancement strategy modeling, stabilization training mechanism, and redundant information filtering.
[0071] In specific implementation, a chatter suppression system can be constructed based on the online chatter suppression method of the robot grinding provided in this embodiment. To make the system have stronger adaptability, in this embodiment, historical data is used for training to automatically adjust and optimize its strategy. As a result, the system can adapt to more diverse and complex working conditions, improving the adaptability and control accuracy of the robot in the complex grinding process.
[0072] Specifically, referring to Figure 3 , during the training process, to reduce the possibility of the algorithm falling into local optimum, in this embodiment, an entropy regularization term is introduced into the loss function of the Critic network. According to , the exploration weight is dynamically adjusted. When the policy entropy is lower than the threshold, the exploration probability is increased to avoid local optimum. And the entropy regularization term is added to the loss function of the Critic network, so that the value output by the Critic network can be maximized, and the Actor network is updated by maximizing the value.
[0073] Furthermore, referring to Figure 3 , this embodiment also introduces an experience replay mechanism based on the priority of time-series error to address the challenges such as overfitting, low random exploration efficiency, and slow training convergence speed that are prone to occur when reinforcement learning processes time-series data. Specifically, this mechanism measures the time-series error of experience samples, assigns priority weights to the experience data, and performs weighted sampling in the experience replay pool according to this weight. For example, weighted sampling is realized using the formula , thereby optimizing the data utilization efficiency, enhancing the model's attention to key experiences, and improving the stability and convergence speed of training. In addition, this method can also alleviate the overfitting problem to a certain extent. It can prevent the model from relying too much on recent experiences, but instead prompts it to learn from a wider range of historical experiences, improving the generalization ability of the strategy.
[0074] In summary, the embodiments of the present application start from the vibration signal and the force error signal to construct a highly recognizable multi-source feature vector through discrete wavelet transform, energy entropy calculation, and statistical feature extraction . After normalization, a convolutional neural network and a multi-head attention module are used for feature fusion and dimensionality reduction to obtain a low-dimensional feature representation . Then, a time-series enhanced policy model is constructed through the DDPG framework embedded with LSTM. The Actor network outputs impedance parameters based on the time-series state , while the Critic network realizes the stable update of the policy by minimizing the loss function and combining entropy regularization. The technical solution of this embodiment does not need to rely on a passive energy absorption structure, constructs a reinforcement learning policy model, which can accurately capture the dynamic characteristics of the flutter phenomenon and provide effective vibration suppression measures for the robot in real time, thereby realizing a robot grinding system integrating measurement and control functions. On the one hand, this embodiment realizes the real-time suppression and efficient control of the complex flutter mechanism during the robot grinding of thin-walled parts, significantly improving the self-adaptability and control accuracy of the system; on the other hand, this embodiment fully solves the key problems such as multi-source data heterogeneity, time-series correlation, and insufficient adaptability of traditional methods
[0075] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously
[0076] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element
[0077] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A robot grinding online chatter suppression method based on reinforcement learning, characterized in that: include: Collecting discrete vibration signals during the robot grinding process, and determining a force error signal during the grinding process, wherein the force error signal is used to characterize a force control deviation during the grinding process; The discrete vibration signal and the force error signal are used as multi-source signals for preprocessing and normalized feature extraction fusion to obtain a normalized feature vector with high recognition; Convolutional neural network and multi-head attention module are used to perform feature fusion and dimensionality reduction on the normalized feature vector to obtain low-dimensional feature representation; Temporal encoding is performed on the low-dimensional feature representation to obtain a temporal state, and a temporal enhanced policy model is constructed through a deep reinforcement learning framework embedded with LSTM, wherein the policy model includes an Actor network and a Critic network of a deep deterministic policy gradient algorithm; Using the timing state as an input of an Actor network to obtain an output impedance parameter, and using the timing state and the impedance parameter as inputs of a Critic network to obtain an output quality evaluation value; Based on the quality evaluation value, the robot control parameters are updated by using a minimum loss function combined with entropy regularization optimization, and the robot control parameters are used as a chatter suppression strategy to regulate the robot end effector; Wherein, the robot control parameters include updated impedance parameters.
2. The method according to claim 1, characterized in that Collecting discrete vibration signals during the robot grinding process and determining the force error signal during the grinding process include: The discrete vibration signal collected during the robot grinding process is defined as , , and define the force error signal during robot grinding as ; in, Indicates time, It is the original discrete vibration signal, which has rich time-frequency information and nonlinear characteristics. is the force expected during robot grinding, is the actual measured force.
3. The method according to claim 1, characterized in that The discrete vibration signal and the force error signal are used as multi-source signals for preprocessing and normalized feature extraction fusion to obtain a normalized feature vector with high recognition, including: Based on the selected wavelet basis function, the discrete vibration signal is subjected to multi-layer discrete wavelet decomposition and energy entropy quantification to obtain vibration signal features characterizing the flutter evolution trend, and original information of the force error signal is retained, and features of the force error signal are statistically analyzed by variance and / or mean difference to obtain force error signal features; The vibration signal feature and the force error signal feature are concatenated, and scale differences are eliminated by normalization to obtain a normalized feature vector.
4. The method according to claim 3, characterized in that The discrete vibration signal is subjected to multi-layer discrete wavelet decomposition and energy entropy quantification based on the selected wavelet basis function to obtain vibration signal features characterizing the flutter evolution trend, and the original information of the force error signal is retained, and the features of the force error signal are statistically analyzed by variance and / or mean difference to obtain the force error signal features, including: according to , combined with the preset number of decomposition layers, the selected wavelet basis function is used to decompose the original vibration signal Perform discrete wavelet decomposition layer by layer to obtain low-frequency approximate coefficients and high-frequency detail coefficients after decomposition of each layer in wavelet transform; Combine the low-frequency approximate coefficients and high-frequency detail coefficients of each layer to get the final decomposition ; For the final decomposition ,according to ,right Each subband Calculating Energy ; Based on the energy calculated for each subband ,according to , calculate the energy proportion of each sub-band, and according to Define wavelet energy entropy , wavelet energy entropy is used to quantify the complexity and nonlinear characteristics of vibration signals, as well as important features that characterize the flutter evolution trend; Determine the vibration signal characteristics based on the wavelet energy entropy and the energy calculated for each sub-band; Retention force error signal Original information, and based on Calculate statistical features to obtain force error signal features; in, is the decomposition layer, The low-frequency coefficients obtained by layer decomposition are , the high frequency coefficient is , As the final decomposition in the wavelet decomposition, it means In the subband wavelet coefficients, is the total number of subbands, is the sampling point.
5. The method according to claim 4, characterized in that The vibration signal feature and the force error signal feature are concatenated, and the scale difference is eliminated by normalization to obtain a normalized feature vector, including: The vibration signal features extracted from the discrete vibration signal and , and the force error signal characteristics extracted from the force error signal and Perform feature cascading and according to Perform normalization processing to eliminate scale differences and obtain a normalized feature vector; in, is the normalized feature vector.
6. The method according to claim 5, characterized in that The convolutional neural network and multi-head attention module are used to perform feature fusion and dimensionality reduction on the normalized feature vector to obtain a low-dimensional feature representation, including: The normalized feature vector is input into the convolution layer, and the output feature map is obtained after a series of convolution and pooling operations. ; Feature map output by convolutional layer As a benchmark, a multi-head attention module is used to perform weighted fusion on the normalized feature vector. For feature maps Perform linear transformation to obtain query ,key With value ; By query ,key With value For input, according to , the attention output of each head is calculated by scaled dot product attention; according to , concatenate the attention outputs of each head, and obtain the final multi-head attention output through linear transformation; according to , the multi-head attention output is further reduced in dimension through a fully connected layer or a pooling layer to obtain a low-dimensional feature representation; in, are all learnable weight matrices, is the dimension of the key vector, is the number of heads, Represents the attention output of the corresponding head.
7. The method according to claim 1, characterized in that The low-dimensional feature representation is temporally encoded to obtain a temporal state, and a temporal enhancement strategy model is constructed through a deep reinforcement learning framework embedded with LSTM, including: The LSTM network is embedded in the reinforcement learning framework to perform temporal encoding on the preprocessed feature sequence and , generate a time series state containing historical dynamic information ; Construct an Actor-Critic framework based on the deep deterministic policy gradient algorithm to obtain a timing-enhanced policy model.
8. The method according to claim 7, characterized in that Using the timing state as an input of an Actor network to obtain an output impedance parameter, and using the timing state and the impedance parameter as an input of a Critic network to obtain an output quality evaluation value, including: In the Actor network, the state is in the form of time series For input, according to Time limit status Perform full connection layer and LSTM layer processing to obtain output action , as a parameter for regulating the terminal impedance of the robot; By time series status and output actions As the input of the Critic network, according to Output corresponding to value , used as a quality evaluation value to evaluate the quality of the current chatter suppression strategy; in, is the stiffness in the impedance parameter, is the damping in the impedance parameter, Represents the parameters of the Actor network, are the parameters of the Critic network.
9. The method according to claim 8, characterized in that Based on the quality evaluation value, the robot control parameters are updated by using a minimum loss function combined with entropy regularization optimization, and the robot control parameters are used as a chatter suppression strategy to regulate the robot end effector, including: When training the policy model, an entropy regularization term is introduced into the loss function of the Critic network. ; Add entropy regularization optimization to the loss function of the Critic network Dynamically adjust the exploration weights to maximize the output of the Critic network value; Based on the maximum Value, use Update the Actor network and obtain updated robot control parameters through the updated Actor network and Critic network; in, is the discount factor, is the target Critic network, The output of the target Actor network is used to process time series data, random exploration efficiency, and training convergence speed in the reinforcement learning process through weights. Perform weighted sampling to enhance the model.
Citation Information
Patent Citations
Structural vibration control method based on reinforcement learning, medium and equipment
CN112698572A
Deep reinforcement learning vibration suppression system and method based on unknown mechanical arm model
CN114932546A
Reward function and vibration suppression reinforcement learning algorithm using same
CN115327927A
Multi-attention mechanism gearbox fault diagnosis method, system and device
CN116124449A
Milling cutter damage state monitoring device and monitoring method thereof
CN118123583A
Cited By
Robot multi-frequency vibration composite suppression method and system based on inertia disturbance analysis
CN120921412A
Biodegradable film production line control method and system based on machine learning
CN121028706A
Feeder fault detection and processing method and system
CN121523293A