Robot operation state identification method based on SGMD-LSTM-Transform fusion network
Through the SGMD-LSTM-Transformer fusion network, the modeling incompatibility and noise interference problems in industrial robot state recognition are solved, and efficient and accurate state recognition is achieved.
Patent Information
- Application Number
- CN202510481747.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-08
Smart Images

Figure CN120448866A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of industrial robot status monitoring, and specifically provides a robot operation status recognition method based on an SGMD-LSTM-Transformer fusion network. Background Art
[0002] Industrial robots undertake various types of processing tasks, such as grasping objects with different loads and executing different motion trajectories. Their operating conditions vary greatly as the tasks they undertake change. Once a fault occurs, the distribution and statistical characteristics of the fault data under different working conditions are different. Each working condition requires its corresponding fault diagnosis model, which poses a great challenge to the monitoring and diagnosis of industrial robots. Therefore, accurate identification of the operating status is very necessary for the establishment of subsequent diagnostic models.
[0003] In the entire life cycle of an industrial robot, steady state, that is, the working state in which the robot is performing a task, occupies the vast majority. Unlike transient state with high randomness and static state with low information availability, in steady state, robots have different characteristics for different tasks. Early research on robot state recognition mainly focused on data analysis and modeling. Starting from the state of the robot itself, robot state recognition was achieved through mathematical or physical analysis, automation technology, deep learning, etc. However, when dealing with complex environments or robot applications with multiple states, there are the following disadvantages: (1) Single modeling and kinematic analysis cannot adapt to rapidly changing dynamic environments, especially when the robot operation involves multiple joints and is subject to external force interference, it is impossible to accurately predict the real-time state of the robot. (2) Analysis based on current signals often stops at the working state of the motor, and there is no in-depth study of the overall operating state of the robot. 3) The current sampling recognition method is more suitable for small-scale data, while robots often run for a long time, the sample size is large, and the recognition accuracy is low. (4) Since the robot system is highly coupled and sensitive to external environmental interference, there is non-negligible noise in the collected experimental signals. How to process the signal and build a suitable network model for state recognition is a problem that needs to be solved.
[0004] Symplectic Geometric Modal Decomposition (SGMD) is a signal decomposition method based on symplectic geometry analysis. It aims to preserve the geometric characteristics of the signal and separate its modal components. This method achieves multi-scale signal decomposition by constructing a trajectory matrix, matrix decomposition, and modal extraction, making it suitable for the analysis and processing of non-stationary signals. Therefore, this patent proposes a new method for robot operation state identification based on SGMD. It constructs a combined modal component selection strategy to determine the optimal component and constructs a model evaluation metric to measure the accuracy of the result classification. This has important implications for predictive maintenance of robots. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a robot operation state recognition method based on SGMD-LSTM-Transformer fusion network, aiming to improve the efficiency and accuracy of robot operation state recognition.
[0006] To achieve the above object, the technical solution adopted by the present invention is:
[0007] To solve the above technical problems, the present invention proposes a robot operation state recognition method based on SGMD-LSTM-Transformer fusion network, aiming to improve the efficiency and accuracy of robot operation state recognition.
[0008] To achieve the above object, the technical solution adopted by the present invention is:
[0009] The robot operation state recognition method based on the SGMD-LSTM-Transformer fusion network includes the following steps:
[0010] S1: Build a robot joint motor current signal acquisition platform to obtain joint motor current signals;
[0011] Build a robot joint motor current signal acquisition platform, determine the sensor model and layout, and set the sampling frequency to f s , N times in time T s Sampling times to obtain the joint motor current signal x;
[0012] S2: Perform SGMD decomposition on the motor current signal to obtain a single modal component set;
[0013] Based on symplectic geometric mode decomposition, the motor current signal x is decomposed to obtain a single modal component set SCGs = [C1, C2, C3, ..., C N ];
[0014] S3: Construct a set of combined modal components and select the optimal combined modal components based on the screening strategy of weighted mutual information;
[0015] The N components in the single modal component set SCGs are reorganized and merged according to the number of combinations, and the combined modal component set is expanded:
[0016] SCGs extend =[C1,C2,...,C N ,C 12 ,C 13 ,...,C 123 ,C 124 ,...,C 12...(N-1) ,...,C 12...N ]
[0017] It includes the extended modal component set and the original single modal component set, a total of 2 N -1 modal component. By calculating the weighted mutual information MI of each combined modal component W , and with the maximum weighted mutual information MI W_MAX The combined modal components of are selected as the components X selected ;
[0018] S4: Construct an SGMD-LSTM-Transformer fusion network to model and analyze the current signal;
[0019] An SGMD-LSTM-Transformer fusion network is constructed to process the combined modal components. Based on the LSTM gate mechanism and before the fully connected layer, a dual encoder module is introduced to further enhance the modeling and learning capabilities of the current signal. At the same time, the combined modal components obtained by SGMD decomposition and expansion are Hilbert transformed to obtain the signal envelope Env selected , as the input of the fusion network X input ;
[0020] S5: The selected optimal modal components are used as the input of the fusion network to obtain the robot operation status information;
[0021] Fusion network based on S4 and input X input , get the label prediction result Label out , and based on the predicted label Label out The original current signal x is classified to obtain the robot operation status information.
[0022] Its further technical solution is:
[0023] Step S2, performing signal decomposition on the motor current signal based on symplectic geometric mode decomposition (SGMD), includes the following steps:
[0024] The current signal x is downsampled by a factor of N ds The downsampling process is performed to obtain the downsampled signal x ds ={x1,x2,...,x n};
[0025] Construct trajectory matrix X t :
[0026]
[0027] Where n is the input signal length and d is the delay length of the trajectory matrix, which is calculated by the following formula:
[0028]
[0029] Among them, f s and f max Represent the sampling frequency and the main frequency of the power spectrum of the input signal respectively. m is calculated by the following formula:
[0030] m=n-(d-1)
[0031] Calculate the covariance approximation matrix A:
[0032] A=X t T X t
[0033] Extract the basis vectors in the covariance approximation matrix A based on Schur decomposition:
[0034] A=QRQ T
[0035] Where Q = {q1,q2,...,q m} is a unitary matrix, that is, the complex expansion form of the orthogonal matrix, satisfying QQ T =I, its column vectors {q i |i=1,2,...,m} are basis vectors, R is an upper triangular matrix containing all eigenvalues of matrix A;
[0036] The timing representation of Q is Z={Z1, Z2, ..., Z i ,...,Z d}, basis vector q i The timing representation of Z i :
[0037]
[0038] Calculate the primary modal component matrix Y={Y1,Y2,...,Y i ,...,Y d},
[0039] Y i =mean(diag(flip(Z i )))
[0040] Among them, flip(·) means matrix flipping, diag(·) means diagonal element extraction, and mean(·) means taking the mean.
[0041] Calculate the primary modal component matrix Y i and Y j The correlation coefficient ρ ij :
[0042] ρ ij =corrcoef(Yi ,Y j )(i≠j)
[0043] Among them, corrcoef(·) represents the correlation coefficient calculation. When ρ ij ≥ρ threshold When it is greater than or equal to the set threshold ρ threshold When , the two can be superimposed to obtain Yij, generating the high-level modal component set Y advance ={Y ij |ρ ij ≥ρ threshold};
[0044] Calculate the reconstruction error:
[0045]
[0046] Among them, NMSE is the normalized mean square error, is the signal x ds DC component, is the sum of the currently extracted high-level modal components, H is the number of components in the current high-level modal component set, and ||·|| represents the Euclidean norm. When NMSE≤NMSE threshold , that is, less than or equal to the threshold NMSE threshold , then the modal components are considered to have been fully reconstructed, and the current high-level modal component set Y advance As the final modal component set SCGs=[C1,C2,C3,...,C N ].
[0047] Step S3, constructing a weighted mutual information index to evaluate the combined modalities before training the network model, and selecting the combined modal component with the maximum weighted mutual information as the input of the network model, including the following steps:
[0048] Compute each combined modal component:
[0049] SCGs extend =[C1,C2,...,C N ,C 12 ,C 13 ,...,C 123 ,C 124 ,...,C 12...(N-1) ,...,C 12...N ]’s average mutual information:
[0050]
[0051] Where L is the signal label, each combined modal component corresponds to the same label, c and l are the combined modal component value and label value respectively, P(c,l) is the joint distribution probability, P(c) and P(l) are the marginal distribution probabilities, and N c is the number of modes in the i-th combined mode;
[0052] Calculate standard mutual information MI and weighted mutual information MI W :
[0053]
[0054] MI W =0.5*MI+0.5*MI A
[0055] Among them, MI is the standard mutual information, which is different from the average mutual information MI A Perform weighted processing to obtain the final weighted mutual information MI W .
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] The present invention effectively realizes the operation state recognition of the robot through the SGMD-LSTM-Transformer fusion network. (1) This technology proposes an SGMD combined mode selection method, which selects the optimal mode component through phase space reconstruction, basis vector extraction, modal component reconstruction, combined modal component expansion, and weighted mutual information index construction. It can expand the available data as much as possible before training while improving the analysis and processing efficiency, laying the foundation for the subsequent network model processing. (2) The dual encoder module complements the LSTM network model's modeling ability for long current time series. Compared with the traditional network model, it shows better results in the constructed label jump and fluctuation model evaluation indicators. In label classification, it also achieves higher recognition accuracy and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a schematic diagram of the process of the present invention;
[0059] Figure 2 A joint motor current signal acquisition system constructed for a specific embodiment of the present invention;
[0060] Figure 3 There are three types of working conditions designed for the specific embodiment of the present invention;
[0061] Figure 4 Current signals (with labels) collected according to a specific embodiment of the present invention;
[0062] Figure 5is the combined modal mutual information value obtained in a specific embodiment of the present invention;
[0063] Figure 6 The SGMD-LSTM-Transformer fusion network structure constructed for the specific embodiment of the present invention;
[0064] Figure 7 The optimal combined modal component prediction label obtained in a specific embodiment of the present invention;
[0065] Figure 8 This is the state recognition result obtained in a specific embodiment of the present invention. DETAILED DESCRIPTION
[0066] The following is a further description of the present invention through specific examples:
[0067] For reference Figure 1 , a robot operation state recognition method based on the SGMD-LSTM-Transformer fusion network of the present application includes the following steps:
[0068] S1: Build a robot joint motor current signal acquisition platform to obtain joint motor current signals;
[0069] Build a robot joint motor current signal acquisition platform, determine the sensor model and layout, and set the sampling frequency to f s , N times in time T s Sampling times to obtain the joint motor current signal x;
[0070] S2: Perform SGMD decomposition on the motor current signal to obtain a single modal component set;
[0071] The current signal x is downsampled by a factor of N ds The downsampling process is performed to obtain the downsampled signal x ds ={x1,x2,...,x n};
[0072] Construct trajectory matrix X t :
[0073]
[0074] Where n is the input signal length and d is the delay length of the trajectory matrix, which is calculated by the following formula:
[0075]
[0076] Among them, f s and f max Represent the sampling frequency and the main frequency of the power spectrum of the input signal respectively. m is calculated by the following formula:
[0077] m=n-(d-1)
[0078] Calculate the covariance approximation matrix A:
[0079] A=X t T X t
[0080] Extract the basis vectors in the covariance approximation matrix A based on Schur decomposition:
[0081] A=QRQ T
[0082] Where Q = {q1,q2,...,q m} is a unitary matrix, that is, the complex expansion form of the orthogonal matrix, satisfying QQ T =I, its column vectors {q i |i=1,2,...,m} are basis vectors, R is an upper triangular matrix containing all eigenvalues of matrix A;
[0083] The timing representation of Q is Z={Z1, Z2, ..., Z i ,...,Z d}, basis vector q i The timing representation of Z i :
[0084]
[0085] Calculate the primary modal component matrix Y={Y1,Y2,...,Y i ,...,Y d},
[0086] Y i =mean(diag(flip(Z i )))
[0087] Among them, flip(·) means matrix flipping, diag(·) means diagonal element extraction, and mean(·) means taking the mean.
[0088] Calculate the primary modal component matrix Y i and Y j The correlation coefficient ρ ij :
[0089] ρ ij =corrcoef(Y i ,Y j )(i≠j)
[0090] Among them, corrcoef(·) represents the correlation coefficient calculation. When ρij ≥ρ threshold When it is greater than or equal to the set threshold ρ threshold When , the two can be superimposed to obtain Y ij , generate high-level modal component set Y advance ={Y ij |ρ ij ≥ρ threshold};
[0091] Calculate the reconstruction error:
[0092]
[0093] Among them, NMSE is the normalized mean square error, is the signal x ds DC component, is the sum of the currently extracted high-level modal components, H is the number of components in the current high-level modal component set, and ||·|| represents the Euclidean norm. When NMSE≤NMSE threshold , that is, less than or equal to the threshold NMSE threshold , then the modal components are considered to have been fully reconstructed, and the current high-level modal component set Y advance As the final modal component set SCGs=[C1,C2,C3,...,C N ].
[0094] S3: Construct a set of combined modal components and select the optimal combined modal components based on the screening strategy of weighted mutual information;
[0095] Compute each combined modal component:
[0096] SCGs extend =[C1,C2,...,C N ,C 12 ,C 13 ,...,C 123 ,C 124 ,...,C 12...(N-1) ,...,C 12...N ]’s average mutual information:
[0097]
[0098] Where L is the signal label, each combined modal component corresponds to the same label, c and l are the combined modal component value and label value respectively, P(c,l) is the joint distribution probability, P(c) and P(l) are the marginal distribution probabilities, and N c is the number of modes in the i-th combined mode;
[0099] Calculate standard mutual information MI and weighted mutual information MI W :
[0100]
[0101] MI W =0.5*MI+0.5*MI A
[0102] Among them, MI is the standard mutual information, which is different from the average mutual information MI A Perform weighted processing to obtain the final weighted mutual information MI W By calculating the weighted mutual information MI of each combined modal component W , and with the maximum weighted mutual information MI W_MAX The combined modal components of are selected as the components X selected ;
[0103] S4: Construct an SGMD-LSTM-Transformer fusion network to model and analyze the current signal;
[0104] An SGMD-LSTM-Transformer fusion network is constructed to process the combined modal components. Based on the LSTM gate mechanism and before the fully connected layer, a dual encoder module is introduced to further enhance the modeling and learning capabilities of the current signal. At the same time, the combined modal components obtained by SGMD decomposition and expansion are Hilbert transformed to obtain the signal envelope Env selected , as the input of the fusion network X input ;
[0105] S5: The selected optimal modal components are used as the input of the fusion network to obtain the robot operation status information;
[0106] Fusion network based on S4 and input X input , get the label prediction result Label out , and based on the predicted label Label out The original current signal x is classified to obtain the robot operation status information.
[0107] The technical solution of the present application is further illustrated below with specific examples.
[0108] Figure 2 The joint motor current signal acquisition system is demonstrated. The experimental system consists of a data collector, a motion controller, a platform base and a SCARA robot. For the three rotary joint motors, the following Figure 3The three types of trajectory working conditions shown in the figure have a sampling frequency of 10 kHz and a sampling time of 5 minutes. The current signal data is obtained from the industrial robot through three current transformers. The data set is divided into samples with a duration of 100 seconds. Each sample contains 1000K data points. The time domain waveform is shown in the figure below. Figure 4 shown.
[0109] The original current signal is preliminarily decomposed based on SGMD, that is, the current signal data is subjected to phase space reconstruction, basis vector extraction, modal component reconstruction and other operations to obtain SCGs = [C1, C2, C3, C4, C5].
[0110] By combining the numbers, the single modal components are spliced and merged row by row to expand the combined modal component set:
[0111] SCGs extend =[C1,C2,...,C5,C 12 ,C 13 ,...,C 45 ,C 123 ,C 124 ,...,C 1234 ,...,C 2345 ,C 12345 ],
[0112] Calculate the MI of each combined mode W , and the combined modal component with the maximum weighted mutual information is selected as the component X selected . Figure 5 The specific changes in the three mutual information values of different combination modes are shown. The combination mode component with the largest weighted mutual information value is C0123 shown in the figure.
[0113] Figure 6 The basic structure of the SGMD-LSTM-Transformer fusion network is shown, which is a key step in state recognition. The C0123 combined modal component with the best performance is used as the input of the SGMD-LSTM-Transformer fusion network, and the following is obtained: Figure 7 The label prediction situation shown in the figure is as follows. According to the label value, the original signal is cut according to the time step, and the signals with the same label are classified into one category. The result is as follows Figure 8 As shown, labels 1, 2, and 3 correspond to Figure 3 The A, B, and C states are marked in the figure, while 0 corresponds to the unmarked part, that is, state D. It can be seen that the current signal collected from the robot joint has been successfully cut and displayed according to the label category, realizing the accurate classification of the robot's working state and achieving good recognition effect.
[0114] The above description is merely a preferred embodiment of the present invention and does not constitute any other form of limitation to the present invention. Any modification or equivalent variation based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.
Claims
1. A robot operation state recognition method based on SGMD-LSTM-Transformer fusion network, characterized by: The following steps are involved: S1: Build a robot joint motor current signal acquisition platform to obtain joint motor current signals; Build a robot joint motor current signal acquisition platform, determine the sensor model and layout, and set the sampling frequency to f s , N times in time T s Sampling times to obtain the joint motor current signal x; S2: Perform SGMD decomposition on the motor current signal to obtain a single modal component set; Based on symplectic geometric mode decomposition, the motor current signal x is decomposed to obtain a single modal component set SCGs = [C1, C2, C3, ..., C N ]; S3: Construct a set of combined modal components and select the optimal combined modal components based on the screening strategy of weighted mutual information; The N components in the single modal component set SCGs are reorganized and merged according to the number of combinations, and the combined modal component set SCGs is expanded. extend =[C1,...,C 12 ,...,C 123 ,...,C 12...N ], which includes the extended modal component set and the original single modal component set, a total of 2 N -1 modal component, by calculating the weighted mutual information MI of each combined modal component W , and with the maximum weighted mutual information MI W_MAX The combined modal components of are selected as the components X selected ; S4. Construct an SGMD-LSTM-Transformer fusion network to model and analyze the current signal; An SGMD-LSTM-Transformer fusion network is constructed to process the combined modal components. Based on the LSTM gate mechanism, a dual encoder module is introduced before the fully connected layer to further enhance the modeling and learning capabilities of the model for the current signal. At the same time, in order to adapt to the network model, the combined modal components obtained by SGMD decomposition and expansion are Hilbert transformed to highlight the amplitude change pattern of the signal and obtain the signal envelope Env selected , as the input of the fusion network X input ; S5. Using the selected optimal modal components as the input of the fusion network to obtain the robot operation status information; Based on the fusion network input X mentioned in S4 input , get the label prediction result Label out , and based on the predicted label Label out The original current signal x is classified to obtain the robot operation status information.
2. The robot operation state recognition method based on the SGMD-LSTM-Transformer fusion network according to claim 1 is characterized in that: Step S2, performing signal decomposition on the motor current signal based on symplectic geometric mode decomposition (SGMD), includes the following steps: The current signal x is downsampled by a factor of N ds The downsampling process is performed to obtain the downsampled signal x ds ={x1,x2,...,x n }; Construct trajectory matrix X t : Where n is the input signal length and d is the delay length of the trajectory matrix, which is calculated as follows: Among them, f s and f max Represent the sampling frequency and the main frequency of the power spectrum of the input signal respectively, and m is calculated by the following formula: m=n-(d-1) Calculate the covariance approximation matrix A: A=X t T X t Extract the basis vectors in the covariance approximation matrix A based on Schur decomposition: A=QRQ T Among them, Q is a unitary matrix, that is, the complex expansion form of the orthogonal matrix, satisfying QQ T =I, its column vectors {q i |i=1,2,...,m} are basis vectors; R is an upper triangular matrix containing all eigenvalues of matrix A; The timing representation of Q is Z={Z1, Z2, ..., Z i ,...,Z d }, basis vector q i The timing representation of Z i : Calculate the primary modal component matrix Y={Y1,Y2,...,Y i ,...,Y d }, Y i =mean(diag(flip(Z i ))) Among them, flip(·) means matrix flipping, diag(·) means diagonal element extraction, and mean(·) means taking the mean; Calculate the primary modal component matrix Y i and Y j The correlation coefficient ρ ij : ρ ij =corrcoef(Y i ,AND j )(i≠j) Among them, corrcoef(·) represents the correlation coefficient calculation. When ρ ij ≥ρ threshold When it is greater than or equal to the set threshold ρ threshold When , the two can be superimposed to obtain Y ij , generate high-level modal component set Y advance ={Y ij |ρ ij ≥ρ threshold }; Calculate the reconstruction error: Among them, NMSE is the normalized mean square error, is the signal x ds DC component, is the sum of the currently extracted high-level modal components, H is the number of components in the current high-level modal component set, ||·|| represents the Euclidean norm, and when NMSE≤NMSE threshold , that is, less than or equal to the threshold NMSE threshold , then the modal components are considered to have been fully reconstructed, and the current high-level modal component set Y advance As the final modal component set SCGs=[C1,C2,C3,...,C N ].
3. The robot operation state recognition method based on the SGMD-LSTM-Transformer fusion network according to claim 1 is characterized in that: Step S3, constructing a weighted mutual information index of the combined modal component set to evaluate the combined modalities before training the network model, and selecting the combined modal component with the maximum weighted mutual information as the input of the network model, including the following steps: Compute each combined modal component: Where L is the signal label, P(c,l) is the joint distribution probability, P(c) and P(l) are the marginal distribution probabilities, and N c is the number of modes in the i-th combined mode. The average modal mutual information can eliminate the influence of the number of modes, but does not take into account the cumulative effect of the modes. Therefore, it is necessary to further calculate the standard mutual information MI and the weighted mutual information MI W : MI W =0.5*MI+0.5*MI A Among them, MI is the standard mutual information, which is different from the average mutual information MI A Perform weighted processing to obtain the final weighted mutual information MI W .