Robot Cable Fault Classification Method and System Based on Deep Tree Learning

Through the deep tree learning method, combining the robot's motion state and cable bending characteristics, a multi-layer decision tree classifier is built, which solves the problem of unstable accuracy of robot cable fault classification in the existing technology, and achieves efficient fault identification and robustness improvement in different motion states.

CN120217126BActive Publication Date: 2025-07-29NINGBO RIYUE ELECTRIC WIRE & CABLES MFG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510704297.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-29
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

The prior art fails to effectively consider the robot motion state and cable bending characteristics in the classification of robot cable faults, resulting in unstable classification accuracy, especially in different motion states, and it is difficult to identify minor faults, and lacks the ability to fusion of multi-scale features.

Method used

A deep tree learning method is adopted to collect robot cable signals and joint position information, divide the motion state, and build a deep tree learning classifier with a multi-layer decision tree structure, combining attention connection mechanism and multi-scale splitting criterion for feature transmission and fusion, and dynamically adjust feature weights to adapt to different time windows.

Benefits of technology

It improves the accuracy and robustness of robot cable fault classification, can effectively identify cable faults in complex industrial environments, and enhances the perception of key features and generalization capabilities of models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217126B_ABST
    Figure CN120217126B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for classifying robot cable faults based on deep tree learning, which relates to the technical field of robots. The method includes collecting cable signals and joint position information, and calculating bending degree features; extracting voltage, current, power, and impedance features according to the motion state of the robot to form a state-sensitive feature set; constructing a deep tree learning classifier, and performing feature transfer through an attention connection mechanism; and adopting a multi-scale splitting criterion for local optimization and global optimization for training. The present invention can effectively identify cable faults in different motion states, and improve the classification accuracy and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to robot technology, and particularly to a method and system for classifying robot cable faults based on deep tree learning. Background Art

[0002] As a core equipment of modern intelligent manufacturing, the industrial robot, as one of the key components of the robot, the cable bears continuous bending stress and electrical load during the movement of the robot, and is prone to faults such as aging, wear, and wire breakage. At present, the diagnosis of robot cable faults mainly relies on signal analysis and machine learning methods. Traditional methods mainly rely on technologies such as time-frequency domain analysis and wavelet transform to extract cable state features, and then combine classification algorithms such as support vector machines and random forests for fault identification. In recent years, deep learning technologies such as convolutional neural networks and long short-term memory networks have also been applied to the field of cable fault diagnosis, improving the diagnosis performance by automatically learning the feature hierarchy.

[0003] However, the existing technologies still have obvious deficiencies in the classification of robot cable faults. First of all, most methods ignore the influence of the robot's motion state on the cable signal features. There are obvious differences in the feature distributions of the data collected under different motion states, resulting in unstable classification accuracy. Secondly, traditional feature extraction methods are difficult to effectively capture the fault features of the cable under different bending degrees, especially the boundary between the slight fault state and the normal state is blurred, which is prone to misjudgment. Finally, the existing classification models lack the ability of multi-scale analysis of time series features and are difficult to adapt to the dynamic changes of fault features under different time windows. Especially during the state transition processes such as acceleration and deceleration, the fault diagnosis accuracy is significantly reduced.

[0004] Therefore, there is an urgent need to develop a fault classification method that can comprehensively consider the robot's motion state and the cable bending characteristics and has the ability of multi-scale feature fusion to improve the accuracy and robustness of robot cable fault diagnosis. Summary of the Invention

[0005] The embodiments of the present invention provide a method and system for classifying robot cable faults based on deep tree learning, which can solve the problems in the existing technologies.

[0006] In the first aspect of the embodiments of the present invention, a method for classifying robot cable faults based on deep tree learning is provided, including:

[0007] Collect the voltage signal and current signal of the robot cable, obtain the joint position information of the robot, and calculate the cable bending degree feature based on the joint position information;

[0008] Divide the robot motion state into a stationary state, a uniform motion state, and an acceleration / deceleration state, extract voltage, current, power, and impedance features for each state, and combine them with the cable bending degree feature to form a state-sensitive feature set;

[0009] Construct a deep tree learning classifier with a multi-layer decision tree structure; input the state-sensitive feature set into the first-layer decision tree unit, and the decision tree units of adjacent layers perform feature transfer through an attention connection mechanism, where the between-class scatter and within-class scatter of features are calculated based on the Fisher discriminant criterion, an attention weight matrix is constructed, and the transferred features are weighted and combined using the attention weight matrix and the adaptive fusion weight to obtain the fused features;

[0010] Train the deep tree learning classifier, including: locally optimizing the decision tree unit using a multi-scale splitting criterion, which calculates the information gain and dynamically adjusts the window weight under different time windows; performing global optimization and introducing regularization constraints to control the feature differences between adjacent-layer decision tree units;

[0011] Input the robot cable signal to be classified into the trained deep tree learning classifier to obtain the fault type.

[0012] In an alternative embodiment,

[0013] The formation of the state-sensitive feature set includes:

[0014] Based on the joint angular velocity vector, joint angular acceleration vector, and end effector motion parameters of the robot, the robot motion state is divided into a stationary state, a uniform motion state, and an acceleration / deceleration state;

[0015] Extract electrical features for each motion state respectively, including: calculating the mean, effective value, standard deviation, peak value, and valley value of the voltage signal to obtain the voltage feature vector, calculating the mean, effective value, standard deviation, peak value, valley value, and maximum change rate of the current signal to obtain the current feature vector, calculating the active power, reactive power, and power factor to obtain the power feature vector, calculating the impedance amplitude and phase to obtain the impedance feature vector; respectively assigning weight coefficients corresponding to the motion state to the voltage feature vector, current feature vector, power feature vector, and impedance feature vector for weighted combination to obtain the state-enhanced features;

[0016] Calculate the local bending angle and cumulative bending amount of the cable based on the cable bending degree feature and perform feature fusion with the state-enhanced features to obtain the state perception features; perform feature selection on the state perception features based on the mutual information evaluation method, and select the features with mutual information values greater than the feature screening threshold to form the state-sensitive feature set.

[0017] In an alternative embodiment,

[0018] The voltage feature vector, current feature vector, power feature vector, and impedance feature vector are respectively assigned weight coefficients corresponding to the motion states for weighted combination to obtain state enhancement features, including:

[0019] Obtain parameter feature quantities, where the parameter feature quantities include voltage parameter features, current parameter features, power parameter features, and impedance parameter features;

[0020] Calculate the correlation degree between the motion state of the robot and each parameter feature, including: calculating the Pearson correlation coefficient to obtain the feature-state correlation coefficient, calculating the mutual information to obtain the feature-state mutual information value, and weighting the feature-state correlation coefficient and the feature-state mutual information value to obtain the state sensitivity index;

[0021] Perform normalization processing on the state sensitivity index to obtain the normalized sensitivity, perform exponential transformation on the normalized sensitivity to obtain the basic weight matrix, and dynamically adjust the basic weight matrix through a state conversion factor with an exponential decay relationship with the state duration to obtain the adaptive weight coefficient;

[0022] Use the adaptive weight coefficient to perform weighted feature fusion on the parameter feature quantities in the stationary state, uniform motion state, and acceleration / deceleration state respectively to obtain state enhancement features; among them, calculate the smooth transition feature based on the state transition probability during the state transition, and the state transition probability has an exponential relationship with the feature difference between states.

[0023] In an alternative embodiment,

[0024] Construct a deep tree learning classifier including a multi-layer decision tree structure; input the state-sensitive feature set into the first-layer decision tree unit, and the decision tree units of adjacent layers perform feature transfer through an attention connection mechanism, where the between-class scatter and within-class scatter of the features are calculated based on the Fisher discriminant criterion, an attention weight matrix is constructed, and the transferred features are weighted and combined using the attention weight matrix and the adaptive fusion weight to obtain the fusion features, including:

[0025] Construct a multi-layer decision tree structure, each layer contains multiple decision tree units, and a feature transfer channel is established between the decision tree units of adjacent layers;

[0026] Input the state-sensitive feature set into the first-layer decision tree unit, calculate the between-class scatter between the mean of each category feature and the overall feature mean based on the Fisher discriminant criterion, calculate the within-class scatter between the within-class feature and the category feature mean, use the ratio of the between-class scatter to the within-class scatter as the Fisher discriminant ratio, and obtain the feature importance score through exponential normalization;

[0027] Calculate the covariance between features and obtain the correlation coefficient between feature pairs, and perform exponential weighted operation on the feature importance scores and the correlation coefficient to generate an attention weight matrix;

[0028] Introduce a temperature parameter to perform exponential adjustment on the feature importance scores to obtain an adaptive fusion weight, perform weighted combination on the current layer features and the attention weight matrix to obtain transmitted features, and perform weighted combination on the transmitted features using the adaptive fusion weight to obtain fused features;

[0029] Calculate the Fisher discriminant ratio gain and information gain of the fused features relative to the transmitted features, and adaptively adjust the temperature parameter according to the Fisher discriminant ratio gain and information gain. The fused features are used as input features for the next layer of decision tree units for classification and judgment of fault types.

[0030] In an alternative embodiment,

[0031] The adaptive adjustment of the temperature parameter includes:

[0032] The initial value of the temperature parameter is the mean of the logarithmic values of the feature importance scores;

[0033] Calculate the sensitivity of each feature to the temperature parameter, and the sensitivity is calculated by the product of the adaptive fusion weight and the feature importance score and the product of the adaptive fusion weight and its complementary value;

[0034] Perform exponential transformation and normalization on the sensitivity to obtain a feature enhancement factor, and enhance the transmitted features using the feature enhancement factor to obtain enhanced features;

[0035] Perform weighted combination on the enhanced features using the adaptive fusion weight to obtain optimized fused features, calculate the ratio of the Fisher discriminant ratio of the optimized fused features to the transmitted features to obtain the discriminant ratio gain, and calculate the difference between the information entropy of the optimized fused features and the transmitted features to obtain the information gain;

[0036] Use the weighted sum of the discriminant ratio gain and the information gain as the update basis for the temperature parameter, and update the temperature parameter using the gradient descent method;

[0037] Calculate the change rate of the adaptive fusion weights in two adjacent rounds and the maximum value of the discriminant ratio gain as the convergence criterion, and determine the final temperature parameter when the convergence criterion is less than a preset convergence threshold.

[0038] In an alternative embodiment,

[0039] Training the deep tree learning classifier includes:

[0040] Generate a multi-scale time window sequence, where the multi-scale time window sequence includes a base window and extended windows that are multiplied and extended based on the base window, and divide training samples based on the base window and the extended windows to obtain multiple window sample sets;

[0041] Calculate the probability distribution of samples of each category in each window sample set, calculate the window information entropy based on the probability distribution, calculate the conditional information entropy for each candidate splitting feature, and use the difference between the window information entropy and the conditional information entropy as the window information gain;

[0042] Calculate the time decay weight based on the time difference between the current moment and the central moment of each window, normalize the time decay weight to obtain the window weight, and use the window weight to weight the information gain of each window to obtain the comprehensive information gain;

[0043] Calculate the splitting information entropy of each candidate splitting feature, use the ratio of the comprehensive information gain to the splitting information entropy as the splitting gain ratio, and select the candidate splitting feature with the largest splitting gain ratio as the optimal splitting feature;

[0044] Calculate the difference degree of corresponding features between adjacent layer decision tree units, use the weighted sum of the difference degrees as the regularization constraint term, and use the cumulative sum of the splitting gain ratios of each layer of decision tree units minus the regularization constraint term as the optimization objective function;

[0045] Calculate the classification performance metrics of the deep tree learning classifier on the validation set, adaptively adjust the time decay coefficient, regularization coefficient, and base window length based on the classification performance metrics, and repeat the feature splitting optimization process until the preset convergence standard is reached.

[0046] In an alternative embodiment,

[0047] Calculating the difference degree of corresponding features between adjacent layer decision tree units, using the weighted sum of the difference degrees as the regularization constraint term, and using the cumulative sum of the splitting gain ratios of each layer of decision tree units minus the regularization constraint term as the optimization objective function includes:

[0048] Calculate the feature mapping matrix between adjacent layer decision tree units, where the feature mapping matrix is used to establish the conversion relationship between the feature vectors of the decision tree units of the i-th layer and the i+1-th layer;

[0049] Calculate the Euclidean distance, cosine similarity, and information entropy difference between the feature vector of the decision tree unit of the i-th layer and the feature vector of the decision tree unit of the i+1-th layer after being transformed by the feature mapping matrix, and weight and combine them to obtain the feature difference vector;

[0050] The unit difference degree is obtained by weighted averaging the feature difference vector and the splitting gain ratio of the corresponding feature. The hierarchical weight is set based on the depth of the decision tree unit in the classification path, and the hierarchical weight decays exponentially as the depth increases;

[0051] The global difference degree is obtained by weighted summing the unit difference degrees between all adjacent-layer decision tree units and the corresponding hierarchical weights, and the regularization constraint term is obtained by multiplying the global difference degree by the regularization coefficient;

[0052] The feature mapping matrix and the decision tree structure are optimized based on the optimization objective function.

[0053] In the second aspect of the embodiments of the present invention, a robot cable fault classification system based on deep tree learning is provided, including:

[0054] A first unit, configured to collect voltage signals and current signals of a robot cable, obtain joint position information of the robot, and calculate cable bending degree features based on the joint position information;

[0055] A second unit, configured to divide the robot motion state into a stationary state, a uniform motion state, and an acceleration / deceleration state, extract voltage, current, power, and impedance features for each state, and combine them with the cable bending degree features to form a state-sensitive feature set;

[0056] A third unit, configured to construct a deep tree learning classifier including a multi-layer decision tree structure; input the state-sensitive feature set into the first-layer decision tree unit, and the decision tree units of adjacent layers perform feature transfer through an attention connection mechanism, where the between-class scatter and within-class scatter of the features are calculated based on the Fisher discriminant criterion, an attention weight matrix is constructed, and the transferred features are weighted and combined using the attention weight matrix and an adaptive fusion weight to obtain fusion features;

[0057] A fourth unit, configured to train the deep tree learning classifier, including: locally optimizing the decision tree unit using a multi-scale splitting criterion, where the multi-scale splitting criterion calculates the information gain and dynamically adjusts the window weight under different time windows; performing global optimization, and introducing a regularization constraint to control the feature difference between adjacent-layer decision tree units;

[0058] A fifth unit, configured to input the robot cable signal to be classified into the trained deep tree learning classifier to obtain the fault type.

[0059] In the third aspect of the embodiments of the present invention,

[0060] An electronic device is provided, including:

[0061] A processor;

[0062] A memory for storing processor-executable instructions;

[0063] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0064] In the fourth aspect of the embodiments of the present invention,

[0065] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0066] The robot cable fault classification method based on deep tree learning provided by the present invention can effectively capture the fault characteristics of the cable under different motion conditions and improve the discrimination degree and classification accuracy of the fault characteristics by dividing the robot motion state into a stationary state, a uniform motion state, and an acceleration / deceleration state, and extracting state-sensitive feature sets for different states.

[0067] The deep tree learning classifier constructed by the present invention combines the interpretability of the decision tree and the expressive ability of deep learning, realizes the effective transfer and fusion of features through the attention connection mechanism, and constructs an attention weight matrix by using the Fisher discriminant criterion, enhancing the algorithm's perception ability of key features and improving the recognition accuracy of complex cable fault patterns.

[0068] The multi-scale splitting criterion and global optimization strategy adopted by the present invention can dynamically adjust the feature weights under different time windows, and at the same time control the model complexity through regularization constraints, effectively solving the problem of scale change of cable fault characteristics in the time domain, enhancing the robustness and generalization ability of the algorithm, and being applicable to the robot cable fault diagnosis in complex industrial environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a schematic flowchart of the method of the embodiments of the present invention;

[0070] Figure 2 It is a flowchart of feature mapping and regularization optimization of the deep tree learning classifier. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0072] The technical solution of the present invention will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0073] Figure 1 As shown in the flowchart of the robot cable fault classification method based on deep tree learning in the embodiments of the present invention, Figure 1 as shown, the method includes:

[0074] Collect the voltage signal and current signal of the robot cable, obtain the joint position information of the robot, and calculate the cable bending degree feature based on the joint position information;

[0075] Divide the robot motion state into a stationary state, a uniform motion state, and an acceleration / deceleration state. Extract voltage, current, power, and impedance features for each state, and combine them with the cable bending degree feature to form a state-sensitive feature set;

[0076] Construct a deep tree learning classifier including a multi-layer decision tree structure; input the state-sensitive feature set into the first-layer decision tree unit, and the decision tree units of adjacent layers perform feature transfer through an attention connection mechanism, where the between-class scatter and within-class scatter of the features are calculated based on the Fisher discriminant criterion, an attention weight matrix is constructed, and the transferred features are weighted and combined using the attention weight matrix and an adaptive fusion weight to obtain a fusion feature;

[0077] Train the deep tree learning classifier, including: locally optimize the decision tree unit using a multi-scale splitting criterion, and the multi-scale splitting criterion calculates the information gain and dynamically adjusts the window weight under different time windows; perform global optimization, and introduce regularization constraints to control the feature differences between the decision tree units of adjacent layers;

[0078] Input the robot cable signal to be classified into the trained deep tree learning classifier to obtain the fault type.

[0079] In an alternative embodiment, the formation of the state-sensitive feature set includes:

[0080] Based on the joint angular velocity vector, joint angular acceleration vector, and end effector motion parameters of the robot, divide the robot motion state into a stationary state, a uniform motion state, and an acceleration / deceleration state;

[0081] Extract electrical features for each motion state respectively, including: calculating the mean value, root mean square value, standard deviation, peak value and valley value of the voltage signal to obtain a voltage feature vector, calculating the mean value, root mean square value, standard deviation, peak value, valley value and maximum change rate of the current signal to obtain a current feature vector, calculating the active power, reactive power and power factor to obtain a power feature vector, and calculating the impedance amplitude and phase to obtain an impedance feature vector; assigning weight coefficients corresponding to the motion state to the voltage feature vector, current feature vector, power feature vector and impedance feature vector respectively for weighted combination to obtain a state enhanced feature;

[0082] Calculate the local bending angle and cumulative bending amount of the cable based on the cable bending degree feature and perform feature fusion with the state enhanced feature to obtain a state perception feature; perform feature selection on the state perception feature based on the mutual information evaluation method, and select features with mutual information values greater than the feature screening threshold to form a state sensitive feature set.

[0083] Exemplarily, during the operation of the robot, the system collects the robot joint angular velocity vector, joint angular acceleration vector and end effector motion parameters. These parameters are used to determine the current motion state of the robot. Specifically, when the magnitudes of both the robot joint angular velocity vector and the joint angular acceleration vector are less than the preset threshold, for example, the joint angular velocity is less than 0.01 radian / second and the joint angular acceleration is less than 0.005 radian / second², it is determined that the robot is in a stationary state. When the magnitude of the joint angular velocity vector is greater than the first threshold and the magnitude of the joint angular acceleration vector is less than the second threshold, for example, the joint angular velocity is greater than 0.05 radian / second and the joint angular acceleration is less than 0.01 radian / second², it is determined that the robot is in a uniform motion state. When the magnitude of the joint angular acceleration vector is greater than the second threshold, for example, the joint angular acceleration is greater than 0.01 radian / second², it is determined that the robot is in an acceleration / deceleration state.

[0084] For different determined motion states, the system extracts electrical features respectively. For the voltage signal, its mean value, effective value, standard deviation, peak value and valley value are calculated. Taking a set of actual collected data as an example, the mean value of the voltage signal collected under a certain stationary state is 220.5 V, the effective value is 219.8 V, the standard deviation is 0.75 V, the peak value is 221.3 V, and the valley value is 219.2 V. These values form the voltage feature vector. For the current signal, its mean value, effective value, standard deviation, peak value, valley value and maximum change rate are calculated. Under the same stationary state, the mean value of the current signal is 1.25 A, the effective value is 1.28 A, the standard deviation is 0.12 A, the peak value is 1.45 A, the valley value is 1.05 A, and the maximum change rate is 0.08 A / s. These values form the current feature vector. The system also calculates the active power, reactive power and power factor to obtain the power feature vector. Under the above stationary state, the active power is 275.6 W, the reactive power is 52.8 var, and the power factor is 0.98. At the same time, the impedance amplitude and phase are calculated through the complex ratio of voltage and current to obtain the impedance feature vector. Under this state, the impedance amplitude is 176.5 Ω and the phase is 3.2 degrees.

[0085] For different motion states, the system uses different weight coefficients to perform weighted combination on each feature vector to obtain the state-enhanced feature.

[0086] The extraction of the cable bending degree feature is based on the positions of each joint of the robot and the cable routing path. The system calculates the local bending angle of the cable through the joint position information. For example, the bending angle of the cable section between the first joint and the second joint is 28.5 degrees, and the bending angle of the cable section between the second joint and the third joint is 42.3 degrees. The cumulative bending amount is the weighted sum of all local bending angles, which is 105.7 degrees in this example. These cable bending degree features are combined with the state-enhanced features obtained previously through the feature fusion algorithm to form the state perception feature. During the feature fusion process, the weight of the state-enhanced feature is 0.7, and the weight of the cable bending degree feature is 0.3.

[0087] The final formation of the state-sensitive feature set is achieved through the mutual information evaluation method. The system calculates the mutual information value between each state perception feature and the robot fault state. Taking a certain fault state as an example, the mutual information value of the current effective value is 0.85, the mutual information value of the power factor is 0.92, and the mutual information value of the cable cumulative bending amount is 0.78. The system sets the feature screening threshold to 0.7, and selects the features with mutual information values greater than this threshold to form the state-sensitive feature set.

[0088] Through the division of joint motion states and the extraction of state-sensitive features, the present invention correlates cable fault features with the motion states of the robot, significantly improving the accuracy of fault identification under different motion states. The feature selection mechanism based on mutual information effectively reduces feature redundancy and enhances the model's perception ability of key fault features, making fault diagnosis more accurate and reliable.

[0089] In an alternative embodiment, weight coefficients corresponding to the motion states are respectively assigned to the voltage feature vector, current feature vector, power feature vector, and impedance feature vector for weighted combination to obtain state-enhanced features, including:

[0090] Obtain parameter feature quantities, where the parameter feature quantities include voltage parameter features, current parameter features, power parameter features, and impedance parameter features;

[0091] Calculate the correlation degree between the motion state of the robot and each parameter feature, including: calculating the Pearson correlation coefficient to obtain the feature-state correlation coefficient, calculating the mutual information to obtain the feature-state mutual information value, and weighting the feature-state correlation coefficient and the feature-state mutual information value to obtain the state sensitivity index;

[0092] Normalize the state sensitivity index to obtain the normalized sensitivity, perform an exponential transformation on the normalized sensitivity to obtain the basic weight matrix, and dynamically adjust the basic weight matrix through a state transition factor that has an exponential decay relationship with the state duration to obtain the adaptive weight coefficient;

[0093] Use the adaptive weight coefficient to perform weighted feature fusion on the parameter feature quantities in the stationary state, uniform motion state, and acceleration / deceleration state respectively to obtain state-enhanced features; wherein, during the state transition, smooth transition features are calculated based on the state transition probability, and the state transition probability has an exponential relationship with the feature difference between states.

[0094] Exemplarily, electrical parameter data of the robot in different motion states are respectively collected, and the data are obtained based on the aforementioned voltage feature vector, current feature vector, power feature vector, and impedance feature vector. For example, in practical applications, the system can collect data every 100 milliseconds and calculate these statistics within a 1-second time window to obtain a set of parameter feature vectors containing 20 dimensions.

[0095] In the stage of calculating the correlation degree, the Pearson correlation coefficient is calculated to obtain the feature-state correlation coefficient. Taking the root mean square current feature as an example, by analyzing the corresponding relationship between this feature and the stationary state, the correlation coefficient is obtained as 0.82, indicating that the root mean square current has a strong positive correlation with the stationary state. Subsequently, the mutual information value is calculated to measure the non-linear relationship between the feature and the state. Still taking the root mean square current as an example, the mutual information value with the acceleration and deceleration state is 0.76, indicating that this feature has a high discrimination ability for identifying the acceleration and deceleration state. The feature-state correlation coefficient and the mutual information value are weighted according to a ratio of 6:4 to obtain the state sensitivity index. In an actual application case, the sensitivity index of the voltage parameter in the stationary state is 0.65, the sensitivity index of the current parameter in the acceleration and deceleration state is 0.88, and the sensitivity index of the power parameter in the uniform motion state is 0.72.

[0096] The state sensitivity index is normalized. The normalization process uses the maximum-minimum normalization method to map the sensitivity index of each feature into the range of 0-1. For example, the normalized sensitivities of the voltage feature in the three states are: 0.76 in the stationary state, 0.51 in the uniform motion state, and 0.44 in the acceleration and deceleration state. The normalized sensitivities are subjected to an exponential transformation to obtain the basic weight matrix. The exponential transformation uses an exponential function with a base of 2 to enhance the weight difference of high-sensitivity features. In the above case, the basic weight of the voltage feature in the stationary state is 1.69, and the basic weight of the current feature in the acceleration and deceleration state is 1.82. To adapt to the change of the robot's motion state, a state transition factor is introduced to dynamically adjust the basic weight matrix. The state transition factor has an exponential decay relationship with the state duration and is expressed as a negative exponential function of the state duration. For example, when the state duration reaches 5 seconds, the state transition factor drops to 0.37; when the state lasts for 10 seconds, the state transition factor drops to 0.14. By adjusting the basic weight matrix with the state transition factor, the adaptive weight coefficient is obtained. In actual applications, for a robot that has just entered the uniform motion state for 2 seconds, the adaptive weight coefficient of the current feature is adjusted from the basic weight of 1.65 to 1.52.

[0097] The adaptive weight coefficient is used to perform weighted feature fusion on the parameter feature quantities in different states. In the stationary state, the adaptive weight coefficients of the voltage feature, current feature, power feature, and impedance feature are 1.69, 1.21, 1.35, and 1.42 respectively; in the uniform motion state, the adaptive weight coefficients of the four features are 1.18, 1.57, 1.65, and 1.29 respectively; in the acceleration and deceleration state, the adaptive weight coefficients of the four features are 1.09, 1.82, 1.45, and 1.37 respectively. Each feature vector is multiplied by the corresponding weight coefficient and accumulated to obtain the state-enhanced feature.

[0098] During state transition, the system calculates smooth transition features based on state transition probabilities to ensure the continuity of feature changes. The state transition probability has an exponential relationship with the feature difference between states. For example, during the transition from a stationary state to a uniform velocity state, the feature difference value is 0.68, and the corresponding state transition probability is 0.51. The original state features and target state features are weighted by the state transition probability to obtain smooth transition features. For example, at the initial stage when the robot transitions from a stationary state to a uniform velocity state, the smooth transition feature is calculated as the sum of 0.49 times the stationary state feature and 0.51 times the uniform velocity state feature. As the transition continues, the weights are continuously adjusted until it fully enters the uniform velocity state.

[0099] The present invention introduces a state sensitivity index and an adaptive weight coefficient, enabling the feature extraction process to be dynamically adjusted according to the motion state, greatly improving the adaptability of features to changes in the motion state. In particular, the smooth feature transition mechanism during state transition effectively eliminates the feature mutation caused by state switching, ensuring the stability and continuity of fault identification during state transition.

[0100] In an optional implementation, a deep tree learning classifier including a multi-layer decision tree structure is constructed; the state-sensitive feature set is input into the first-layer decision tree unit, and the decision tree units of adjacent layers perform feature transfer through an attention connection mechanism, where the between-class scatter and within-class scatter of features are calculated based on the Fisher discriminant criterion, an attention weight matrix is constructed, and the transferred features are weighted and combined using the attention weight matrix and an adaptive fusion weight to obtain fused features, including:

[0101] Construct a multi-layer decision tree structure, with each layer containing multiple decision tree units, and establish a feature transfer channel between the decision tree units of adjacent layers;

[0102] Input the state-sensitive feature set into the first-layer decision tree unit, calculate the between-class scatter of the mean of each category feature and the overall feature mean based on the Fisher discriminant criterion, calculate the within-class scatter of the within-class feature and the category feature mean, use the ratio of the between-class scatter to the within-class scatter as the Fisher discriminant ratio, and obtain the feature importance score through exponential normalization;

[0103] Calculate the covariance between features to obtain the correlation coefficient between feature pairs, and perform an exponential weighted operation on the feature importance score and the correlation coefficient to generate an attention weight matrix;

[0104] Introduce a temperature parameter to perform exponential adjustment on the feature importance score to obtain an adaptive fusion weight, perform weighted combination on the current layer features and the attention weight matrix to obtain transferred features, and use the adaptive fusion weight to perform weighted combination on the transferred features to obtain fused features;

[0105] Calculate the Fisher discriminant ratio gain and information gain of the fusion feature relative to the transfer feature, and adaptively adjust the temperature parameter according to the Fisher discriminant ratio gain and information gain. The fusion feature is used as the input feature of the next-level decision tree unit for the classification and judgment of fault types.

[0106] Exemplarily, a multi-layer decision tree structure is constructed as the infrastructure of the deep tree learning classifier. Each layer contains multiple decision tree units, and each decision tree unit consists of a splitting node and a leaf node. The splitting node assigns samples to child nodes according to the feature values, and the leaf node contains the class probability distribution. In this embodiment, a 5-layer deep tree structure is constructed. The first layer contains 1 decision tree unit, the second layer contains 2 decision tree units, the third layer contains 4 decision tree units, the fourth layer contains 8 decision tree units, and the fifth layer contains 16 decision tree units. A feature transfer channel is established between the decision tree units of adjacent layers to form a connected tree-shaped network structure.

[0107] The state-sensitive feature set is input into the first-layer decision tree unit. This feature set contains electrical features and cable bending degree features extracted under different motion states. For example, the dimension of the feature set is 25, including 10-dimensional static state features, 8-dimensional constant velocity state features, and 7-dimensional acceleration and deceleration state features. Calculate the discriminant ability of each feature based on the Fisher discriminant criterion. For each feature in the feature set, calculate the difference between the mean of each class feature and the overall feature mean to obtain the between-class scatter; calculate the difference between the samples within each class and the class feature mean to obtain the within-class scatter. Take the ratio of the between-class scatter to the within-class scatter as the Fisher discriminant ratio. In this embodiment, first calculate the mean of each fault class in each feature dimension, and then calculate the overall mean of all fault classes. The between-class scatter is calculated by the sum of the squares of the differences between each class mean and the overall mean, and the within-class scatter is calculated by the sum of the squares of the differences between the samples within the class and the class mean. The larger the Fisher discriminant ratio value, the stronger the ability of the feature to distinguish fault classes.

[0108] Divide the Fisher discriminant ratio of each feature by the maximum value of the Fisher discriminant ratios of all features, and then perform normalization after exponential function transformation to obtain a feature importance score ranging from 0 to 1. For example, for the 25-dimensional feature, the obtained feature importance scores are distributed between 0.12 and 0.95, where the importance score of the current change rate feature under the acceleration and deceleration state is the highest, reaching 0.95.

[0109] Further calculate the mutual relationship between features. First, calculate the covariance between features, and then divide the covariance by the product of the corresponding feature standard deviations to obtain the correlation coefficient, forming a feature correlation coefficient matrix. For strongly correlated feature pairs, the correlation coefficient is close to 1 or -1; for weakly correlated feature pairs, the correlation coefficient is close to 0. Perform an exponential weighted operation on the feature importance scores and the feature correlation coefficients to generate an attention weight matrix. Specifically, for feature i and feature j, calculate its attention weight coefficient as the result of multiplying the importance score of feature i by the exponentiated correlation coefficient of the two features. In this embodiment, the attention weight between the current-related features and the impedance features is relatively high, and the attention weight between the voltage feature and the power feature is also relatively high.

[0110] Introduce a temperature parameter to perform an exponential adjustment on the feature importance scores to obtain an adaptive fusion weight. The temperature parameter controls the smoothness of the feature importance distribution. A lower temperature parameter makes important features more prominent, while a higher temperature parameter makes the feature importance distribution more uniform. The initial temperature parameter is the mean of the logarithmic values of the feature importance scores, usually between -1.2 and -0.5. The adaptive fusion weight is obtained by performing an exponential operation on the feature importance scores divided by the temperature parameter and then normalizing. In this embodiment, when the temperature parameter is -0.8, the fusion weight of important features reaches above 0.8, while the fusion weight of secondary features is below 0.2.

[0111] For each feature vector of the current layer, multiply it by the attention weight matrix to obtain a weighted feature vector considering the correlation between features, which serves as the transmitted feature. Then, use the adaptive fusion weight to perform a weighted combination on the transmitted feature again to obtain the fused feature. In actual implementation, the dot product of the transmitted feature and the adaptive fusion weight is used as the fused feature, which retains the information of highly discriminative features and suppresses the influence of low discriminative features.

[0112] The Fisher discriminant ratio gain is calculated as the ratio of the Fisher discriminant ratio of the fused feature to the Fisher discriminant ratio of the transmitted feature, and the information gain is calculated as the information entropy of the transmitted feature minus the conditional entropy of the fused feature. The temperature parameter is adaptively adjusted according to the Fisher discriminant ratio gain and the information gain. When the gain is large, lower the temperature parameter to further strengthen the role of important features; when the gain is small, increase the temperature parameter to increase the uniformity of the feature distribution.

[0113] The fused features are used as the input features for the next-layer decision tree unit to classify and judge the fault types. Through this inter-layer feature transfer mechanism, the deep tree learning classifier can maintain and enhance the role of discriminative features in the deep network, while suppressing the interference of redundant and noisy features. For each layer of decision tree unit, a decision tree is constructed according to the input fused features and the optimal splitting criterion. The decision tree unit of the last layer outputs the probability distribution of the fault types, and the category with the highest probability is selected as the final fault classification result.

[0114] Through the combination of the deep tree learning architecture and the attention connection mechanism, the present invention realizes the complementary advantages of traditional decision trees and deep learning; the feature evaluation based on the Fisher discriminant criterion and the adaptive fusion weight adjustment mechanism enable the model to dynamically identify and enhance highly discriminative features, significantly improving the accuracy and robustness of cable fault classification.

[0115] In an optional implementation manner, the adaptive adjustment of the temperature parameter includes:

[0116] The initial value of the temperature parameter is the mean of the logarithmic values of the feature importance scores;

[0117] Calculate the sensitivity of each feature to the temperature parameter, where the sensitivity is calculated by the product of the adaptive fusion weight and the feature importance score and the product of the adaptive fusion weight and its complementary value;

[0118] Perform an exponential transformation and normalization on the sensitivity to obtain a feature enhancement factor, and use the feature enhancement factor to enhance the transfer feature to obtain an enhanced feature;

[0119] Use the adaptive fusion weight to perform a weighted combination on the enhanced feature to obtain an optimized fusion feature, calculate the ratio of the Fisher discriminant ratio of the optimized fusion feature to the transfer feature to obtain a discriminant ratio gain, and calculate the difference in information entropy between the optimized fusion feature and the transfer feature to obtain an information gain;

[0120] Use the weighted sum of the discriminant ratio gain and the information gain as the update basis for the temperature parameter, and update the temperature parameter using the gradient descent method;

[0121] Calculate the change rate of the adaptive fusion weight in two adjacent rounds and the maximum value of the discriminant ratio gain as the convergence criterion, and determine the final temperature parameter when the convergence criterion is less than a preset convergence threshold.

[0122] Exemplarily, the initialization of the temperature parameter is the first step in feature fusion optimization. To ensure that the initial temperature parameter can reasonably reflect the overall situation of the feature importance distribution, the natural logarithm of the importance score of each feature is taken, and then the arithmetic mean of all the logarithmic values is calculated to obtain the initial temperature parameter. In practical applications, for a 25-dimensional state-sensitive feature set, assuming that the feature importance scores are distributed between 0.1 and 0.9, the initial temperature parameter is usually between -1.2 and -0.5.

[0123] Calculate the sensitivity of each feature to the temperature parameter. This step aims to quantify the impact of temperature parameter changes on the feature fusion weights. For feature i, its sensitivity is the product of the adaptive fusion weight of the feature, the feature importance score, and the complement of the adaptive fusion weight of the feature (i.e., 1 minus the adaptive fusion weight). When the adaptive fusion weight is close to 0 or 1, the sensitivity is low; when the adaptive fusion weight is close to 0.5, the sensitivity is high. In this embodiment, the features with the highest sensitivity often appear near the medium importance scores, and these features are the most sensitive to changes in the temperature parameter.

[0124] Perform an exponential transformation on the sensitivity. The exponential transformation is achieved by multiplying the sensitivity by a magnification factor and then performing an exponential operation. The magnification factor is set to 2.0 in this embodiment. The normalization process ensures that the sum of all feature enhancement factors is 1, facilitating subsequent feature enhancement operations. The feature enhancement factor reflects the dynamic adjustment amplitude of each feature in the fusion process, and features with high sensitivity obtain larger enhancement factors.

[0125] Multiply each dimension of the transmitted feature by the corresponding feature enhancement factor and then add the original transmitted feature to achieve differential enhancement. While retaining the original feature information, the enhanced feature strengthens the feature dimensions that are sensitive to changes in the temperature parameter. In practical applications, for key features such as the current change rate, the enhancement amplitude can reach 1.5 times the original value, while for less important features, the enhancement amplitude may only be 1.1 times the original value.

[0126] Obtain the optimized fusion feature vector through the dot product operation of the enhanced feature and the adaptive fusion weight.

[0127] Calculate the performance differences between the optimized fusion features and the transferred features, including the ratio of Fisher discriminant ratios and the difference in information entropy. The ratio of Fisher discriminant ratios is called the discriminant ratio gain, and is calculated by dividing the Fisher discriminant ratio of the optimized fusion features by the Fisher discriminant ratio of the transferred features. The Fisher discriminant ratio is obtained by calculating the ratio of between-class scatter to within-class scatter. The information gain is obtained by calculating the information entropy of the transferred features minus the conditional entropy of the optimized fusion features. The information entropy is calculated based on the probability distribution of samples of each class in the feature space. In this embodiment, effective feature fusion optimization usually produces a discriminant ratio gain of more than 1.2 and an information gain of more than 0.15.

[0128] Use the weighted sum of the discriminant ratio gain and the information gain as the basis for updating the temperature parameter. In this embodiment, the discriminant ratio gain weight is set to 0.7 and the information gain weight is set to 0.3. Use the gradient descent method to update the temperature parameter. The specific update formula is the current temperature parameter minus the learning rate multiplied by the weighted gain. The learning rate is initially set to 0.1 and gradually decreases as the number of iterations increases. When the weighted gain is positive, decrease the temperature parameter to strengthen the role of important features; when the weighted gain is negative, increase the temperature parameter to increase the uniformity of the feature distribution.

[0129] To determine whether the optimization of the temperature parameter reaches the convergence state, calculate the change rate of the adaptive fusion weights between two adjacent rounds and the maximum value of the discriminant ratio gain as the convergence criterion. The change rate of the adaptive fusion weights is obtained by calculating the maximum absolute value of the differences in all feature fusion weights between two rounds of iteration. When the change rate is less than the preset convergence threshold (set to 0.01 in this embodiment) and the maximum value of the discriminant ratio gain does not change by more than 0.5% for three consecutive rounds, it is considered that the optimization process converges, and the temperature parameter at this time is determined as the final temperature parameter. In practical applications, the optimization of the temperature parameter usually converges within 10 to 20 rounds of iteration.

[0130] After the temperature parameter is determined, the adaptive fusion weights calculated based on this parameter are used for the final feature fusion process. By applying the final adaptive fusion weights to the transferred features, fusion features with optimal discriminant ability are obtained. These fusion features are used as the input for the next layer of the decision tree for accurate classification of fault types.

[0131] Existing technologies mostly adopt fixed-window feature extraction and uniform feature weight assignment strategies, lacking the ability to adaptively perceive signal features at different time scales. At the same time, feature transmission in traditional multi-layer decision tree structures is usually achieved through simple concatenation or average fusion, unable to effectively distinguish feature importance and capture complex correlation relationships between features. Starting from the problem of differentiating cable fault features in different motion states of the robot, the present invention introduces an attention connection mechanism based on the Fisher discriminant criterion and an adaptive temperature parameter adjustment mechanism. By calculating the between-class scatter and within-class scatter, the accurate quantification of feature discrimination ability is realized; the adaptive adjustment of temperature parameters transforms the feature fusion process into a dynamic learning task, and the system can automatically adjust the feature weight distribution according to the discrimination ratio gain and information gain. The improvement also includes sensitivity-driven feature enhancement and multi-index convergence criteria, forming a feature optimization system with closed-loop feedback. The present invention significantly improves the accuracy and robustness of robot cable fault classification; the combination of the Fisher discriminant criterion and the attention mechanism enables the model to automatically focus on highly discriminative features, reducing the interference of noise features; the adaptive adjustment mechanism of temperature parameters realizes the dynamic enhancement of feature importance, effectively coping with the fault feature changes in different motion states.

[0132] In an alternative embodiment, training the deep tree learning classifier includes:

[0133] Generating a multi-scale time window sequence, the multi-scale time window sequence including a base window and extended windows obtained by multiplying the base window, and dividing training samples based on the base window and the extended windows to obtain multiple window sample sets;

[0134] Calculating the probability distribution of samples of each category in each window sample set, calculating the window information entropy based on the probability distribution, calculating the conditional information entropy for each candidate splitting feature, and taking the difference between the window information entropy and the conditional information entropy as the window information gain;

[0135] Calculating the time decay weight based on the time difference between the current moment and the central moment of each window, normalizing the time decay weight to obtain the window weight, and weighting the information gain of each window with the window weight to obtain the comprehensive information gain;

[0136] Calculating the splitting information entropy of each candidate splitting feature, taking the ratio of the comprehensive information gain to the splitting information entropy as the splitting gain ratio, and selecting the candidate splitting feature with the largest splitting gain ratio as the optimal splitting feature;

[0137] Calculating the difference degree of corresponding features between adjacent layer decision tree units, taking the weighted sum of the difference degrees as the regularization constraint term, and taking the cumulative sum of the splitting gain ratios of each layer decision tree unit minus the regularization constraint term as the optimization objective function;

[0138] Calculate the classification performance metrics of the depth tree learning classifier on the validation set, adaptively adjust the time decay coefficient, regularization coefficient, and the basic window length based on the classification performance metrics, and repeat the feature splitting optimization process until a preset convergence criterion is reached.

[0139] Exemplarily, generate a multi-scale time window sequence as the basis for capturing fault features at different time scales. The multi-scale time window sequence consists of a basic window and an extended window. The length of the basic window is determined according to the minimum characteristic period of the cable fault signal, which is set to 10 milliseconds in this embodiment and can capture the high-frequency change features in the cable signal. The extended window is obtained by multiplying the basic window by a multiple to form a complete multi-scale window sequence. Specifically, construct 5 time windows, namely 10 milliseconds, 20 milliseconds, 40 milliseconds, 80 milliseconds, and 160 milliseconds in sequence, covering various cable fault features from fast transient changes to slow trend changes. Divide the training samples based on these windows to obtain multiple window sample sets. For each sample point, take the moment it is located as the center and extract the corresponding time period data according to the length of each window to form the corresponding window sample. For example, for the sample point at time t, the basic window sample contains the data from t - 5 milliseconds to t + 5 milliseconds, while the 160-millisecond window sample contains the data from t - 80 milliseconds to t + 80 milliseconds.

[0140] Perform class statistics on the samples in each window sample set and calculate the probability distribution of samples in each class. In this embodiment, the robot cable faults are divided into 5 categories: normal state, insulation aging, poor contact, short circuit, and open circuit. For the 10-millisecond window sample set, the statistics show that the normal state samples account for 35%, the insulation aging samples account for 20%, the poor contact samples account for 25%, the short circuit samples account for 12%, and the open circuit samples account for 8%. The window information entropy is represented by the negative value of the sum of the product of the probability of each class and its logarithm, reflecting the uncertainty of the sample set. For example, the information entropy of the 10-millisecond window is 2.14, while the information entropy of the 160-millisecond window is 1.92, indicating that the sample distribution is more concentrated under the long-time window. Calculate the conditional information entropy for each candidate splitting feature, and the conditional information entropy is calculated by the weighted average of the subset information entropy when the feature takes different values. Take the difference between the window information entropy and the conditional information entropy as the window information gain, reflecting the contribution of the candidate feature to sample classification. In the 10-millisecond window, the information gain of the current change rate feature is 0.75, while in the 160-millisecond window, the information gain of the impedance amplitude feature is 0.68, indicating that the effective features are different at different time scales.

[0141] Calculate the time decay weight based on the time difference between the current moment and the central moment of each window, which reflects the influence of time correlation on the importance of features. The time decay weight is calculated by the exponential function of the time difference. The smaller the time difference, the larger the weight, indicating that the recent data is more important for the current decision. The time decay coefficient is initially set to 0.05 and is dynamically adjusted during the training process. Normalize the time decay weight to obtain the window weight so that the sum of all window weights is 1. In practical applications, the window weight of the current moment is approximately 0.4, the window weight of the adjacent moment is approximately 0.25, and the window weights of the earlier moments gradually decrease. Use the window weight to weight the information gain of each window to obtain the comprehensive information gain, making full use of the multi-scale window information. For the current change rate feature, its comprehensive information gain is 0.72; for the impedance amplitude feature, its comprehensive information gain is 0.65.

[0142] Calculate the split information entropy of each candidate split feature. The split information entropy reflects the balance degree of the feature in splitting the sample set and is represented by the negative value of the sum of the products of the subset sample ratio of different values of the feature and its logarithm. The smaller the split information entropy, the more unbalanced the feature split. Take the ratio of the comprehensive information gain to the split information entropy as the split gain ratio, comprehensively considering the information contribution and split balance of the feature. In this embodiment, the split information entropy of the current change rate feature is 1.25, and the split gain ratio is 0.576; the split information entropy of the impedance amplitude feature is 1.42, and the split gain ratio is 0.458. Select the candidate split feature with the largest split gain ratio as the optimal split feature for the construction of the decision tree.

[0143] Calculate the difference degree of the corresponding features between adjacent-layer decision tree units to achieve the global optimization of the deep tree structure. Subtract the regularization constraint term from the cumulative sum of the split gain ratios of each layer of decision tree units to obtain the optimization objective function. The optimization objective function is used as an evaluation index when updating the decision tree structure and the feature mapping relationship. The larger the value, the better the model performance.

[0144] Calculate the classification performance metrics of the deep tree learning classifier on the validation set, including accuracy, precision, recall, and F1-score. On the validation set, the initial accuracy of the model is 87.5%, the precision is 86.2%, the recall is 87.8%, and the F1-score is 87.0%. Adaptive adjustment is made to the time decay coefficient, regularization coefficient, and base window length based on these performance metrics. When the increase in accuracy is less than 1%, increase the time decay coefficient to strengthen the influence of recent data; when the precision is lower than 85%, increase the regularization coefficient to improve the generalization ability of the model; when the difference in recall between different classes is greater than 15%, adjust the base window length to balance the performance of each class. In this embodiment, the time decay coefficient is adjusted from the initial 0.05 to 0.08, the regularization coefficient is adjusted from 0.1 to 0.15, and the base window length is adjusted from 10 milliseconds to 12 milliseconds. Repeat the feature splitting optimization process until the preset convergence criterion is reached. The convergence criterion is set as the increase in accuracy on the validation set being less than 0.5% for three consecutive rounds or reaching the maximum number of iterations of 30 rounds.

[0145] The deep tree learning classifier of the present invention fully learns the temporal patterns of cable fault features under multi-scale time windows and establishes a hierarchical representation from low-level basic features to high-level abstract features; each layer of decision tree units effectively divides samples through optimal splitting features, and the feature transfer and regularization constraints between adjacent layers ensure the coherence and consistency of feature representations; the finally trained classifier can accurately identify various cable faults and provide reliable support for the cable health status monitoring of robots.

[0146] The present invention realizes the feature extraction and optimization of cable fault signals at different time scales through a multi-scale time window sequence and a time decay weight mechanism, effectively capturing the full-spectrum fault features from transient changes to long-term trends. The splitting criterion based on information gain and split information entropy takes into account the balance of feature segmentation while ensuring the classification performance of the decision tree, improving the model's ability to identify minority-class faults; the introduction of regularization constraints and a global optimization objective function solves the feature consistency problem in the deep decision tree structure, significantly enhancing the generalization ability and anti-noise ability of the model, and providing an efficient and reliable technical support for the real-time monitoring and preventive maintenance of robot cable faults.

[0147] In an optional implementation manner, calculate the difference degree of corresponding features between adjacent layers of decision tree units, use the weighted sum of the difference degrees as a regularization constraint term, and use the cumulative sum of the split gain ratios of each layer of decision tree units minus the regularization constraint term as an optimization objective function, including:

[0148] Calculate the feature mapping matrix between adjacent layers of decision tree units, and the feature mapping matrix is used to establish the conversion relationship between the feature vectors of the decision tree units of the i-th layer and the (i + 1)-th layer;

[0149] Calculate the Euclidean distance, cosine similarity, and information entropy difference between the feature vectors of the decision tree units in the i-th layer and the feature vectors of the decision tree units in the (i + 1)-th layer after being transformed by the feature mapping matrix, and perform weighted combination to obtain a feature difference vector;

[0150] Perform weighted average on the feature difference vector and the splitting gain ratio of the corresponding feature to obtain a unit difference degree, and set a hierarchical weight based on the depth of the decision tree unit in the classification path, where the hierarchical weight decays exponentially as the depth increases;

[0151] Perform weighted summation on the unit difference degrees between all adjacent-layer decision tree units and the corresponding hierarchical weights to obtain a global difference degree, and multiply the global difference degree by a regularization coefficient to obtain a regularization constraint term;

[0152] Optimize the feature mapping matrix and the decision tree structure based on the optimization objective function.

[0153] Exemplarily, in combination with Figure 2 The flowchart of feature mapping and regularization optimization of a deep tree learning classifier is described as follows: Calculate the feature mapping matrix between adjacent-layer decision tree units, which is used to establish the transformation relationship between the feature vectors of the decision tree units in the i-th layer and the (i + 1)-th layer. The feature mapping matrix is obtained through backpropagation training, indicating the transformation method of features during the inter-layer transmission process. Specifically, for each pair of adjacent-layer decision tree units, collect their input feature vectors and output feature vectors to construct a training sample set. Use the gradient descent method to minimize the difference between the input features after being transformed by the mapping matrix and the actual output features, and train to obtain the feature mapping matrix. In this embodiment, for the feature mapping between the first layer and the second layer, 5000 pairs of feature vector samples are used for training, the learning rate is set to 0.01, and the number of iterations is 200 times. The dimension of the obtained feature mapping matrix matches the feature dimensions of the two layers. For example, if the feature dimension of the first layer is 25 and the feature dimension of the second layer is 20, then the dimension of the mapping matrix is 25×20. The values in the mapping matrix reflect the correlation strength between different features, and the numerical range is usually between -1 and 1. For example, the mapping coefficient between the current feature and the current-related feature in the next layer is 0.85, while the mapping coefficient with the unrelated feature is close to 0.

[0154] Calculate the difference between the feature vectors of the decision tree units in the i-th layer and the feature vectors of the decision tree units in the (i + 1)-th layer after being transformed by the feature mapping matrix. Three different measurement methods are used for the difference calculation: Euclidean distance, cosine similarity, and information entropy difference. The Euclidean distance reflects the absolute distance of the feature vectors in space and is obtained by calculating the square root of the sum of the squares of the differences of each dimension of the two vectors. The cosine similarity represents the degree of closeness of the directions of the two vectors, is calculated by dividing the dot product of the two vectors by the product of their respective magnitudes, and the value range is between -1 and 1. The closer the value is to 1, the more similar the directions are. The information entropy difference measures the difference in the uncertainty of the two feature distributions and is obtained by calculating the absolute value of the difference between the information entropies of the probability distributions of the two feature vectors. In practical applications, the Euclidean distance, cosine similarity, and information entropy difference are weighted and combined into a feature difference vector according to the ratio of 6:3:1. For example, if the Euclidean distance of a pair of adjacent layer features is 0.42, the cosine similarity is 0.78 (the difference value is 0.22), and the information entropy difference is 0.15, then the weighted combined feature difference value is 0.42×0.6 + 0.22×0.3 + 0.15×0.1 = 0.348.

[0155] The split gain ratio reflects the importance of the feature in the decision tree split. Feature differences with high importance should be given higher weights. In this embodiment, for each dimension value of the feature difference vector, multiply it by the normalized value of the split gain ratio of the corresponding feature as the weight coefficient, and then calculate the weighted average to obtain the unit difference degree. For example, the current change rate feature in the state-sensitive feature set has a relatively high split gain ratio of 0.58, and its corresponding feature difference value is 0.31. Then, when calculating the unit difference degree, the weighted contribution of this feature is 0.31×0.58÷total weight = 0.18.

[0156] The purpose of setting the layer weights is to make the features in the shallow layers maintain higher consistency, while the features in the deep layers have greater flexibility to adapt to complex classification boundaries. In a five-layer deep tree structure, the layer weights of each layer are set as follows: the first layer is 1.0, the second layer is 0.7, the third layer is 0.49, the fourth layer is 0.343, and the fifth layer is 0.24. This exponentially decaying weight setting makes the feature differences between adjacent shallow layers account for a larger proportion in the regularization constraint.

[0157] The global dissimilarity comprehensively reflects the degree of change in feature representations throughout the deep tree structure. For all pairs of decision tree units in each pair of adjacent layers, their unit dissimilarities are calculated, summed after multiplying by the corresponding layer weights, and the global dissimilarity is obtained. For example, if the sum of the weighted unit dissimilarities between the first and second layers is 0.41, between the second and third layers is 0.38, between the third and fourth layers is 0.35, and between the fourth and fifth layers is 0.32, then the global dissimilarity is 0.41×1.0 + 0.38×0.7 + 0.35×0.49 + 0.32×0.343 = 0.997. The global dissimilarity is multiplied by the regularization coefficient to obtain the regularization constraint term. The regularization coefficient controls the strength of the feature consistency constraint, with an initial value set to 0.15 and dynamically adjusted during the training process. The regularization constraint term is used to penalize excessive changes in feature representations between adjacent layers, prompting the model to learn more coherent feature representations.

[0158] Subtract the regularization constraint term from the cumulative sum of the splitting gain ratios of the decision tree units in each layer to form the optimization objective function. The cumulative sum of the splitting gain ratios reflects the local classification performance of each decision tree unit. In this embodiment, the cumulative sum of the splitting gain ratios for the five-layer structure is 2.85. The value of the optimization objective function after subtracting the regularization constraint term is 2.70, and this value is used as an indicator to evaluate the overall performance of the model. The larger the value, the better the classification performance of the model while maintaining feature consistency.

[0159] Optimize the feature mapping matrix and the decision tree structure based on the optimization objective function. The optimization process adopts an alternating optimization strategy: first, fix the decision tree structure and optimize the feature mapping matrix by the gradient ascent method to maximize the objective function value; then, fix the feature mapping matrix and adjust the splitting features and splitting points in the decision tree structure to further improve the objective function value. After each round of optimization, recalculate the regularization constraint term and the value of the optimization objective function until the convergence condition or the maximum number of iterations is reached. In this embodiment, the optimization process usually converges within 15 rounds, and the final value of the optimization objective function increases by approximately 12% to reach 3.03.

[0160] The present invention improves the problems of the lack of an effective inter-layer feature transfer mechanism and easy overfitting in the existing deep integrated tree model. From the perspectives of feature mapping relationship and inter-layer feature consistency, the present invention proposes a multi-dimensional feature difference calculation and a hierarchical decay weight mechanism, controls the feature changes in the deep tree structure through regularization constraints, and realizes the feature abstraction ability similar to deep learning while maintaining the interpretability of the decision tree. The present invention establishes a clear feature conversion relationship between adjacent layer decision tree units, realizing effective feature transfer in the deep tree structure; the multi-dimensional feature difference calculation mechanism comprehensively considers the distance, direction, and distribution characteristics of the feature vector, comprehensively evaluating the changes in the inter-layer feature representation; the exponential decay hierarchical weight strategy based on the classification path depth ensures the stability of shallow features and the flexibility of deep features; the balance mechanism of regularization constraints and splitting gain improves the local classification performance while maintaining the coherence of the global feature representation, significantly enhancing the recognition ability of the deep tree learning classifier for robot cable faults, especially performing well in complex working conditions and multi-category fault scenarios, providing high-precision and high-reliability technical support for the preventive maintenance of industrial robots.

[0161] In the second aspect of the embodiments of the present invention, a robot cable fault classification system based on deep tree learning is provided, including:

[0162] The first unit is used to collect the voltage signal and current signal of the robot cable, obtain the joint position information of the robot, and calculate the cable bending degree feature based on the joint position information;

[0163] The second unit is used to divide the robot motion state into a stationary state, a uniform motion state, and an acceleration / deceleration state, extract voltage, current, power, and impedance features for each state, and combine them with the cable bending degree feature to form a state-sensitive feature set;

[0164] The third unit is used to construct a deep tree learning classifier including a multi-layer decision tree structure; input the state-sensitive feature set into the first-layer decision tree unit, and the decision tree units of adjacent layers perform feature transfer through an attention connection mechanism, where the between-class scatter and within-class scatter of the features are calculated based on the Fisher discriminant criterion, an attention weight matrix is constructed, and the transferred features are weighted and combined using the attention weight matrix and an adaptive fusion weight to obtain a fusion feature;

[0165] The fourth unit is used to train the deep tree learning classifier, including: locally optimizing the decision tree unit using a multi-scale splitting criterion, where the multi-scale splitting criterion calculates the information gain and dynamically adjusts the window weight under different time windows; performing global optimization, introducing regularization constraints to control the feature differences between adjacent layer decision tree units;

[0166] The fifth unit is configured to input the robot cable signals to be classified into the trained deep tree learning classifier to obtain the fault types.

[0167] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including:

[0168] A processor;

[0169] A memory for storing instructions executable by the processor;

[0170] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0171] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0172] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A robot cable fault classification method based on deep tree learning, characterized by: Including: Collect the voltage signal and current signal of the robot cable, obtain the joint position information of the robot, and calculate the cable bending degree feature based on the joint position information; Divide the robot motion state into a stationary state, a uniform motion state, and an acceleration / deceleration state. Extract voltage, current, power, and impedance features for each state, and combine them with the cable bending degree feature to form a state-sensitive feature set; Construct a deep tree learning classifier containing a multi-layer decision tree structure; input the state-sensitive feature set into the first-layer decision tree unit, and the decision tree units of adjacent layers perform feature transfer through an attention connection mechanism. Among them, calculate the between-class scatter and within-class scatter of the features based on the Fisher discriminant criterion, construct an attention weight matrix, and use the attention weight matrix and the adaptive fusion weight to perform weighted combination on the transferred features to obtain the fusion features; Train the deep tree learning classifier, including: locally optimize the decision tree unit using a multi-scale splitting criterion, and the multi-scale splitting criterion calculates the information gain and dynamically adjusts the window weight under different time windows; perform global optimization, and introduce regularization constraints to control the feature differences between the decision tree units of adjacent layers; Input the robot cable signal to be classified into the trained deep tree learning classifier to obtain the fault type.

2. The method according to claim 1, wherein The formation of the state-sensitive feature set includes: Based on the joint angular velocity vector, joint angular acceleration vector, and end effector motion parameters of the robot, divide the robot motion state into a stationary state, a uniform motion state, and an acceleration / deceleration state; Extract electrical features for each motion state respectively, including: calculate the mean, effective value, standard deviation, peak value, and valley value of the voltage signal to obtain the voltage feature vector, calculate the mean, effective value, standard deviation, peak value, valley value, and maximum change rate of the current signal to obtain the current feature vector, calculate the active power, reactive power, and power factor to obtain the power feature vector, calculate the impedance amplitude and phase to obtain the impedance feature vector; perform weighted combination on the voltage feature vector, current feature vector, power feature vector, and impedance feature vector by assigning weight coefficients corresponding to the motion state to obtain the state-enhanced feature; Calculate the local bending angle and cumulative bending amount of the cable based on the cable bending degree feature and perform feature fusion with the state-enhanced feature to obtain the state perception feature; perform feature selection on the state perception feature based on the mutual information evaluation method, and select the features with mutual information values greater than the feature screening threshold to form the state-sensitive feature set.

3. The method according to claim 2, characterized in that Performing weighted combination on the voltage feature vector, current feature vector, power feature vector, and impedance feature vector by assigning weight coefficients corresponding to the motion state to obtain the state-enhanced feature includes: Obtain parameter feature quantities, and the parameter feature quantities include voltage parameter features, current parameter features, power parameter features, and impedance parameter features; Calculate the correlation degree between the robot motion state and each parameter feature, including: calculate the Pearson correlation coefficient to obtain the feature-state correlation coefficient, calculate the mutual information to obtain the feature-state mutual information value, and perform weighting on the feature-state correlation coefficient and the feature-state mutual information value to obtain the state sensitivity index; Normalize the state sensitivity index to obtain the normalized sensitivity. Perform an exponential transformation on the normalized sensitivity to obtain the basic weight matrix. Dynamically adjust the basic weight matrix through a state transition factor that has an exponential decay relationship with the state duration to obtain the adaptive weight coefficient. Use the adaptive weight coefficient to perform weighted feature fusion on the parameter feature quantities in the stationary state, uniform motion state, and acceleration / deceleration state respectively to obtain the state enhancement feature. Among them, calculate the smooth transition feature based on the state transition probability during the state transition, and the state transition probability has an exponential relationship with the feature difference between states.

4. The method according to claim 1, wherein Construct a deep tree learning classifier that includes a multi-layer decision tree structure. Input the state-sensitive feature set into the first-layer decision tree unit. The decision tree units of adjacent layers perform feature transfer through an attention connection mechanism. Among them, calculate the between-class scatter and within-class scatter of the features based on the Fisher discriminant criterion, construct an attention weight matrix, and use the attention weight matrix and the adaptive fusion weight to perform weighted combination on the transferred features to obtain the fusion features, including: Construct a multi-layer decision tree structure, with each layer containing multiple decision tree units, and establish a feature transfer channel between the decision tree units of adjacent layers. Input the state-sensitive feature set into the first-layer decision tree unit. Calculate the between-class scatter of the mean of each category feature and the overall feature mean based on the Fisher discriminant criterion, calculate the within-class scatter of the within-class feature and the category feature mean, take the ratio of the between-class scatter to the within-class scatter as the Fisher discriminant ratio, and obtain the feature importance score through exponential normalization. Calculate the covariance between features and obtain the correlation coefficient between feature pairs. Perform an exponential weighted operation on the feature importance score and the correlation coefficient to generate the attention weight matrix. Introduce a temperature parameter to perform exponential adjustment on the feature importance score to obtain the adaptive fusion weight. Perform a weighted combination of the current layer features and the attention weight matrix to obtain the transferred features, and use the adaptive fusion weight to perform a weighted combination on the transferred features to obtain the fusion features. Calculate the Fisher discriminant ratio gain and information gain of the fusion features relative to the transferred features, and adaptively adjust the temperature parameter according to the Fisher discriminant ratio gain and information gain. The fusion features are used as the input features of the next-layer decision tree unit for the classification and judgment of the fault type.

5. The method according to claim 4, characterized in that, The adaptive adjustment of the temperature parameter includes: The initial value of the temperature parameter is the mean of the logarithm values of each feature importance score. Calculate the sensitivity of each feature to the temperature parameter, and the sensitivity is calculated through the product of the adaptive fusion weight and the feature importance score and the product of the adaptive fusion weight and its complementary value. Perform an exponential transformation and normalization on the sensitivity to obtain the feature enhancement factor, and use the feature enhancement factor to enhance the transferred features to obtain the enhanced features. The enhanced features are weighted and combined using the adaptive fusion weights to obtain optimized fusion features. The ratio of the Fisher discriminant ratio of the optimized fusion features to the transfer features is calculated to obtain the discriminant ratio gain, and the difference between the information entropy of the optimized fusion features and the transfer features is calculated to obtain the information gain; The weighted sum of the discriminant ratio gain and the information gain is used as the update basis for the temperature parameter, and the temperature parameter is updated using the gradient descent method; The change rate of the adaptive fusion weights between two adjacent rounds and the maximum value of the discriminant ratio gain are calculated as the convergence criterion. When the convergence criterion is less than the preset convergence threshold, the final temperature parameter is determined.

6. The method according to claim 1, characterized in that Training the deep tree learning classifier includes: Generating a multi-scale time window sequence, where the multi-scale time window sequence includes a base window and extended windows obtained by multiplying the base window. Based on the base window and the extended windows, the training samples are divided to obtain multiple window sample sets; Calculating the probability distribution of samples of each category in each window sample set, calculating the window information entropy based on the probability distribution, calculating the conditional information entropy for each candidate splitting feature, and taking the difference between the window information entropy and the conditional information entropy as the window information gain; Calculating the time decay weight based on the time difference between the current moment and the central moment of each window, normalizing the time decay weight to obtain the window weight, and weighting the information gain of each window using the window weight to obtain the comprehensive information gain; Calculating the splitting information entropy of each candidate splitting feature, taking the ratio of the comprehensive information gain to the splitting information entropy as the splitting gain ratio, and selecting the candidate splitting feature with the largest splitting gain ratio as the optimal splitting feature; Calculating the difference degree of corresponding features between adjacent layer decision tree units, taking the weighted sum of the difference degrees as the regularization constraint term, and subtracting the regularization constraint term from the cumulative sum of the splitting gain ratios of each layer of decision tree units as the optimization objective function; Calculating the classification performance metrics of the deep tree learning classifier on the validation set, adaptively adjusting the time decay coefficient, regularization coefficient, and base window length based on the classification performance metrics, and repeating the feature splitting optimization process until the preset convergence standard is reached.

7. The method according to claim 6, characterized in that Calculating the difference degree of corresponding features between adjacent layer decision tree units, taking the weighted sum of the difference degrees as the regularization constraint term, and subtracting the regularization constraint term from the cumulative sum of the splitting gain ratios of each layer of decision tree units as the optimization objective function includes: Calculating the feature mapping matrix between adjacent layer decision tree units, where the feature mapping matrix is used to establish the conversion relationship between the feature vectors of the decision tree units in the i-th layer and the (i + 1)-th layer; Calculating the Euclidean distance, cosine similarity, and information entropy difference between the feature vector of the decision tree unit in the i-th layer and the feature vector of the decision tree unit in the (i + 1)-th layer after being transformed by the feature mapping matrix, and weighting and combining them to obtain the feature difference vector; Weighted averaging the feature difference vector and the splitting gain ratio of the corresponding feature to obtain the unit difference degree, and setting the hierarchical weight based on the depth of the decision tree unit in the classification path, where the hierarchical weight decays exponentially with the increase in depth; The unit difference degrees between all adjacent-layer decision tree units are weighted and summed with the corresponding layer weights to obtain a global difference degree, and the global difference degree is multiplied by a regularization coefficient to obtain a regularization constraint term; Based on the optimization objective function, the feature mapping matrix and the decision tree structure are optimized.

8. A robot cable fault classification system based on deep tree learning for implementing the method according to any one of the preceding claims 1-7, characterized in that, Including: The first unit is configured to collect the voltage signal and current signal of the robot cable, obtain the joint position information of the robot, and calculate the cable bending degree feature based on the joint position information; The second unit is configured to divide the robot motion state into a stationary state, a uniform motion state, and an acceleration / deceleration state, extract voltage, current, power, and impedance features for each state, and combine them with the cable bending degree feature to form a state-sensitive feature set; The third unit is configured to construct a deep tree learning classifier including a multi-layer decision tree structure; input the state-sensitive feature set into the first-layer decision tree unit, and the decision tree units of adjacent layers perform feature transfer through an attention connection mechanism, where the between-class scatter and within-class scatter of the features are calculated based on the Fisher discriminant criterion, an attention weight matrix is constructed, and the transferred features are weighted and combined using the attention weight matrix and an adaptive fusion weight to obtain a fusion feature; The fourth unit is configured to train the deep tree learning classifier, including: locally optimizing the decision tree unit using a multi-scale splitting criterion, where the multi-scale splitting criterion calculates the information gain and dynamically adjusts the window weight under different time windows; performing global optimization by introducing a regularization constraint to control the feature difference between adjacent-layer decision tree units; The fifth unit is configured to input the robot cable signal to be classified into the trained deep tree learning classifier to obtain the fault type.

9. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Cable fault edge diagnosis method fusing parallel network and attention mechanism

    CN119226770A

  • Intelligent fault diagnosis method for wrapping robot

    CN119418487A