Artificial Intelligence-Based Quality Control and Fault Prediction Method for Surgical Robots

Through a neural network based on dual-path feature separation and enhancement, the problems of high-frequency noise interference, transient impact feature loss and frequency band time-degeneration of surgical robot fault detection are solved, and accurate fault prediction and model generalization capabilities are improved.

CN120183644BActive Publication Date: 2025-08-01SHANDONG PROVINCIAL HOSPITAL AFFILIATED TO SHANDONG FIRST MEDICAL UNIVERSITY (SHANDONG PROVINCIAL HOSPITAL)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510652858.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-01
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the transient impact characteristics of surgical robot failures, and the model generalization ability is insufficient to adapt to frequency drift and high-frequency noise interference.

Method used

A neural network based on dual-path feature separation and enhancement is adopted, and a convolutional neural network is constructed through wavelet decomposition and dynamic normalization of frequency band energy. The time sequence feature enhancement module, dynamic sparse self-attention mechanism and residual adaptive frequency modulation module are used to extract and fusion, and a fault-aware loss function and an adaptive gradient clipping optimization model are used.

Benefits of technology

Accurately capture the fault characteristics of surgical robots, improve the accuracy of fault prediction and the generalization ability of the model, suppress high-frequency noise, adaptively track time-varying frequency band faults, and improve training stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183644B_ABST
    Figure CN120183644B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of artificial intelligence and data processing, and specifically relates to a method for quality control and fault prediction of a surgical robot based on artificial intelligence, which is as follows: collecting data of the surgical robot through sensors, manually annotating the collected data to form a sample data set; performing multi-scale normalization processing on the sample data set; constructing a fault prediction model of the surgical robot by using a convolutional neural network, inputting the normalized data into the model for training to obtain a trained fault prediction model; inputting the newly collected and preprocessed data into the trained fault prediction model, and the fault prediction model outputs the fault category prediction result corresponding to the input data; constructing a three-dimensional quality control matrix, and the three-dimensional quality control matrix triggers a corresponding automatic calibration program based on the category prediction result output by the fault prediction model. The present invention can accurately capture the fault characteristics of the surgical robot, improve the accuracy of fault prediction and the generalization ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and data processing, and particularly relates to a method for quality control and fault prediction of a surgical robot based on artificial intelligence. Background Art

[0002] With the wide development of modern minimally invasive surgeries, surgical robots, as high-precision and highly flexible operation platforms, have been widely used in multiple fields such as neurosurgery and urology. Their complex joint structures and multi-sensor control systems enable them to have high controllability and repeatability when performing precise tasks. However, due to their high system integration and fast operation rhythm, any minor abnormality (such as mechanical jamming, gear meshing error, sensor drift) may have an inestimable impact on surgical safety. Especially during the surgical process, some faults are manifested as instantaneous shocks, frequency drift, or hysteresis phenomena, which are characterized by suddenness, locality, and non-stationarity. Traditional quality control and fault detection methods are difficult to capture and accurately identify such complex signal features in a timely manner.

[0003] The Chinese invention patent with the publication number CN118938882A proposes an intelligent fault monitoring and diagnosis system and method for an industrial robot, deploying target sensors at the monitoring parts of the target industrial robot to obtain real-time data; determining and analyzing the diagnostic faults based on machine learning technology and the real-time data, and outputting a diagnostic report, which can specifically determine the monitoring parts to collect fault detection data and improve the efficiency of subsequent fault detection.

[0004] The Chinese invention patent with the publication number CN114897102A proposes an industrial robot fault diagnosis method, system, device, and storage medium, mainly using a twin neural network structure to achieve automatic feature extraction and fault recognition.

[0005] The Chinese invention patent with the publication number CN115099268A proposes a method and system for intelligent fault diagnosis of a wheeled robot based on a graph convolutional network. The spatio-temporal difference graph convolutional network is used to calculate multi-order backward difference features of the wheeled robot data relationship graph, the local difference characteristics are used to enhance the features of the nodes, and the spatio-temporal graph convolutional module is used to obtain spatio-temporal related features. The constructed robot data relationship graph is beneficial to fault classification, and the developed STDGCN has the most advanced performance.

[0006] The objective drawbacks of the prior art are as follows: Conventional time-domain normalization methods fail to consider the frequency-band energy differences, amplify high-frequency noise, compress low-frequency features, and are prone to causing diagnostic failures; Conventional convolutional networks can only extract steady-state trends and cannot effectively capture transient impact features, and are prone to misjudging serious faults; The fixed-structure attention mechanism has serious redundant calculations and cannot focus on the fault active period, resulting in low training efficiency; Fixed-bandwidth convolutions / filters cannot adapt to faults with frequency drift and lack frequency-band adaptability; Using cross-entropy loss ignores the imbalance of fault categories and misjudgment costs, and the model tends to identify the majority class, losing its practicality; High-dimensional non-stationary signals lead to non-convergence or gradient explosion during training and poor generalization ability.

[0007] Therefore, the present invention proposes an artificial intelligence-based method for quality control and fault prediction of surgical robots to solve the above problems. Summary of the Invention

[0008] In view of the deficiencies of the prior art, the present invention develops an artificial intelligence-based method for quality control and fault prediction of surgical robots. Through a neural network based on dual-path feature separation and enhancement, the present invention can accurately capture the fault features of surgical robots and improve the accuracy of fault prediction and the generalization ability of the model.

[0009] The technical solution for the present invention to solve the technical problems is an artificial intelligence-based method for quality control and fault prediction of surgical robots, including the following steps:

[0010] S1. Data acquisition: Collect the vibration signal data and operation feedback data of the surgical robot through sensors installed on the surgical robot. At the same time, the edge acquisition device synchronously records the operation state of the surgical robot, and manually annotates the collected data in combination with the surgical log and expert suggestions. Then, the collected data is unified into time-window sample data of a fixed length, and finally a labeled sample data set is formed;

[0011] S2. Data preprocessing: Perform multi-scale normalization processing on the data in the sample data set. Specifically, a normalization method based on frequency-band energy distribution is adopted, and at the same time, wavelet decomposition is combined to perform multi-scale energy equalization on the data to obtain the sample data after normalization processing;

[0012] S3. Construct and train a fault prediction model: Use a convolutional neural network to construct a fault prediction model for the surgical robot, input the sample data after normalization processing into the fault prediction model for training, preset the number of iterative training times, and finally obtain a trained fault prediction model;

[0013] The training process of the fault prediction model includes: constructing a convolutional neural network, performing feature extraction through a temporal feature enhancement module, performing attention enhancement through a dynamic sparse self-attention mechanism, performing feature extraction through a residual adaptive frequency modulation module, performing feature fusion and alignment operations, performing multi-granularity temporal pooling operations, calculating a fault perception loss function, and updating the parameters of the neural network;

[0014] S4. After preprocessing the newly collected data, input it into the trained fault prediction model of the surgical robot, and the fault prediction model outputs the fault category prediction result corresponding to the input data;

[0015] S5. Construct a three-dimensional quality control matrix, including a real-time decision monitoring dimension, a maintenance decision dimension, and a process optimization dimension. The three-dimensional quality control matrix triggers the corresponding automatic calibration program based on the category prediction result output by the fault prediction model.

[0016] S1 is specifically as follows:

[0017] The sensors installed on the surgical robot include a high-precision three-axis acceleration sensor and a force / torque sensor;

[0018] The installation positions of the sensors are the key motion units, joint bearing parts, and execution ends of the surgical robot;

[0019] The sensors continuously collect data at high frequency;

[0020] The operating states of the surgical robot include, but are not limited to, joint angles, speeds, and control commands;

[0021] Manual annotation means that experts mark faults according to abnormal working conditions, and the abnormal working conditions include, but are not limited to, jamming, hysteresis, and abnormal vibration, and classify the fault marks, including, but are not limited to, normal, slightly abnormal, and severely faulty;

[0022] Slice the collected data, unify the collected data into time window samples of a fixed length, and at the same time label the corresponding fault category labels, and finally form a labeled sample data set.

[0023] S2 is specifically as follows:

[0024] Normalize the data in the sample dataset. The data in the sample dataset are vibration signals with characteristics of high-frequency noise, non-stationarity, and time-varying amplitude. A normalization method based on frequency-band energy distribution is adopted, and wavelet decomposition is combined to perform multi-scale energy equalization on the vibration signals. Specifically, the input vibration signal is decomposed into sub-band coefficients of different frequency bands through wavelet decomposition, the energy mean and standard deviation of each layer of wavelet coefficients are calculated, after normalizing each layer of wavelet coefficients, the normalized coefficient distribution is dynamically adjusted through learnable scaling parameters and offset parameters, and finally the normalized sample data is obtained.

[0025] Construct a convolutional neural network as follows:

[0026] The structure of the convolutional neural network specifically adopts a hierarchical architecture design. The input layer receives the multi-scale normalized vibration signals and is sequentially connected to the time-series feature enhancement module, the dynamic sparse self-attention layer, the residual adaptive frequency modulation module, the feature fusion layer, and the fully connected layer;

[0027] The time-series feature enhancement module separates and enhances the steady-state and transient features through a dual-path structure; the dynamic sparse self-attention layer uses a gating mechanism to focus on the active fault periods; the residual adaptive frequency modulation module performs frequency-band adaptive focusing through an adjustable filter; after the feature fusion layer integrates the input vibration signals and filtered features, multi-granularity pooling is used to extract cross-scale statistical features, and finally the fault probability distribution is output through the fully connected layer.

[0028] The feature extraction through the time-series feature enhancement module is as follows:

[0029] Adopt the time-series feature enhancement module to separate the steady-state and transient components of the signal through a dual-path structure, then extract the long-term trend from the steady-state component through convolution operations, and capture the local mutations in the transient component through wavelet transform;

[0030] Specifically, the steady-state component is extracted through low-pass filtering convolution, the transient residual is obtained by subtracting the steady-state component from the input vibration signal, the steady-state component is further enhanced through high-pass filtering, and at the same time, the time-frequency features of the transient component are extracted using discrete wavelet transform. Finally, the features of the two paths are concatenated along the channel dimension.

[0031] The attention enhancement through the dynamic sparse self-attention mechanism is as follows:

[0032] In the dynamic sparse self-attention mechanism module, dynamic sparse gating is introduced. The attention score matrix is sorted row by row through dynamic sparse gating, and only the Top-k significant correlations are retained. Specifically, the first n maximum values in each row are retained, and the rest are set to zero, where n is initially 1 and linearly increases to the preset upper limit k with the number of training epochs, thereby gradually expanding the receptive field.

[0033] Feature extraction is performed through a residual adaptive frequency modulation module as follows:

[0034] The residual adaptive frequency modulation module is used to embed a tunable band filter in the residual connection, enabling the network to adaptively focus on the fault-sensitive frequency band. The basic frequency band features are extracted through convolution in the residual path, and a multi-layer perceptron is used to dynamically generate a band modulation factor to adaptively adjust the center frequency of the band-pass filter and focus on the fault-related frequency band.

[0035] Feature fusion and alignment operations and multi-granularity time pooling operations are performed as follows:

[0036] Feature fusion and alignment operations: The band-filtered features are fused with the input vibration signal and normalized into a sequence of the same length as the input vibration signal. Specifically, a learnable gated convolutional upsampling module is used to expand the time dimension of the band-filtered features to T and perform gated feature fusion with the original vibration signal, dynamically allocating residual weights to align the time axis.

[0037] Multi-granularity time pooling operations: Multi-scale time pooling operations are adopted, and local statistical features are extracted using multi-scale sliding windows. Different window sizes capture short-term shock and long-term trend features. Specifically, three different lengths of sliding windows are preset, with a window step size of 1. When sliding point by point on the time axis, two statistical metrics, namely the maximum value and standard deviation within the window, are maintained in real time. After the scan is completed, the statistical features obtained from the three window lengths are concatenated in the channel direction, and redundant scale information is automatically suppressed through a layer with learnable gating, thereby forming comprehensive multi-scale time features, which can avoid blurring the fault occurrence time period during pooling.

[0038] Calculation of the fault-aware loss function is performed as follows:

[0039] The fault-aware loss function is adopted, combining class-sensitive weights and the fault characteristics of the surgical robot for weighted focusing and misclassification penalty of fault class samples. Specifically, through dynamic class weight allocation and misclassification cost-sensitive modulation, higher loss weights are imposed on fault class samples, and additional penalties are imposed on key misclassification combinations. The calculation formula of the fault-aware loss function is as follows:

[0040] ,

[0041] ,

[0042] Among them, represents the fault-aware loss function; represents the total number of sample data; represents the total number of fault classes; represents the class weight; represents the The true class of a sample; Indicates the penalty coefficient when the true class of the th sample is misclassified as the cth class; Is the misclassification penalty factor; Indicates the logarithmic function with base 10; Indicates that the focus factor suppresses easy-to-classify samples; Indicates the predicted probability that the nth sample belongs to the cth class.

[0043] The parameter update of the neural network is specifically as follows:

[0044] The parameter calculation formula for updating the fault prediction model through adaptive gradient clipping and the AdamW optimizer is as follows:

[0045] ,

[0046] ,

[0047] ,

[0048] ,

[0049] ,

[0050] Among them, Indicates the gradient of the th iteration; Indicates the gradient of the fault perception loss function with respect to the model parameters; Indicates the parameters of the convolutional neural network; Indicates the th iteration of the clipped gradient; Is the threshold parameter for gradient clipping; Indicates the L2 norm; Indicates the th iteration of the first moment estimate; Indicates the th iteration of the first moment estimate; Indicates the th iteration of the second moment estimate; Indicates the th iteration of the second moment estimate; Indicates the exponential decay rate of the second moment estimate; Indicates the th iteration of the parameters of the convolutional neural network; Indicates the th iteration of the parameters of the convolutional neural network; Indicates the learning rate parameter; Indicates the weight decay coefficient.

[0051] The effects provided in the invention content are only the effects of the embodiments, rather than all the effects of the invention. The above technical solutions have the following advantages or beneficial effects:

[0052] The present invention provides a quality control and fault prediction method for surgical robots based on artificial intelligence. By using wavelet decomposition combined with dynamic normalization of frequency band energy, it can effectively retain low-frequency features and suppress high-frequency noise. By splitting the signal into steady-state and transient components, processing them separately and then fusing them, the problem that transient shocks are smoothed in convolution processing can be solved. By adopting a dynamic Top-k gating strategy, the problem of information loss in the initial stage of training can be prevented, the receptive field can be expanded in the later stage, and local fault-related features can be adaptively extracted. By embedding an adjustable filter in the residual path and dynamically adjusting the center frequency in combination with a fully connected layer, continuous tracking of time-varying frequency band faults can be achieved. By adopting multi-scale sliding window maximum and standard deviation pooling and combining with learnable gating, redundant scales can be automatically suppressed. By adopting a dual modulation mechanism of class weight and misclassification cost, higher weights can be assigned to specific fault class samples. By combining adaptive gradient clipping with the AdamW optimizer, the training stability and model generalization ability can be improved.

[0053] In summary, the present invention proposes a quality control and fault prediction method for surgical robots based on artificial intelligence. By means of a neural network based on dual-path feature separation and enhancement, problems such as high-frequency noise interference, loss of transient shock features, difficulty in extracting local fault features, and time-variability of fault frequency bands in surgical robot fault detection are solved, and the fault features of surgical robots can be accurately captured, improving the accuracy of fault prediction and the generalization ability of the model. Brief Description of the Drawings

[0054] The drawings are used to provide further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention.

[0055] Figure 1 It is a schematic flow chart of the method of the present invention.

[0056] Figure 2 It is a schematic diagram of the input vibration signal.

[0057] Figure 3 It is a comparison diagram of the time-frequency distribution of wavelet coefficients between the present invention and conventional methods.

[0058] Figure 4 It is a three-dimensional spectral waterfall diagram of the minimum-maximum normalization method.

[0059] Figure 5 It is a three-dimensional spectral waterfall diagram of the Z-score normalization method.

[0060] Figure 6 This is the three-dimensional spectral waterfall plot of the normalization method of the present invention.

[0061] Figure 7 This is the experimental graph of the dynamic sparse attention matrix.

[0062] Figure 8 This is the smooth curve graph of the attention sparsity.

[0063] Figure 9 This is the trend graph of the effective feature energy.

[0064] Figure 10 This is the experimental graph of the frequency response of the adaptive filter bank. Detailed implementation manners

[0065] In order to clearly illustrate the technical features of this solution, the present invention will be elaborated in detail below through specific implementation manners and in combination with its accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below.

[0066] Embodiment 1

[0067] A quality control and fault prediction method for a surgical robot based on artificial intelligence, comprising the following steps:

[0068] S1. Data acquisition: Vibration signal data and operation feedback data of the surgical robot are collected through sensors installed on the surgical robot. At the same time, the operation state of the surgical robot is synchronously recorded by the edge acquisition device, and the collected data is manually labeled in combination with the operation log and expert suggestions. Then, the collected data is unified into time window sample data of a fixed length, and finally a labeled sample data set is formed;

[0069] S2. Data preprocessing: The data in the sample data set is subjected to multi-scale normalization processing. Specifically, a normalization method based on the energy distribution of frequency bands is adopted, and at the same time, wavelet decomposition is combined to perform multi-scale energy equalization on the data to obtain the sample data after normalization processing;

[0070] S3. Constructing and training a fault prediction model: A fault prediction model of the surgical robot is constructed by using a convolutional neural network. The sample data after normalization processing is input into the fault prediction model for training, and the number of iterative training times is preset in advance. Finally, a trained fault prediction model is obtained;

[0071] The training process of the fault prediction model includes: constructing a convolutional neural network, performing feature extraction through a time series feature enhancement module, performing attention enhancement through a dynamic sparse self-attention mechanism, performing feature extraction through a residual adaptive frequency modulation module, performing feature fusion and alignment operations, performing multi-granularity time pooling operations, calculating a fault perception loss function, and updating the parameters of the neural network;

[0072] S4. After preprocessing the newly collected data, input it into the trained fault prediction model of the surgical robot, and the fault prediction model outputs the fault category prediction result corresponding to the input data;

[0073] S5. Construct a three-dimensional quality control matrix, including a real-time decision monitoring dimension, a maintenance decision dimension, and a process optimization dimension. The three-dimensional quality control matrix triggers the corresponding automatic calibration program based on the category prediction result output by the fault prediction model.

[0074] In the specific implementation manner, S1 is specifically as follows:

[0075] The sensors installed on the surgical robot include a high-precision three-axis acceleration sensor and a force / torque sensor;

[0076] The installation positions of the sensors are the key motion units, joint bearing parts, and execution ends of the surgical robot;

[0077] The sensors continuously collect data at high frequency;

[0078] The operating states of the surgical robot include but are not limited to joint angles, speeds, and control commands;

[0079] Manual annotation means that experts mark faults according to abnormal working conditions. The abnormal working conditions include but are not limited to jamming, hysteresis, and abnormal vibration, and grade the fault marks, including but not limited to normal, slightly abnormal, and severe faults;

[0080] Slice the collected data, unify the collected data into time window samples of a fixed length, and at the same time label the corresponding fault category labels, and finally form a labeled sample data set.

[0081] In the specific implementation manner, S2 is specifically as follows:

[0082] Normalize the data in the sample dataset. The data in the sample dataset are vibration signals with high-frequency noise, non-stationarity, and time-varying amplitude characteristics. A normalization method based on frequency-band energy distribution is adopted, and wavelet decomposition is combined to perform multi-scale energy equalization on the vibration signals. Specifically, the input vibration signal is decomposed into sub-band coefficients of different frequency bands through wavelet decomposition, the energy mean and standard deviation of each layer of wavelet coefficients are calculated, after normalizing each layer of wavelet coefficients, the normalized coefficient distribution is dynamically adjusted through learnable scaling parameters and offset parameters, and finally the normalized sample data are obtained.

[0083] The calculation formula for normalizing each layer of wavelet coefficients is as follows:

[0084] ,

[0085] where, represents the normalized output of the wavelet decomposition coefficients of the th layer, representing the vibration energy of a specific frequency band; represents the th layer coefficient after the original vibration signal is decomposed by wavelet; represents the energy mean of the th layer coefficient; represents the energy standard deviation of the th layer coefficient; represents the trainable scaling parameter of the th layer coefficient, represents the trainable offset parameter of the th layer coefficient. The training method is gradient descent combined with the backpropagation algorithm, and and are updated through the chain rule to make them adapt to the feature distributions of different frequency bands; represents the minimum constant to prevent division by zero;

[0086] The energy mean rather than the amplitude mean is calculated because calculating the energy mean is more suitable for fault impact identification. The calculation formula for the energy mean is as follows:

[0087] ,

[0088] where, represents the length of the time series; represents the value of the th layer wavelet decomposition coefficient at time point ;

[0089] By calculating the standard deviation of the energy, the instability of the vibration amplitude can be better captured. The calculation formula for the energy standard deviation is as follows:

[0090] ;

[0091] It should be noted that by selecting the calculation based on the mean and standard deviation of the frequency band energy , rather than the traditional normalized time-domain statistic, it can dynamically adjust the normalization coefficient according to the frequency band energy for the non-stationarity of the vibration signal of the surgical robot, suppress high-frequency noise and retain low-frequency fault characteristics.

[0092] In the specific implementation manner, the convolutional neural network is constructed as follows:

[0093] The convolutional neural network structure specifically adopts a hierarchical architecture design. The input layer receives the vibration signal after multi-scale normalization and is sequentially connected to the temporal feature enhancement module, the dynamic sparse self-attention layer, the residual adaptive frequency modulation module, the feature fusion layer, and the fully connected layer;

[0094] The temporal feature enhancement module separates and enhances the steady-state and transient features through a dual-path structure; the dynamic sparse self-attention layer uses a gating mechanism to focus on the active fault period; the residual adaptive frequency modulation module performs frequency-band adaptive focusing through an adjustable filter; after the feature fusion layer integrates the input vibration signal and the filtered features, it extracts cross-scale statistical features through multi-granularity pooling, and finally outputs the fault probability distribution through the fully connected layer.

[0095] It should be noted that several (such as 3 layers) fully connected layers can be added to the middle layer of the convolutional neural network designed by the present invention, or the modules / layers proposed by the present invention can be deleted, and it should still fall within the protection scope of the present invention.

[0096] In the specific implementation manner, the feature extraction through the temporal feature enhancement module is as follows:

[0097] The temporal feature enhancement module is adopted to separate the steady-state and transient components of the signal through a dual-path structure, and then the long-term trend is extracted from the steady-state component through convolution operation, and the local mutation in the transient component is captured through wavelet transform. Mechanical jamming vibration signals, including anomalies such as instant collisions and tool tip hysteresis in surgical robots, are only visible in local windows. Using the temporal feature enhancement module for feature extraction can solve the problem of transient feature loss and enhance the pertinence of the transient impact features of mechanical jamming vibration signals;

[0098] Specifically, the steady-state component is extracted through low-pass filter convolution, the transient residual is obtained by subtracting the steady-state component from the input vibration signal, the steady-state component is enhanced by high-pass filtering, and at the same time, the time-frequency features of the transient component are extracted by discrete wavelet transform. Finally, the two-path features are concatenated along the channel dimension.

[0099] The calculation formula of the temporal feature enhancement module is as follows:

[0100] ,

[0101] ,

[0102] ,

[0103] Among them, represents a low-pass filter with a size of 5 and a cut-off frequency of 50 Hz. The steady-state components of the vibration signal of the surgical robot are usually concentrated in the low frequency, usually with a frequency lower than 50 Hz, while transient impacts such as mechanical jamming are manifested as high-frequency mutations; represents the vibration signal after multi-scale normalization; represents the steady-state component, which is the signal after low-pass filtering; characterizes a one-dimensional convolution with a step size of 1, using a high-pass filter to extract the low-frequency steady-state component; represents a one-dimensional convolution operation; is the residual between the input vibration signal and the steady-state component, characterizing the transient impact component and emphasizing transient anomalies such as mechanical jamming; represents a high-pass filter. By enhancing the high-frequency details of the steady-state component through the high-pass filter, it is possible to avoid the long-term trend from masking the transient features; characterizes a one-dimensional convolution with a step size of 3, using a low-pass filter to extract the low-frequency steady-state component; represents feature concatenation; represents the enhanced time-series features, fusing the time-frequency information of the steady-state and transient components; represents discrete wavelet transform, extracting the time-frequency features of the transient component, separating and locating the time-frequency mutation points of the transient.

[0104] In the specific implementation, the attention enhancement is performed through the dynamic sparse self-attention mechanism as follows:

[0105] In the dynamic sparse self-attention mechanism module, dynamic sparse gating is introduced. The attention score matrix is sorted row by row through dynamic sparse gating, and only the Top-k significant associations are retained. Specifically, the first n maximum values in each row are retained, and the rest are set to zero. Here, n is initially 1 and linearly increases to the preset upper limit k with the number of training rounds, thereby gradually expanding the receptive field and avoiding the loss of key features due to excessive sparsity in the early training.

[0106] The related calculations of the dynamic sparse self-attention mechanism are as follows:

[0107] ,

[0108] ,

[0109] ,

[0110] ,

[0111] ,

[0112] ,

[0113] Among them, represents a query matrix, which characterizes the dynamic association pattern between time points, learns how to extract query patterns from input features, characterizes which feature dimensions need to be concerned at different time points, and is optimized through the backpropagation algorithm; represents a key matrix, which characterizes the feature encoding pattern of time points, learns how to encode key patterns, characterizes how the features of time points are retrieved by other positions, and is optimized through the backpropagation algorithm; represents a value matrix, which learns how to extract value patterns from input features, that is, the semantic representation of fault features, and is optimized through the backpropagation algorithm; represents a query vector matrix; represents a key vector matrix; represents the th query vector of the time point; represents the th key vector of the time point; represents transpose; represents the dimension of the key vector, which characterizes the dimension scaling factor; represents the th time point and the th attention score between time points; represents the element in the th row and th column of the sparse gating mask; represents taking the first rows of the attention scores sorted and taking the first maximum values. Taking the first maximum values for each row. It should be noted that n is dynamically adjusted during training. Initially, n = 1 and gradually expands to a preset maximum threshold, such as 5, to avoid losing key features due to excessive sparsity in early training and gradually expand the receptive field to capture long-range dependencies; represents the Softmax function; represents the sparse attention score matrix, and its elements are ; represents the value vector matrix; is the feature representation after attention weighting.

[0114] In the specific implementation manner, the feature extraction is specifically as follows through the residual adaptive frequency modulation module:

[0115] Due to the wear of surgical robot components and the gear meshing error, the fault frequency band often changes at any time. The present invention embeds an adjustable band filter in the residual connection by using a residual adaptive frequency modulation module, enabling the network to adaptively focus on the fault-sensitive frequency band. The basic frequency band features are extracted through the convolution of the residual path, and a multi-layer perceptron is used to dynamically generate a frequency band modulation factor to adaptively adjust the center frequency of the band-pass filter and focus on the fault-related frequency band.

[0116] The calculation formula of the residual adaptive frequency modulation module is as follows:

[0117] ,

[0118] ,

[0119] ,

[0120] Among them, represents the one-dimensional convolution of the residual path, and a residual convolution kernel is used as the filter bank to extract the basic frequency band features; represents the residual convolution kernel with a size of 3; represents the residual feature; represents the reference frequency band range; represents using a multi-layer perceptron, with the input to generate a frequency band modulation factor; represents the multi-layer perceptron with a hidden layer dimension of 64; represents the Sigmoid activation function; represents the Butterworth band-pass filter. By dynamically adjusting the filtering range, it can avoid the inability of traditional fixed-bandwidth convolution to adapt to the time-varying fault frequency band, and realize the joint optimization of frequency modulation and feature enhancement through the residual path; represents dynamically filtering the residual feature in a frequency band, and the center frequency is controlled by ; represents the center frequency, which is dynamically generated by the multi-layer perceptron according to the input features. In response to the time-variation of the fault-related frequency band, such as the resonance frequency shift caused by bearing wear, it adaptively focuses on the sensitive frequency band; represents the reference frequency band, such as 50 to 150 Hz; represents the frequency band filtered feature.

[0121] In the specific implementation manner, the feature fusion and alignment operation and the multi-granularity time pooling operation are as follows:

[0122] Feature fusion and alignment operation: Feature fusion is performed on the band-filtered features and the input vibration signal, and it is normalized into a sequence of the same length as the input vibration signal. Specifically, a learnable gated convolutional upsampling module is used to extend the time dimension of the band-filtered features to T and perform gated feature fusion with the original vibration signal, dynamically allocating residual weights to align the time axes;

[0123] The calculation formula for the feature fusion and alignment operation is as follows:

[0124] ,

[0125] ,

[0126] ,

[0127] where, represents a one-dimensional transposed convolution operation, using a convolution kernel to perform upsampling, with a stride of 2, indicating that the band-filtered features are upsampled to length , retaining the time localization information of the fault features, such as the exact moment when the fault occurs; represents the upsampling convolution kernel, with a size of 4 and a stride of 2; represents the upsampled features; represents concatenating the upsampled features and the original vibration signal along the channel dimension; represents the gated convolution kernel, with a size of 3; represents performing a one-dimensional convolution operation on the concatenated features using the convolution kernel to extract the gated weights; represents the aligned features after gated feature fusion; represents the gated weight matrix generated by the Sigmoid function, used to dynamically adjust the fusion ratio of the upsampled features and the original signal, and solve the time axis alignment problem between the band-filtered features and the original signal, such as the timing synchronization requirements of the surgical robot's actions; represents feature concatenation; represents element-wise multiplication.

[0128] Multi-granularity Time Pooling Operation: The multi-scale time pooling operation is adopted to extract local statistical features using multi-scale sliding windows. Different window sizes capture short-term shock and long-term trend features. Specifically, three sliding windows with different lengths are preset, the window step size is 1. When sliding point by point on the time axis, two statistical indicators, namely the maximum value and the standard deviation within the window, are maintained in real time. After the scanning is completed, the statistical features obtained from the three window lengths are concatenated in the channel direction, and the redundant scale information is automatically suppressed through a layer with learnable gating, thus forming comprehensive multi-scale time features. This can avoid blurring the time period when the fault occurs during pooling.

[0129] The calculation formula of the multi-granularity time pooling operation is as follows:

[0130] ,

[0131] ,

[0132] Among them, represents the pooling result of the feature; represents the maximum value within the calculation window; represents the standard deviation within the calculation window; represents the subsequence of the feature of within the window represents the window size; represents the concatenation result of the multi-scale pooling features; Characterizes the concatenation of the pooling results of different window sizes along the feature dimension, with sizes of 5, 10, and 20 respectively. For the surgical robot, the fault may manifest as instantaneous anomalies or long-term performance decline. Instantaneous anomalies such as gear fractures, and long-term performance decline such as motor wear. The multi-scale can capture short-term shocks and long-term trends.

[0133] In the specific implementation manner, the calculation of the fault perception loss function is as follows:

[0134] The fault perception loss function is adopted, combined with class-sensitive weights and the fault characteristics of the surgical robot, for weighted focusing and misclassification penalty of fault class samples; specifically, through dynamic class weight allocation and misclassification cost-sensitive modulation, a higher loss weight is imposed on fault class samples, and an additional penalty is imposed on key misclassification combinations. The calculation formula of the fault perception loss function is as follows:

[0135] ,

[0136] ,

[0137] Among them, represents the fault perception loss function; represents the total number of sample data; Indicates the total number of fault categories; Indicates the category weight; Indicates the true category of the Indicates the penalty coefficient when the true category of the sample is misclassified as the c-th category; ; Indicates the logarithmic function with base 10; Indicates that the focus factor suppresses easy-to-classify samples, set ; Indicates the predicted probability that the n-th sample belongs to the c-th category.

[0138] Calculate the category weight through the frequency of the category , the category weight is inversely proportional to the square root of the frequency, suppressing the dominant loss of high-frequency categories. mainly for the scarcity of surgical robot fault samples, such as intraoperative emergency shutdown events, which can enhance the attention to small-sample categories. The category weight calculation formula is as follows:

[0139] [[ID=�2]];

[0140] where Indicates the frequency of the -th fault category appearing,

[0141] Indicates the minimum constant to prevent division by zero.

[0142] In the specific implementation manner, the parameter update of the neural network is specifically as follows:

[0143] ;

[0144] ;

[0145] ;

[0146] ;

[0147] ;

[0148] where Indicates the gradient of the -th iteration; Indicates the gradient of the fault perception loss function with respect to the model parameters; Indicates the The gradient after the $i$-th iteration of clipping; is the threshold parameter for gradient clipping, set , to limit the gradient norm and prevent gradient explosion caused by abnormal samples in the surgical robot data; represents the L2 norm; represents the first moment estimate of the $i$-th iteration, characterizing the exponentially weighted moving average of the current gradient, used to smooth the gradient fluctuations; represents the first moment estimate of the $i$-th iteration, characterizing the smoothed gradient result of the previous iteration; is the exponential decay rate of the first moment estimate; represents the second moment estimate of the $i$-th iteration, characterizing the exponentially weighted moving average of the current squared gradient, used to adaptively adjust the learning rate; represents the second moment estimate of the $i$-th iteration, characterizing the weighted average of the squared gradients in the previous iteration; represents the exponential decay rate of the second moment estimate, set ; represents the parameters of the convolutional neural network at the $i$-th iteration; represents the parameters of the convolutional neural network at the $i$-th iteration; represents the learning rate parameter, set ; represents the weight decay coefficient, set .

[0149] Example 2

[0150] As Figure 2 and Figure 3 described, by using the wavelet scale map to compare the transient shock feature capture capabilities of the time series feature enhancement module and the traditional convolutional network, and by analyzing the simulated signal containing transient pulses and observing the time-frequency distribution of the wavelet coefficients after different methods of processing, the short-time shock component is clearly visible in the time-domain waveform of the input vibration signal. However, in the wavelet coefficient map of the traditional method, only a weak response is presented in this period, indicating that the conventional convolutional operation smooths the transient features. The dual-path structure adopted in the present invention shows a strong red response region during the shock occurrence period, the coefficient amplitude is significantly increased and the boundary is clear. At the same time, the coefficient amplitudes in the background noise region are suppressed. The improvement of the time-frequency localization ability verifies the design advantage of separating the steady-state and transient components. By combining low-pass filtering and residual analysis, the sensitivity of the network to transient fault features such as mechanical jamming is enhanced.

[0151] As Figure 4 , Figure 5 , Figure 6As shown, the regulation effects of different normalization methods on the energy of the vibration signal bands are compared through three-dimensional spectral waterfall diagrams. According to the energy distribution of the vibration signal input to the surgical robot in the time-frequency domain, the energy change trends of high-frequency noise and low-frequency fault characteristics are mainly observed. The input vibration signal shows energy fluctuations in the main fault characteristic bands, and at the same time, noise energy accumulation appears in the high-frequency region. After using the traditional normalization method, the high-frequency noise energy is further amplified, forming an obvious red energy band, while the low-frequency characteristic energy shows irregular attenuation. In contrast, the multi-scale normalization processing of the present invention shows a more balanced energy distribution in the three-dimensional time-frequency diagram, maintaining a stable energy density in the low-frequency characteristic region, and presenting a significant blue attenuation band in the high-frequency noise region. The energy regulation characteristics verify that the normalization strategy based on the dynamic adjustment of band energy can effectively suppress the amplitude difference of non-stationary signals, avoid high-frequency noise interference, and at the same time retain the integrity of key fault characteristics.

[0152] As Figure 7 , Figure 8 , Figure 9 shown, the working principle and its advantages of the dynamic sparse self-attention mechanism are analyzed through heat maps and training curves. The visualization of the attention matrix shows that the traditional method forms a diffuse energy distribution within the entire sequence range, with a large number of inefficient weak associations. The matrix generated by the present invention presents a diagonal banded focusing pattern, forming a highlighted area only during the active period of fault characteristics. During the training process, the effective feature energy curve of the traditional method shows violent fluctuations, and the sparsity reaches the extreme value prematurely, resulting in the loss of key features. The present invention dynamically adjusts the sparse gating threshold, so that the attention sparsity increases progressively with the training process, and finally, while retaining the key time associations, the effective feature energy is stably maintained at a high level. The adaptive focusing mechanism not only avoids waste of computing resources but also ensures accurate feature extraction during the fault-sensitive period.

[0153] As Figure 10 shown, the dynamic tracking ability of the residual adaptive frequency modulation module is verified through the frequency response curves of the filter bank. The experimental scenario of the fault characteristic frequency drifting with time is analyzed, and the fixed-bandwidth filter is compared with the dynamic filter bank generated by the present technology. The experimental results show that the passband position of the fixed filter is constant, and its gain drops rapidly when the fault frequency shifts. However, the center frequency of the filter bank generated by the adaptive modulation can migrate continuously, always maintaining a high-gain coverage of the characteristic frequency band, and the passband width is stable without distortion. In terms of the out-of-band noise suppression ability, the filter of the present invention shows a steep attenuation characteristic in the non-sensitive frequency band. The dynamic frequency focusing characteristic proves that the residual structure embedded with adjustable filters can effectively cope with the time-varying working conditions of mechanical systems and improve the signal ratio of fault characteristics.

[0154] Although the specific embodiments of the invention have been described above in conjunction with the accompanying drawings, they are not intended to limit the scope of protection of the present invention. Based on the technical solutions of the present invention, various modifications or variations that can be made by those skilled in the art without creative efforts are still within the scope of protection of the present invention.

Claims

1. A quality control and fault prediction method for a surgical robot based on artificial intelligence, characterized in that, It includes the following steps: S1. Data acquisition: The vibration signal data and operation feedback data of the surgical robot are collected through sensors installed on the surgical robot. At the same time, the edge acquisition device synchronously records the operation state of the surgical robot, and the collected data is manually annotated in combination with the surgical log and expert suggestions. Then, the collected data is unified into time window sample data of a fixed length, and finally a labeled sample data set is formed. S2. Data preprocessing: The data in the sample data set is subjected to multi-scale normalization processing. Specifically, a normalization method based on frequency band energy distribution is adopted, and at the same time, wavelet decomposition is combined to perform multi-scale energy equalization on the data to obtain the sample data after normalization processing. S2 is specifically as follows: The data in the sample data set is normalized. The data in the sample data set is a vibration signal with characteristics of high-frequency noise, non-stationarity, and time-varying amplitude. A normalization method based on frequency band energy distribution is adopted, and at the same time, wavelet decomposition is combined to perform multi-scale energy equalization on the vibration signal. Specifically, the input vibration signal is decomposed into sub-band coefficients of different frequency bands through wavelet decomposition, the energy mean and standard deviation of each layer of wavelet coefficients are calculated, and after each layer of wavelet coefficients is normalized, the normalized coefficient distribution is dynamically adjusted through learnable scaling parameters and offset parameters, and finally the sample data after normalization processing is obtained. S3. Construct a fault prediction model and train it: A convolutional neural network is used to construct a fault prediction model for the surgical robot, and the sample data after normalization processing is input into the fault prediction model for training. The number of iterative training times is preset in advance, and finally a trained fault prediction model is obtained. The training process of the fault prediction model includes: constructing a convolutional neural network, performing feature extraction through a temporal feature enhancement module, performing attention enhancement through a dynamic sparse self-attention mechanism, performing feature extraction through a residual adaptive frequency modulation module, performing feature fusion and alignment operations, performing multi-granularity time pooling operations, calculating a fault-aware loss function, and updating the parameters of the neural network. Constructing the convolutional neural network is specifically as follows: The convolutional neural network structure specifically adopts a hierarchical architecture design. The input layer receives the vibration signal after multi-scale normalization and is sequentially connected to the temporal feature enhancement module, the dynamic sparse self-attention layer, the residual adaptive frequency modulation module, the feature fusion layer, and the fully connected layer. The temporal feature enhancement module separates and enhances the steady-state and transient features through a dual-path structure; the dynamic sparse self-attention layer uses a gating mechanism to focus on the fault active period; the residual adaptive frequency modulation module performs frequency band adaptive focusing through an adjustable filter; after the feature fusion layer integrates the input vibration signal and the filtered features, cross-scale statistical features are extracted through multi-granularity pooling, and finally the fault probability distribution is output through the fully connected layer. Performing feature extraction through the temporal feature enhancement module is specifically as follows: A temporal feature enhancement module is adopted to separate the steady-state and transient components of the signal through a dual-path structure, and then the long-term trend is extracted from the steady-state component through convolutional operations, and the local mutations in the transient component are captured through wavelet transform. Specifically, the steady-state component is extracted by low-pass filter convolution. The transient residual is obtained by subtracting the steady-state component from the input vibration signal. Then, the steady-state component is enhanced by high-pass filtering. At the same time, the time-frequency features of the transient component are extracted by discrete wavelet transform. Finally, the features of the two paths are concatenated along the channel dimension. The attention enhancement is carried out through the dynamic sparse self-attention mechanism as follows: In the dynamic sparse self-attention mechanism module, dynamic sparse gating is introduced. The attention score matrix is sorted row by row through dynamic sparse gating, and only the Top-k significant associations are retained. Specifically, the first n maximum values in each row are retained, and the rest are set to zero, where n is initially 1 and linearly increases to the preset upper limit k with the number of training rounds, thereby gradually expanding the receptive field. The feature extraction is carried out through the residual adaptive frequency modulation module as follows: The residual adaptive frequency modulation module is used to embed an adjustable bandpass filter in the residual connection, enabling the network to adaptively focus on the fault-sensitive frequency band. The basic frequency band features are extracted by convolution in the residual path. A multi-layer perceptron is used to dynamically generate the band modulation factor to adaptively adjust the center frequency of the bandpass filter and focus on the fault-related frequency band. S4. The newly collected data is preprocessed and then input into the trained fault prediction model of the surgical robot, and the fault prediction model outputs the fault category prediction result corresponding to the input data. S5. A three-dimensional quality control matrix is constructed, including the real-time decision monitoring dimension, the maintenance decision dimension, and the process optimization dimension. The three-dimensional quality control matrix triggers the corresponding automatic calibration program based on the category prediction result output by the fault prediction model.

2. The quality control and fault prediction method for the surgical robot based on artificial intelligence according to claim 1, characterized in that, S1 is specifically as follows: The sensors installed on the surgical robot include a high-precision three-axis acceleration sensor and a force / torque sensor. The installation positions of the sensors are the key motion units, joint bearing parts, and execution ends of the surgical robot. The sensors continuously collect data at a high frequency. The operating states of the surgical robot include, but are not limited to, joint angles, speeds, and control commands. Manual annotation is that experts mark faults according to abnormal working conditions, which include, but are not limited to, jamming, hysteresis, and abnormal vibration, and grade the fault marks, including, but are not limited to, normal, slightly abnormal, and severely faulty. The collected data is sliced, and the collected data is unified into time window samples of a fixed length. At the same time, the corresponding fault category labels are marked, and finally a labeled sample data set is formed.

3. The quality control and fault prediction method for the surgical robot based on artificial intelligence according to claim 2, wherein The feature fusion and alignment operation and the multi-granularity time pooling operation are specifically as follows: Feature fusion and alignment operation: The bandpass filter features and the input vibration signal are fused, and normalized into a sequence of the same length as the input vibration signal. Specifically, a learnable gated convolutional upsampling module is used to expand the time dimension of the bandpass filter features to T and perform gated feature fusion with the original vibration signal, dynamically allocating residual weights to align the time axes. Multi-granularity time pooling operation: Adopt multi-scale time pooling operation, use multi-scale sliding windows to extract local statistical features. Different window sizes capture short-term shock and long-term trend features. Specifically, three sliding windows with different lengths are preset, the window stride is 1. When sliding point by point on the time axis, two statistical indicators, namely the maximum value and the standard deviation within the window, are maintained in real time. After the scanning is completed, the statistical features obtained from the three window lengths are concatenated in the channel direction, and the redundant scale information is automatically suppressed through a layer with learnable gating, thereby forming comprehensive multi-scale time features. This can avoid blurring the time period when the fault occurs during pooling.

4. The quality control and fault prediction method for the surgical robot based on artificial intelligence according to claim 3, characterized in that, The calculation of the fault-aware loss function is specifically as follows: Adopt the fault-aware loss function, combine the class-sensitive weight and the fault characteristics of the surgical robot, and perform weighted focusing and misclassification penalty on the fault class samples; specifically, through dynamic class weight assignment and misclassification cost-sensitive modulation, apply a higher loss weight to the fault class samples, and impose an additional penalty on the key misclassification combinations. The calculation formula of the fault-aware loss function is as follows: , , Among them, represents the fault perception loss function; represents the total number of sample data; represents the total number of fault categories; represents the class weight; represents the true class of the th sample; represents the penalty coefficient when the true class of the th sample is misclassified as the c-th class; is the misclassification penalty factor; represents the logarithmic function with base 10; represents the focus factor to suppress easy-to-classify samples; represents the predicted probability that the n-th sample belongs to the c-th class.

5. The quality control and fault prediction method for the surgical robot based on artificial intelligence according to claim 4, characterized in that, The parameter update of the neural network is specifically as follows: Update the parameters of the fault prediction model through adaptive gradient clipping and the AdamW optimizer. The calculation formula is as follows: , , , , , Among them, represents the gradient of the th iteration; represents the gradient of the fault perception loss function with respect to the model parameters; represents the parameters of the convolutional neural network; represents the gradient after clipping at the th iteration; is the threshold parameter for gradient clipping; represents the L2 norm; represents the first moment estimate at the th iteration; represents the first moment estimate at the th iteration; represents the second moment estimate at the th iteration; represents the second moment estimate at the th iteration; represents the exponential decay rate of the second moment estimate; represents the parameters of the convolutional neural network at the th iteration; represents the parameters of the convolutional neural network at the th iteration; represents the learning rate parameter; represents the weight decay coefficient.

Citation Information

Patent Citations

  • Industrial robot fault diagnosis method, system and device and storage medium

    CN114897102A

  • Wheeled robot intelligent fault diagnosis method and system based on graph convolutional network

    CN115099268A

  • Intelligent fault monitoring and diagnosis system and method for industrial robot

    CN118938882A

  • Intelligent low-voltage moulded case circuit breaker

    CN109346386A

  • Artificial intelligence data acquisition system for pipeline state monitoring

    CN113775942A