Robot holographic compliant assembly health diagnosis method based on deep learning

By using dynamic DS evidence reasoning and graph neural network topology modeling, the problem of unstable multi-source data fusion in the robot holographic compliant assembly system was solved, achieving higher diagnostic accuracy and robustness, and improving fault location sensitivity and the ability to identify minor faults.

CN121881277BActive Publication Date: 2026-05-15GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU UNIVERSITY
Filing Date
2026-03-20
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing health diagnosis methods for robotic holographic compliant assembly systems suffer from problems such as high uncertainty in multi-source data fusion, difficulty in adaptively quantifying the credibility of evidence, unstable fusion due to evidence conflicts, and difficulty in effectively modeling complex dependencies, resulting in poor recognition accuracy and robustness.

Method used

By employing a deep learning-based approach, dynamic DS evidence reasoning based on multi-source sensor signals and graph neural network topology modeling, adaptive weighting and conflict suppression fusion of multi-source evidence are achieved. Combined with a progressive redundancy removal strategy, a sparse representation is constructed to improve the expression and recognition capabilities of fault features.

Benefits of technology

It effectively reduces the uncertainty of multi-source data, improves the accuracy and robustness of health diagnosis of the robot holographic compliant assembly system under varying working conditions, enhances the sensitivity of fault location, improves the ability to identify minor faults, and reduces the impact of noise propagation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121881277B_ABST
    Figure CN121881277B_ABST
Patent Text Reader

Abstract

The present application relates to a robot holographic compliant assembly health diagnosis method based on deep learning, comprising: obtaining acoustic and vibration window sequences based on robot holographic compliant assembly process acoustic signals and vibration signals; extracting features of each sequence and mapping them into basic probability distribution functions to construct acoustic and vibration evidence corresponding to the window; calculating the confidence of acoustic and vibration evidence, determining the fusion weight and correcting the basic probability distribution function; obtaining the fusion evidence based on the conflict degree between the evidence and introducing the fuzzy entropy optimization D-S evidence combination process, and then constructing the graph; based on the dynamic similarity, the spatial semantics and time dimension of the node time sequence features are aggregated to obtain the node importance representation; after sparse processing, the classification and decision layer outputs the health state to obtain a model for assembly health diagnosis. The present application reduces the uncertainty of multi-source data and effectively improves the expression and recognition ability of weak fault features under variable working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of equipment status monitoring, fault diagnosis and health assessment technology for robot holographic compliant assembly systems, and in particular to a deep learning-based method for health diagnosis of robot holographic compliant assembly. Background Technology

[0002] As a key assembly execution unit in intelligent production lines, robotic holographic compliant assembly systems operate under complex conditions such as high-speed reciprocating motion, frequent starts and stops, impact contact, load fluctuations, changes in assembly cycle time, and lubrication degradation. If vulnerable components such as critical parts of the robotic holographic compliant assembly system experience assembly misalignment, loosening or wear of fixtures / end-effectors, increased joint transmission clearance, or degradation of reducers or bearings, it can lead to drastic performance degradation or even catastrophic failures, causing line stoppages, equipment cascading damage, and significant economic losses. Therefore, condition monitoring and health assessment of robotic holographic compliant assembly systems are one of the core tasks for achieving predictive maintenance and intelligent operation and maintenance. With the development of the Industrial Internet of Things (IIoT), sensors and data acquisition systems can continuously acquire large-scale, multi-source data during equipment operation, providing a data foundation for intelligent diagnostics. However, this also places higher demands on the robustness, real-time performance, and generalization capabilities of diagnostic algorithms.

[0003] Existing methods for fault diagnosis and health assessment of robotic holographic compliant assembly systems can be broadly categorized into two types: One type relies on feature engineering methods based on mathematical statistics and signal processing. These methods typically extract statistical indicators such as feature entropy, kurtosis, mean, and spectral features from the original signal manually or semi-manually, then combine these with threshold discrimination or traditional classifiers to achieve state recognition. This type of method heavily depends on prior knowledge and the quality of feature selection. When faced with complex operating conditions, noise interference, and non-stationary signals, it is prone to problems such as insufficient feature representation and poor transferability, thus affecting diagnostic accuracy and stability. The other type is based on machine learning or deep learning methods. These methods utilize data-driven nonlinear mapping capabilities to achieve end-to-end feature learning, which can reduce reliance on manual features and improve recognition performance to some extent. However, deep models have relatively insufficient interpretability and still face several challenges in multi-source information fusion, evidence conflict handling, and complex topological relationship modeling.

[0004] In real-world industrial scenarios, a single sensor often struggles to comprehensively characterize the fault evolution process of key components in a robot's holographic compliant assembly process. Acoustic and vibration signals are complementary in terms of frequency domain, time domain, and propagation path sensitivity, respectively, making multi-sensor fusion an important approach to improve diagnostic reliability. However, existing multi-source fusion methods often employ fixed-weight fusion strategies that struggle to adapt to changing operating conditions and noise disturbances. Furthermore, time-varying inconsistencies and conflicting evidence may exist between multiple sources, leading to unstable fusion results. Evidence theory (such as DS evidence reasoning) can be used to process uncertain information and achieve multi-source evidence fusion, but existing DS-based diagnostic methods often rely on manually setting evidence credibility, combination rules, or static weights, making it difficult to achieve adaptive and stable evidence combination under conditions of high noise, high conflict, and significant changes in operating conditions.

[0005] In recent years, deep neural network models have attracted increasing research interest. However, existing deep learning methods still have shortcomings in multi-source information fusion and complex dependency modeling: on the one hand, fixed weights or static fusion are difficult to adapt to operating condition fluctuations and noise disturbances, and time-varying inconsistencies between multi-source information can easily lead to evidence conflicts, thus causing fusion instability; on the other hand, topological dependency modeling often adopts predefined or dense connection structures, which can easily introduce redundant connections and noise propagation, affecting the discriminability and robustness of fault features. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a deep learning-based method for health diagnosis of holographic compliant assembly in robots. This method solves the technical problems of poor recognition accuracy and robustness caused by the strong uncertainty of multi-source data, the difficulty in adaptively quantifying the credibility of evidence, the instability of fusion due to evidence conflicts, and the difficulty in effectively modeling complex dependencies and the easy introduction of redundant propagation during the health diagnosis process of existing holographic compliant assembly systems.

[0007] The technical solution adopted in this invention is as follows:

[0008] A deep learning-based method for health diagnosis of holographic compliant assembly in robots, comprising:

[0009] The system collects multi-source sensor signals that characterize the assembly process state information under the holographic dynamic load spectrum during the holographic compliant assembly process of the robot. These signals include acoustic and vibration signals collected by the corresponding sensors under the same operating conditions.

[0010] The acoustic and vibration signals are normalized, and a sliding window segmentation is used to obtain the acoustic window sequence and vibration window sequence corresponding to each sliding window.

[0011] For each of the acoustic window sequences and vibration window sequences, features are extracted and mapped to a basic probability assignment function. To construct acoustic and vibration evidence for corresponding windows; calculate the confidence levels of the acoustic and vibration evidence within the same sliding window; determine the fusion weights based on the confidence levels and adjust accordingly. Based on the revised The degree of conflict between acoustic and vibration evidence is calculated, and fuzzy entropy is introduced to constrain and optimize the DS evidence combination process to obtain window-level fused evidence.

[0012] Graph construction is performed using the window-level fusion evidence as node features to determine the node set and edge weights, and the node features are initialized and embedded.

[0013] By utilizing graph convolution and structural consistency constraints, local temporal modeling of node temporal features is performed, dynamic similarity is calculated, and spatial semantic aggregation and temporal dimension aggregation are performed to obtain the final feature representation. A node importance representation for fault localization is obtained, which acts as a spatial attention mask on the final feature representation to achieve fault-sensitive area localization of key components in robot holographic compliant assembly.

[0014] Based on the joint evaluation of attention statistics and information gain, redundant connections and redundant nodes are dynamically sparsified to obtain a sparsified graph representation.

[0015] The sparsed graph representation is used as input to the classification and decision layer to predict the health status category, and the model with the optimal parameters is saved for health diagnosis of robot holographic compliant assembly.

[0016] in:

[0017] ,

[0018] In the formula, Indicates the first Within the first sliding window Sensor signals are targeted The basic probability assignment function, Indicates from the first Within the first sliding window Temporal statistical feature vectors extracted from sensor signals; , Preset fault mode hypothesis set In , j One failure mode, For distance measurement, This is the uncertainty adjustment coefficient.

[0019] The preferred technical solution is as follows:

[0020] The modified The degree of conflict between acoustic and vibration evidence is calculated, and fuzzy entropy is introduced to constrain and optimize the DS evidence combination process, resulting in window-level fused evidence, including:

[0021] Calculate the degree of conflict:

[0022]

[0023] In the formula, For the first The degree of conflict between acoustic and vibrational evidence within a sliding window; and For the first Acoustic sensor within a sliding window Vibration sensor Failure modes , The corrected basic probability assignment value;

[0024] The fault mode hypothesis set is calculated based on the standard DS combination rule. A specific fault mode The initial fusion result, i.e., the initial fusion basic probability assignment value. :

[0025]

[0026] The fuzzy entropy, representing the uncertainty of the initial fusion result, is calculated based on the initial fusion result. :

[0027]

[0028] Using the fuzzy entropy The evidence combination process is adjusted to obtain the final window-level fused evidence. :

[0029] .

[0030] The confidence levels of the acoustic and vibration evidence within the same sliding window are calculated, and the fusion weights are adaptively determined and adjusted based on these confidence levels. ,include:

[0031] Calculate the confidence level of the evidence corresponding to the sensor for each sliding window:

[0032]

[0033] In the formula, No. Within the first sliding window The confidence level of evidence corresponding to sensor-like devices. To prevent constants with a denominator of zero; the first Within the first sliding window Shannon entropy of sensor-like evidence , The first characteristic distribution of the signal within the sliding window A probability term, N The total number of probability terms;

[0034] The window-level fusion weights are adaptively updated based on the confidence level. :

[0035]

[0036] In the formula, Indicates the first The fusion weights of evidence corresponding to the s-th type of sensor within a sliding window; An index variable for iterating through sensor types, used to sum over all sensors in the denominator; and These represent acoustic sensors and vibration sensors, respectively.

[0037] use right After correction ,use replace Participate in the calculation of the conflict degree and window-level fusion evidence.

[0038] The method utilizes graph convolution and structural consistency constraints to perform local temporal modeling of node temporal features, calculates dynamic similarity, and performs spatial semantic aggregation and temporal dimension aggregation, including:

[0039] Local temporal modeling is performed on the temporal features of nodes to obtain node... In the time window Local temporal coding features within ;

[0040] Calculate the time window internal nodes With nodes Dynamic similarity between ;

[0041] based on Obtain the joint edge weights of a graph neural network ,in Represents a node With nodes Structural consistency regularization terms between them;

[0042] Spatial semantic aggregation is performed based on the aforementioned joint edge weights:

[0043]

[0044] In the formula, For nodes In the time window Features derived from fine-grained spatial semantic aggregation; Represents a node The set of local neighbor nodes; and Both involve traversing the set of neighbor nodes. index variables; This represents the learnable feature transformation weight matrix; This represents the attention weight after normalization of the joint edge weights;

[0045] Based on the features obtained from the fine-grained spatial semantic aggregation, a time-dimensional weighted aggregation is performed to obtain nodes. The final feature representation after spatiotemporal aggregation , The total number of time windows. Indicates the first Aggregated weights for each time window.

[0046] The obtained node importance representation for fault localization, which acts as a spatial attention mask on the final feature representation, includes:

[0047] Computation node importance score:

[0048]

[0049] in, Represents a node Importance score; the importance score integrates all time windows. internal nodes With all its neighboring nodes Joint border rights between and the feature strength of neighboring nodes ;

[0050] The node importance score is used as a spatial attention mask to apply to the final feature representation, resulting in the enhanced final feature representation. .

[0051] The dynamic similarity is calculated using cosine similarity, which reflects the synchronicity between sound and vibration evidence.

[0052] The local temporal coding features are obtained by using a one-dimensional convolution operation.

[0053] The dynamic sparsification of redundant connections and redundant nodes based on joint evaluation of attention statistics and information gain includes:

[0054] The edge attention score is calculated at each layer of the graph neural network;

[0055] The attention weights are obtained by normalizing the edge attention scores;

[0056] Calculate the average attention weight across multiple time windows;

[0057] Calculate the information gain of nodes to quantify their contribution:

[0058] Construct asymptotic gating functions;

[0059] Based on the average attention and the information gain, redundant edge connections and redundant nodes are gradually weakened or removed. The redundancy removal intensity is gradually enhanced through the asymptotic gating function, and finally a sparse topological representation is obtained.

[0060] The process of inputting the sparsed graph representation into the classification and the decision layer outputting a health status category prediction includes:

[0061] Set threshold To achieve rejection decision-making:

[0062]

[0063] in, This represents the final diagnostic decision result output by the model; This represents the target candidate category with the highest predicted probability among all categories, which is the classification result that the model most favors. This represents the fusion graph features ultimately used for classification decisions; Indicates that, given the fusion graph features Time prediction as category The conditional probability value, y For true category variables; This represents the preset confidence threshold. Indicates a state of rejection;

[0064] When the highest predicted probability is still below the set threshold At that time, the model output It avoids the risk of misjudgment caused by strong noise or unknown new faults by refusing to identify.

[0065] The process of inputting the sparsed graph representation into the classification and the decision layer outputting a health status category prediction includes:

[0066] The probability of each fault mode is output through a fully connected layer and a Softmax function.

[0067] The technical solution of the present invention can achieve at least the following beneficial effects:

[0068] This invention achieves adaptive quantification and conflict suppression fusion of multi-source evidence credibility through dynamic DS evidence reasoning, and combines graph neural networks to model the topological relationships between fused evidence. Simultaneously, it introduces a progressive redundancy removal strategy to achieve structured representation learning from dense to sparse data. This effectively reduces the uncertainty of multi-source data, improves the expression and identification of weak fault features under varying operating conditions, and thus enhances the accuracy and robustness of health diagnosis for robotic holographic compliant assembly systems. Specifically, it has the following advantages:

[0069] (1) This invention achieves adaptive weighting and conflict-suppressible fusion of multi-source acoustic-vibration evidence through dynamic DS evidence reasoning driven by sliding window confidence, thereby improving the fusion stability and recognition robustness under complex working conditions and noise disturbances.

[0070] (2) Based on the fusion evidence to construct a graph structure and introduce a fine-grained spatiotemporal feature extraction mechanism, this invention can simultaneously characterize the dynamic correlation of signals and structural consistency, thereby enhancing the ability to express the fault evolution dependency of key components in robot holographic compliant assembly and improving the sensitivity of fault location.

[0071] (3) This invention proposes a node-edge progressive redundancy removal mechanism, which achieves dynamic sparse modeling through joint evaluation of attention statistics and information gain, effectively suppresses noise propagation caused by redundant connections and reduces computational redundancy, while improving the adaptability to unknown fault modes by combining rejection decision.

[0072] Other features and advantages of the invention will be set forth in the following description or may be learned by practicing the invention. Attached Figure Description

[0073] Figure 1 This is a schematic diagram of the process framework of the method in an embodiment of the present invention.

[0074] Figure 2 This is a time-domain waveform diagram of the multi-source sensing signal in an embodiment of the present invention.

[0075] Figure 3 This is a schematic diagram of the window-level fusion mechanism for dynamic DS evidence reasoning in an embodiment of the present invention.

[0076] Figure 4 This is a network structure diagram of topology modeling and fine-grained spatiotemporal feature extraction based on graph neural networks, as described in an embodiment of the present invention.

[0077] Figure 5 This is a visualization result (t-SNE) of the two-dimensional feature distribution of the method in this embodiment of the invention under different assembly offset levels.

[0078] Figure 6 for Figure 1A schematic diagram of the process framework for specific implementation examples of experimental verification and result visualization analysis. Detailed Implementation

[0079] The specific embodiments of the present invention are described below with reference to the accompanying drawings.

[0080] This embodiment provides a deep learning-based method for holographic compliant assembly health diagnosis of robots. The holographic compliant assembly refers to recording and interacting with the complex dynamic load energy field experienced by the robot during assembly operations through multiple channels, relying on a holographic dynamic load spectrum. The multi-channel information at least characterizes the variation patterns of characteristic vectors such as load amplitude, direction, frequency, phase, and density. Based on this, an assembly process performance index model is constructed to guide the compliant assembly of key core functional components and the entire machine, serving its full-cycle dynamic load service process. In this embodiment, the multi-channel information can consist of acoustic signals, vibration signals, and / or sensor information such as force / torque and displacement, wherein acoustic and vibration signals are used to characterize the dynamic load energy field and its evolution.

[0081] See Figure 1 This embodiment presents a method for health diagnosis of key components in a robot holographic compliant assembly system based on multi-source acoustic-vibration information fusion, dynamic DS evidence reasoning, and graph neural network topology modeling. The method specifically includes the following steps:

[0082] S1. Collect multi-source sensor signals that characterize the assembly process state information under the holographic dynamic load spectrum during the robot's holographic compliant assembly process, including acoustic signals and vibration signals.

[0083] In one specific embodiment, the acoustic signal and vibration signal are synchronously acquired under the same operating condition of the robot's holographic compliant assembly task. The sensor is arranged near the end effector, assembly fixture, or joint transmission. Vibration signals collected by sound pressure sensors The data is collected by an accelerometer.

[0084] See Figure 2 This is a time-domain waveform diagram of a multi-source sensing signal in a specific embodiment. Figure 2 In the figure, (a) and (b) are time-domain waveforms of vibration and acoustic signals under a certain assembly offset state (e.g., 10 mm).

[0085] S2. Normalize the acoustic signal and vibration signal, and use a sliding window to segment them to obtain the acoustic window sequence and vibration window sequence corresponding to each sliding window.

[0086] By standardizing and normalizing the acoustic and vibration signals to eliminate the impact of differences in the dimensions and amplitudes of different sensors on subsequent fusion and modeling, the normalization is specifically performed according to the following formula:

[0087]

[0088] in, , These represent the mean and standard deviation of the acoustic signal, respectively. , These are the mean and standard deviation of the vibration signal, respectively.

[0089] In one specific implementation, the normalized signal is divided into overlapping sliding windows of length L and step size Δ. The acoustic window sequence and vibration window sequence are implemented using the following formula:

[0090]

[0091] in, , Indicates the first The acoustic window sequence and vibration window sequence corresponding to each sliding window; The index of the sliding window is a positive integer. This is the sampling time step index or the starting sampling time point for the normalized signal.

[0092] S3. Extract features from each of the acoustic window sequences and vibration window sequences and map them to the basic probability allocation function. To construct acoustic and vibration evidence for corresponding windows; calculate the confidence levels of the acoustic and vibration evidence within the same sliding window; determine the fusion weights based on the confidence levels and adjust accordingly. Based on the revised The degree of conflict between acoustic and vibration evidence is calculated, and fuzzy entropy is introduced to constrain and optimize the DS evidence combination process, resulting in window-level fused evidence.

[0093] In one specific implementation, the following steps are included:

[0094] S31. Constructing the Failure Mode Hypothesis Set That is, the set of all possible failure modes, extracting the time-domain statistical feature vectors of each sliding window, and mapping the window features to the basic probability allocation function (BPA):

[0095] ,

[0096] In the formula, Indicates the first Within the first sliding window Sensor signals for fault modes The basic probability assignment function is used to quantify the probability distribution of each fault mode. Support level; Indicates from the first Within the first sliding window The time-domain statistical feature vector extracted from sensor signals utilizes a vector composed of statistical features such as root mean square and kurtosis. , Preset fault mode hypothesis set In , j One failure mode, j The index variable is used to iterate through the set of failure mode hypotheses, and is used to sum all failure modes in the denominator; For distance measurement, This is an uncertainty adjustment coefficient used to control the sensitivity of the distance metric to similarity calculation.

[0097] S32. Calculate the confidence levels of the acoustic and vibration evidence within the same sliding window, adaptively determine the fusion weights based on the confidence levels, and adjust accordingly. ,include:

[0098] For each window, the confidence level of the evidence corresponding to the sensor is calculated. This confidence level characterizes the stability of the evidence within that window; a higher confidence level indicates more stable and reliable evidence. In one specific embodiment, the confidence level is constructed based on the reciprocal of the signal's Shannon entropy.

[0099]

[0100] In the formula, For the first Within the first sliding window The confidence level of evidence corresponding to sensor-like devices. To prevent constants with a denominator of zero; the first Within the first sliding window Shannon entropy of sensor-like evidence , The first characteristic distribution of the signal within the sliding window A probability term, N The total number of probability terms;

[0101] The window-level fusion weights are adaptively updated based on the confidence level. This allows highly credible evidence to play a larger role in the fusion process. The update rules are as follows:

[0102]

[0103] In the formula, Indicates the first The fusion weight of the evidence corresponding to the s-th type of sensor within a sliding window is used to achieve dynamic weighting of multi-source evidence. An index variable for iterating through sensor types, used to sum over all sensors in the denominator; and These represent acoustic sensors and vibration sensors, respectively.

[0104] use right Evidence weighting corrections are performed to enhance the dominant role of high-confidence sensor evidence in fusion, resulting in... ,use replace Participate in the calculation of the conflict degree and window-level fusion evidence. And, obtain... , This represents the uncertainty assigned to the global hypothesis set after correction.

[0105] S33. Based on the revised The degree of conflict between acoustic and vibration evidence is calculated, and fuzzy entropy is introduced to constrain and optimize the DS evidence combination process, resulting in window-level fused evidence, including:

[0106] When multiple sources of evidence inconsistently support the mutually exclusive hypothesis, calculate the degree of conflict of evidence:

[0107]

[0108] In the formula, For the first The degree of conflict between acoustic and vibrational evidence within a sliding window; and For the first Acoustic sensor within a sliding window Vibration sensor Failure modes , The modified value of the basic probability assignment function;

[0109] The fault mode hypothesis set is calculated based on the standard DS combination rule. A specific fault mode The initial fusion result, i.e., the initial fusion basic probability assignment value. :

[0110]

[0111] The fuzzy entropy, representing the uncertainty of the initial fusion result, is calculated based on the initial fusion result. :

[0112]

[0113] Using the fuzzy entropy The evidence combination process is adjusted to obtain the final window-level fused evidence. :

[0114] .

[0115] By introducing fuzzy entropy to constrain and optimize the DS evidence combination process, the instability of evidence combination in strong conflict scenarios can be suppressed.

[0116] See Figure 3 This is a schematic diagram illustrating a specific implementation of the window-level fusion mechanism for dynamic DS evidence reasoning described in step S33 above. This mechanism includes confidence calculation, weight update, conflict degree calculation, and fuzzy entropy constraint fusion processes, thereby realizing the construction of the dynamic DS evidence reasoning module. The input to the dynamic DS evidence reasoning module consists of acoustic and vibration evidence within the same sliding window, and the output is window-level fused evidence.

[0117] S4. Using the window-level fusion evidence as node features, construct a graph, determine the node set and edge weights, initialize and embed the node features, map the window-level fusion evidence to a learnable graph representation space, and realize graph neural network topology modeling.

[0118] In one specific embodiment, a node set is constructed. Each node corresponds to a sliding window of fused evidence representation, which is represented as follows:

[0119]

[0120] in, Indicates the first One node; express 3D real space, i.e. Representing the features of each node All are of one dimension eigenvectors.

[0121] Calculate the edge weights between nodes The edge weights are determined by the following formula after characterizing and normalizing the node feature similarity:

[0122]

[0123] in, This is a scale parameter used to control the sensitivity of feature similarity calculation; This is the index variable for traversing other nodes in the graph. This indicates that the set excludes nodes. Sum and normalize all nodes except those in the above order;

[0124] For each node, a feature initialization embedding is performed to map the fused evidence to a learnable graph representation space. The initialization embedding is obtained by the following formula:

[0125]

[0126] in, superscript Indicates the index of the node or the corresponding sliding window. This represents the fused data output through DS evidence reasoning; The feature mapping weight matrix, The bias vector is used to map low-dimensional evidence vectors to a high-dimensional feature space so that graph neural networks can perform feature mining.

[0127] S5. Construct a graph-based fine-grained spatiotemporal feature extraction mechanism: Utilize graph convolution and structural consistency constraints to perform local temporal modeling of node temporal features, calculate dynamic similarity, perform spatial semantic aggregation and temporal dimension aggregation, and obtain node importance representation for fault localization, which is used to locate fault-sensitive areas of key components in robot holographic compliant assembly.

[0128] The fine-grained spatiotemporal feature extraction mechanism is determined by both signal dynamic similarity and structural consistency, and specifically includes the following steps:

[0129] S51. Perform local temporal modeling on the temporal characteristics of nodes to obtain node... In the time window Local temporal coding features within .

[0130] In one specific embodiment, the local temporal coding features are obtained using a one-dimensional convolution operation:

[0131]

[0132] In the formula, node In the time window Local temporal coding features within, This represents a one-dimensional convolution operation used for temporal feature extraction, which is used to capture periodically changing features; Represents a node Corresponding from time arrive The input original signal segment; Indicates the length of the time window.

[0133] S52. Calculate the time window internal nodes With nodes Dynamic similarity between .

[0134] In one specific embodiment, the dynamic similarity is calculated using cosine similarity, which reflects the synchronicity between acoustic and vibration evidence:

[0135]

[0136] In the formula, Indicates time window internal nodes With nodes The dynamic similarity between them, namely cosine similarity, reflects the synchronicity of sound and vibration characteristics; Representing neighboring nodes In the time window Temporal encoding features within; superscript This represents the vector transpose operation; The L2 norm of the eigenvector is represented by the vector length.

[0137] S53. Based on Obtain the joint edge weights of a graph neural network that combines dynamic similarity and structural consistency. ,in Represents a node With nodes The structural consistency regularization term between them. Specifically:

[0138]

[0139] In the formula, Represents a node With nodes Structural consistency regularization terms between them; Represents a node With nodes The topological distance between them, i.e., the time interval at the center of the sliding window; This represents the scale factor used to adjust for distance attenuation.

[0140] S54. Perform spatial semantic aggregation based on the joint edge weights:

[0141]

[0142] In the formula, For nodes In the time window Features derived from fine-grained spatial semantic aggregation; Represents a node The set of local neighbor nodes; and Both involve traversing the set of neighbor nodes. index variables; This represents the learnable feature transformation weight matrix; This represents the attention weight after normalization of the joint edge weights;

[0143] S55. Based on the features obtained after fine-grained spatial semantic aggregation, perform time-dimensional weighted aggregation to obtain nodes. The final characteristics after spatiotemporal aggregation , The total number of time windows. Indicates the first The aggregate weights for each time window can be adaptively determined using statistical rate of change.

[0144] S56. Obtain node importance representations for fault localization to locate fault-sensitive areas of key components in the robot's holographic compliant assembly, including:

[0145] Computation node importance score:

[0146]

[0147] in, Represents a node Importance score; the importance score integrates all time windows. internal nodes With all its neighboring nodes Joint border rights between and the feature strength of neighboring nodes ;

[0148] The node importance score is used as a spatial attention mask to apply to the final feature representation, resulting in the enhanced final feature representation. This leads to an enhanced ability to express oneself.

[0149] Through steps S4 and S5 above, a graph neural network-based topology modeling and fine-grained spatiotemporal feature extraction network was constructed. Its structure diagram is shown below. Figure 4 .like Figure 4 As shown, this embodiment utilizes a graph neural network to aggregate node neighborhood information to extract spatiotemporal correlation features. In the spatial dimension, spatial semantic aggregation is employed within the current time window. Inner nodes local neighbor node set The features are weighted and aggregated. Spatial semantic aggregation can adaptively enhance effective features highly correlated with the central node and suppress irrelevant noise based on the dynamic similarity between nodes. In the time dimension, time-weighted aggregation is used to aggregate nodes across multiple time windows (total). Spatial aggregation features extracted within (number) A time-dimensional fusion is performed. The time-weighted aggregation further captures the global dynamic dependencies of node states as they evolve over time during the assembly process.

[0150] S6. Construct an incremental redundancy removal mechanism: Dynamically sparsify redundant connections and redundant nodes based on joint evaluation of attention statistics and information gain to obtain a sparsified graph representation.

[0151] This embodiment employs a progressive redundancy removal mechanism to gradually weaken or remove redundant edge connections and redundant nodes, achieving a smooth transition from dense to sparse graph structure. This results in a sparse yet efficient topological representation, reducing noise interference caused by redundancy propagation and improving the characterization capability of sensitive regions of key components. Specifically, it includes the following steps:

[0152] S61. Calculate edge attention scores at each layer of the graph neural network:

[0153]

[0154] in, Indicates the first Layered graph networks in time windows Next node With nodes Attention score between edges; For activation functions; This is the transpose of the learnable attention weight vector; The feature weight matrix; and They are nodes With nodes In the The feature vector of the layer; This indicates a vector concatenation operation.

[0155] S62. The attention weights are obtained by scoring and normalizing the edge attention:

[0156]

[0157] in, This represents the normalized attention weights; For nodes The set of neighboring nodes; This is the index variable for traversing the nodes in the neighbor set.

[0158] S63. Calculate the average attention weight over multiple time windows:

[0159]

[0160] in, Represents a node With nodes The average attention weight of the edges between them over multiple time windows; This represents the total number of statistical sliding time windows.

[0161] S64. Calculate node information gain to quantify node contribution:

[0162]

[0163] in, Represents a node The information gain is used to quantify the contribution of the node's features to fault classification; Health status label Information entropy; Representing the features of a given node Conditional entropy under given conditions.

[0164] S65. Constructing asymptotically gated functions:

[0165]

[0166] in, This represents the asymptotic gating function, i.e., the sparsification threshold coefficient; This represents the current iteration number. This represents the total number of iterations or a preset time constant. An exponential factor to control the sparsification rate.

[0167] S66. Based on the average attention and the information gain, redundant edge connections and redundant nodes are gradually weakened or removed. The redundancy removal intensity is gradually enhanced through the asymptotic gating function, and finally a sparse topological representation is obtained.

[0168] S7. Input the sparsed graph representation into the classification and decision layer outputs the health status category prediction, and save the model with the optimal parameters for robot holographic compliant assembly health diagnosis.

[0169] In one specific embodiment, the step of inputting the sparsed graph representation into classification and the decision layer outputting a health status category prediction includes:

[0170] The probabilities of each failure mode are output through a fully connected layer and a Softmax function:

[0171]

[0172] in, Indicates the feature representation of a given graph. Under the condition that the sample belongs to the first Predicted probability of failure modes; This represents the global graph feature vector obtained after extraction and sparsification by a graph neural network. and These represent the corresponding numbers in the fully connected layer. The transpose and bias terms of the class's learnable weight vector; This represents the total number of known failure mode categories; This is an index variable for iterating through all categories.

[0173] Set threshold To achieve rejection decision-making:

[0174]

[0175] in, This represents the final diagnostic decision result output by the model; This represents the target candidate category with the highest predicted probability among all categories, which is the classification result that the model most favors. This represents the fusion graph features ultimately used for classification decisions; Indicates that, given the fusion graph features Time prediction as category The conditional probability value, y For true category variables; This represents the preset confidence threshold. Indicates a state of rejection;

[0176] When the highest predicted probability is still below the set threshold At that time, the model output It avoids the risk of misjudgment caused by strong noise or unknown new faults by refusing to identify.

[0177] To verify the feasibility and effectiveness of the method of the present invention, a robot holographic compliant assembly abnormality health diagnosis experimental platform was constructed in a specific embodiment. Assembly offset levels (0, 5, 10, 15, 20 mm) were set to form different health states, and acoustic and vibration signals were collected simultaneously to form a sample set. The experimental data collection and parameter settings are shown in Table 1.

[0178] Table 1 Experimental Data Acquisition and Parameter Settings

[0179]

[0180] First, according to the settings shown in Table 1, acoustic and vibration signals are simultaneously acquired under stable operation conditions of the robot's holographic compliant assembly system, and the acquired signals are standardized and preprocessed. Then, a sliding window is used to segment the signals, forming window-level samples. Next, acoustic and vibration evidence are constructed for each window, the window-level confidence is calculated, and the fusion weights are adaptively updated. Considering evidence conflict, fuzzy entropy constraints are introduced for evidence combination to obtain window-level fused evidence. Then, the fused evidence from multiple windows is used as node features to construct a graph structure. A graph neural network is used to perform topological modeling of the relationships between nodes and extract fine-grained spatiotemporal features. During training, a progressive redundancy removal mechanism is used to gradually weaken or prune redundant edge connections and redundant nodes, causing the graph structure to gradually transition from dense to sparse, thereby reducing the noise impact of redundancy propagation and improving the effective feature representation. Finally, the sparsed graph representation is input into the classification and decision layer to obtain the predicted health status category results, and the model with the optimal parameters is saved for subsequent testing or online recognition.

[0181] In this embodiment, the collected samples can be divided into training set and test set according to a preset ratio, and a unified evaluation index can be used to evaluate the recognition performance. Figure 5 This is a two-dimensional visualization (t-SNE) of the multi-source signals from the assembly process at different assembly offset levels after feature extraction. Figure 5 As can be seen, the samples corresponding to the offset levels of different health states form relatively independent clustering regions in the feature space. The inter-class separation is obvious and the intra-class distribution is more compact. This indicates that the method of the present invention can extract stable and discriminative state features under the holographic dynamic load spectrum characterization, and has high accuracy in health diagnosis and noise resistance.

[0182] Furthermore, the effectiveness of each core module proposed in this invention under different noise environments was verified through ablation experiments and weight visualization. For the specific implementation framework of the experimental verification and result visualization analysis, please refer to [link to relevant documentation]. Figure 6 . Figure 6 The middle top Figure 5 The t-SNE shown, Figure 6 The middle section is a performance comparison line graph. In (a), (b), and (c), the vertical axis represents the model's diagnostic accuracy, F1 score, and area under the curve (AUC), respectively. The horizontal axis represents the signal-to-noise ratio environment (e.g., -6 dB, 0 dB, 4 dB) to simulate different noise interference intensities in actual industrial environments. The legend "Ours" represents the complete method proposed in this invention; "-DD-SRM" represents a variant model with the dynamic DS evidence reasoning module removed; "-GFEM" represents a variant model with the fine-grained spatiotemporal feature extraction mechanism removed; and "-RRMS" represents a variant model with the progressive redundancy removal mechanism removed.

[0183] Depend on Figure 6 As can be seen, the complete model (Ours) achieved the highest accuracy, F1 score, and AUC value under all signal-to-noise ratio conditions. Particularly noteworthy is that removing any one of the core modules in a strong noise environment of -6 dB leads to a significant performance degradation. This demonstrates that the three core modules in this invention work synergistically, making indispensable technical contributions to improving system noise immunity, resolving evidence conflicts, and accurately locating fault characteristics.

[0184] Figure 6 The bar chart at the bottom is the final layer weight distribution diagram. The horizontal axis represents the different feature components input to the classification decision layer, and the vertical axis represents the weight allocation value corresponding to each feature component. This figure shows the attention distribution of the model in this invention when making a health status diagnosis. Figure 6 It can be seen that the weight distribution exhibits a highly concentrated characteristic. This indicates that after processing with the graph neural network and sparsification mechanism proposed in this invention, the model can accurately focus on the few core features most sensitive to the failure of key components, effectively suppressing the interference of redundant and useless features, and proving the model's efficiency and discriminative ability in complex assembly conditions.

[0185] It will be understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A deep learning-based method for holographic compliant assembly health diagnosis of robots, characterized in that, include: The system collects multi-source sensor signals that characterize the assembly process state information under the holographic dynamic load spectrum during the holographic compliant assembly process of the robot. These signals include acoustic and vibration signals collected by the corresponding sensors under the same operating conditions. The acoustic and vibration signals are normalized, and a sliding window segmentation is used to obtain the acoustic window sequence and vibration window sequence corresponding to each sliding window. For each of the acoustic window sequences and vibration window sequences, features are extracted and mapped to a basic probability assignment function. In order to construct acoustic and vibration evidence corresponding to the window; Calculate the confidence levels of the acoustic and vibration evidence within the same sliding window, determine the fusion weights based on the confidence levels, and then adjust accordingly. Based on the revised The degree of conflict between acoustic and vibration evidence is calculated, and fuzzy entropy is introduced to constrain and optimize the DS evidence combination process to obtain window-level fused evidence. Graph construction is performed using the window-level fused evidence as node features to determine the node set and edge weights, and the node features are initialized and embedded. By leveraging graph convolution and structural consistency constraints, local temporal modeling of node temporal features is performed, dynamic similarity is calculated, and spatial semantic aggregation and temporal dimension aggregation are conducted to obtain the final feature representation. The node importance representation for fault localization is obtained, which acts as a spatial attention mask on the final feature representation to realize the fault-sensitive area localization of key components in the robot's holographic compliant assembly. Based on the joint evaluation of attention statistics and information gain, redundant connections and redundant nodes are dynamically sparsified to obtain a sparsified graph representation. The sparsed graph representation is used as input to the classification and decision layer to predict the health status category, and the model with the optimal parameters is saved for health diagnosis of robot holographic compliant assembly. in: , In the formula, Indicates the first Within the first sliding window Sensor signals are targeted The basic probability assignment function, Indicates from the first Within the first sliding window Temporal statistical feature vectors extracted from sensor signals; , Preset fault mode hypothesis set In , j One failure mode, For distance measurement, This is the uncertainty adjustment coefficient.

2. The method according to claim 1, characterized in that, The basis is the revised The degree of conflict between acoustic and vibration evidence is calculated, and fuzzy entropy is introduced to constrain and optimize the DS evidence combination process, resulting in window-level fused evidence, including: Calculate the degree of conflict: , In the formula, For the first The degree of conflict between acoustic and vibrational evidence within a sliding window; and For the first Acoustic sensor within a sliding window Vibration sensor Failure modes , The corrected basic probability assignment value; The fault mode hypothesis set is calculated based on the standard DS combination rule. A specific fault mode The initial fusion result, i.e., the initial fusion basic probability assignment value. : , The fuzzy entropy, representing the uncertainty of the initial fusion result, is calculated based on the initial fusion result. : , Using the fuzzy entropy The evidence combination process is adjusted to obtain the final window-level fused evidence. : 。 3. The method according to claim 1, characterized in that, The confidence levels of the acoustic and vibration evidence within the same sliding window are calculated, and the fusion weights are adaptively determined and adjusted based on these confidence levels. ,include: Calculate the confidence level of the evidence corresponding to the sensor for each sliding window: In the formula, No. Within the first sliding window The confidence level of evidence corresponding to sensor-like devices. To prevent constants with a denominator of zero; the first Within the first sliding window Shannon entropy of sensor-like evidence , The first characteristic distribution of the signal within the sliding window A probability term, N The total number of probability terms; The window-level fusion weights are adaptively updated based on the confidence level. : , In the formula, Indicates the first The fusion weights of evidence corresponding to the s-th type of sensor within a sliding window; An index variable for iterating through sensor types, used to sum over all sensors in the denominator; and These represent acoustic sensors and vibration sensors, respectively. use right After correction ,use replace Participate in the calculation of the conflict degree and window-level fusion evidence.

4. The method according to claim 1, characterized in that, The method utilizes graph convolution and structural consistency constraints to perform local temporal modeling of node temporal features, calculates dynamic similarity, and performs spatial semantic aggregation and temporal dimension aggregation, including: Local temporal modeling is performed on the temporal features of nodes to obtain node... In the time window Local temporal coding features within ; Calculate the time window internal nodes With nodes Dynamic similarity between ; based on Obtain the joint edge weights of a graph neural network ,in Represents a node With nodes Structural consistency regularization terms between them; Spatial semantic aggregation is performed based on the aforementioned joint edge weights: In the formula, For nodes In the time window Features derived from fine-grained spatial semantic aggregation; Represents a node The set of local neighbor nodes; and Both involve traversing the set of neighbor nodes. index variables; This represents the learnable feature transformation weight matrix; This represents the attention weight after normalization of the joint edge weights; Based on the features obtained after fine-grained spatial semantic aggregation, time-dimensional weighted aggregation is performed to obtain nodes. The final feature representation after spatiotemporal aggregation , The total number of time windows. Indicates the first Aggregated weights for each time window.

5. The method according to claim 4, characterized in that, The obtained node importance representation for fault localization, which acts as a spatial attention mask on the final feature representation, includes: Computation node importance score: , in, Represents a node Importance score; the importance score integrates all time windows. internal nodes With all its neighboring nodes Joint border rights between and the feature strength of neighboring nodes ; The node importance score is used as a spatial attention mask to apply to the final feature representation, resulting in the enhanced final feature representation. .

6. The method according to claim 4, characterized in that, The dynamic similarity is calculated using cosine similarity, which reflects the synchronicity between sound and vibration evidence.

7. The method according to claim 4, characterized in that, The local temporal coding features are obtained by using a one-dimensional convolution operation.

8. The method according to claim 1, characterized in that, The dynamic sparsification of redundant connections and redundant nodes based on joint evaluation of attention statistics and information gain includes: The edge attention score is calculated at each layer of the graph neural network; The attention weights are obtained by normalizing the edge attention scores; Calculate the average attention weight across multiple time windows; Calculate the information gain of nodes to quantify their contribution: Construct asymptotic gating functions; Based on the average attention and the information gain, redundant edge connections and redundant nodes are gradually weakened or removed. The redundancy removal intensity is gradually enhanced through the asymptotic gating function, and finally a sparse topological representation is obtained.

9. The method according to claim 1, characterized in that, The process of inputting the sparsed graph representation into the classification and the decision layer outputting a health status category prediction includes: Set threshold To achieve rejection decision-making: , in, This represents the final diagnostic decision result output by the model; This represents the target candidate category with the highest predicted probability among all categories, which is the classification result that the model most favors. This represents the fusion graph features ultimately used for classification decisions; Indicates that, given the fusion graph features Time prediction as category The conditional probability value, y For true category variables; This represents the preset confidence threshold. Indicates a state of rejection; When the highest predicted probability is still below the set threshold At that time, the model output It avoids the risk of misjudgment caused by strong noise or unknown new faults by refusing to identify.

10. The method according to claim 1, characterized in that, The process of inputting the sparsed graph representation into the classification and the decision layer outputting a health status category prediction includes: The probability of each fault mode is output through a fully connected layer and a Softmax function.