Wire harness assembly robot control method and system based on machine learning
By constructing a contact event tensor and decomposing the signal using a sparse autoencoder model, combined with sparse vector judgment and cluster matching, valid contact events are identified and responded to, thus solving the problem of signal misjudgment in robot wire harness assembly and improving assembly accuracy and efficiency.
Patent Information
- Application Number
- CN202511770925.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-11-28
AI Technical Summary
In the current robot wiring harness assembly process, the amplitude of the signal and the interference signal are similar, which leads to misjudgment by the controller and affects the assembly accuracy and efficiency. Traditional control methods lack dynamic adaptation capabilities, making it difficult to separate the effective contact signal from the interference noise, and thus failing to improve the assembly quality.
A machine learning-based approach is adopted to construct a contact event tensor by real-time acquisition of multi-dimensional sensor data. A sparse autoencoder model is used to decompose the signal into sparse components and residual components. By combining the L1 norm judgment and cluster matching of sparse vectors, valid contact events are identified, and a predefined fine-tuning control strategy is invoked.
It enables accurate identification and response to subtle contact events, improving the automation and intelligence level of assembly robots and enhancing assembly accuracy and efficiency.
Smart Images

Figure CN121315980A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of robot control, and particularly relates to a wire harness assembly robot control method and system based on machine learning. BACKGROUND
[0002] In the robot wire harness assembly operation, the wire harness plugging process is often accompanied by a large number of weak, short-time and non-standardized contact phenomena. The signal generated thereby is similar in amplitude to the interference signal of wire harness shaking and structure resonance, and is mixed with each other, which easily leads to misjudgment of the robot controller and affects the assembly precision and efficiency.
[0003] Existing control methods mostly rely on the "threshold + filtering" mechanism, and can only distinguish based on signal amplitude, which is difficult to effectively separate real contact signals and interference noises of the same amplitude. At the same time, the traditional control strategy lacks dynamic adaptation ability, and is not robust enough in the face of non-standardized micro-contact scenarios of flexible wire harness, and cannot provide reliable basis for assembly quality optimization. In addition, the existing scheme does not construct an efficient feature extraction model for contact event characteristics, which is difficult to accurately capture effective contact information from complex signals, resulting in delayed or mis-triggered control actions, which restricts the improvement of the automation and intelligence level of the wire harness assembly robot. SUMMARY
[0004] To this end, the present application provides a wire harness assembly robot control method and system based on machine learning to solve the above technical problems.
[0005] According to one aspect of the present application, a wire harness assembly robot control method based on machine learning is provided, comprising the following method steps: S1, real-time acquisition of multi-dimensional sensor data of the robot during execution of the wire harness plugging task, and cutting of the multi-dimensional sensor data by a sliding time window to construct a multi-channel contact event tensor, wherein the multi-dimensional sensor data at least includes joint torque, end effector position error and end effector attitude change; S2, inputting the contact event tensor into a pre-trained sparse autoencoder model, the encoder of the sparse autoencoder model encoding the contact event tensor into a sparse vector, the decoder thereof reconstructing a reconstructed tensor based on the sparse vector, and decomposing the input signal into a sparse component corresponding to an effective contact event and a residual component corresponding to interference noise by comparing the contact event tensor with the reconstructed tensor; S3, based on the sparse vector, identifying whether an effective contact event with control significance occurs in the current time window, specifically, by judging the activation pattern of the sparse vector, including judging whether its L1 norm exceeds a first threshold; S4, when an effective contact event is identified, the robot controller invokes a predefined fine control strategy for the event type to execute an action response, which includes adjusting the insertion speed, applying a fine angle, or applying a constant fine pushing force.
[0006] Preferably, in step S1, the multi-dimensional sensor data is pre-processed, including removing the DC offset and normalizing the data for each channel, and / or the contact event tensor further contains first-order difference features and / or second-order difference features calculated from the raw sensor data.
[0007] Preferably, the sparse autoencoder model is trained in an unsupervised manner, and the loss function contains a reconstruction error term and a sparse regularization term, wherein the sparse regularization term is an L1 norm penalty on the sparse vector.
[0008] Preferably, before step S3, it further includes model semanticization, including, after the training of the sparse autoencoder model is completed, using a clustering algorithm to cluster the sparse vectors obtained from the training data set to form multiple clusters, and semantically labeling these clusters, marking the clusters corresponding to effective contact events as effective event types, and marking the clusters corresponding to various types of interference as noise.
[0009] Preferably, step S4 includes dividing the real-time obtained sparse vector into the semantically labeled clusters, thereby directly identifying the specific type of effective contact event, and the effective contact event type includes an entry scratch event, a slot damping event, or a bottom unlocking event.
[0010] Preferably, in step S5, a preset strategy mapping table is included, which defines the mapping relationship between different effective contact event types and specific fine control strategies, when an entry scratch event is identified, a strategy of deceleration and fine angle adjustment is executed, and when a slot damping event is identified, a strategy of applying a constant fine pushing force and correcting the pose is executed.
[0011] Preferably, the signal decomposition specifically includes that the sparse component is represented by a sparse reconstruction tensor obtained by decoding the sparse vector by the decoder, and the residual component is obtained by subtracting the sparse reconstruction tensor from the contact event tensor.
[0012] Preferably, in step S5, the micro-motions generated by the fine control strategy are superimposed on the nominal trajectory of the robot in an additive manner, and are gradually mixed through a smoothing coefficient.
[0013] Preferably, when the joint torque exceeds a safety threshold during the execution of the micro-motion, the current micro-motion is immediately interrupted and a safety stop is triggered.
[0014] In another aspect, the application provides a machine learning-based wire harness assembly robot control system, comprising: a data acquisition and construction module configured to acquire multi-dimensional sensor data of the robot in real time when performing a wire harness insertion task, to intercept the multi-dimensional sensor data in a sliding time window, and to construct a multi-channel contact event tensor, wherein the multi-dimensional sensor data at least includes joint torque, end effector position error, and end effector attitude change; a signal decomposition module configured to input the contact event tensor into a pre-trained sparse autoencoder model, an encoder of the sparse autoencoder model encoding the contact event tensor into a sparse vector, a decoder of the sparse autoencoder model reconstructing a reconstructed tensor based on the sparse vector, and decomposing the input signal into a sparse component corresponding to an effective contact event and a residual component corresponding to interference noise by comparing the contact event tensor with the reconstructed tensor; an event identification module configured to identify whether an effective contact event with control significance occurs in a current time window based on the sparse vector, specifically by judging an activation pattern of the sparse vector, including judging whether an L1 norm of the sparse vector exceeds a first threshold; a control execution module configured to, when an effective contact event is identified, invoke a fine-tuning control strategy predefined for the event type by a robot controller to execute an action response, the fine-tuning control strategy including adjusting an insertion speed, applying a fine-tuning angle, or applying a constant fine-tuning force.
[0015] The application integrates multi-dimensional sensor data by constructing a contact event tensor, realizes semantic decoupling of effective contact signals and interference noise by means of a pre-trained sparse autoencoder, breaks through the limitation of traditional "threshold + filtering" that only distinguishes signals by amplitude, realizes accurate identification of effective contact events by combining L1 norm judgment of the sparse vector and cluster matching, invokes a preset fine-tuning strategy and executes actions smoothly in an additive superposition manner, and has a torque overrun safety shutdown mechanism. The application improves the response accuracy and collaborative continuity of the wire harness assembly robot to weak contact events, and significantly improves the assembly automation and intelligence level. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the application, and therefore should not be considered as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor.
[0017] Other features, objects and advantages of the application will become more apparent from the following detailed description of the non-limiting embodiments, made with reference to the accompanying drawings: Figure 1A machine learning-based wire harness assembly robot control method flowchart is provided for the embodiments of the present application.
[0018] Figure 2 A sparse autoencoder model architecture diagram is provided for the embodiments of the present application.
[0019] Figure 3 A sparse autoencoder model training flowchart is provided for the embodiments of the present application.
[0020] Figure 4 A semantic flowchart is provided for the embodiments of the present application.
[0021] Figure 5 A machine learning-based wire harness assembly robot control system diagram is provided for the embodiments of the present application. DETAILED DESCRIPTION
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0023] It should be noted that all user information (including but not limited to user device information, user personal information, and object information corresponding to device usage data) and data (including but not limited to data for analysis, stored data, displayed data, and device usage data) involved in all embodiments of the present disclosure are information and data authorized by the user or authorized by all parties.
[0024] As shown in Figure 1 The embodiments of the present application disclose a machine learning-based wire harness assembly robot control method 100, which comprises the following method steps: S1, real-time acquisition of multi-dimensional sensor data of a robot when performing a wire harness plugging task, interception of the multi-dimensional sensor data by a sliding time window, and construction of a multi-channel contact event tensor, wherein the multi-dimensional sensor data at least includes joint torque, end effector position error, and end effector attitude change; S2, inputting the contact event tensor into a pre-trained sparse autoencoder model, encoding the contact event tensor into a sparse vector by an encoder of the sparse autoencoder model, reconstructing a reconstructed tensor based on the sparse vector by a decoder of the sparse autoencoder model, and decomposing an input signal into a sparse component corresponding to an effective contact event and a residual component corresponding to interference noise by comparing the contact event tensor with the reconstructed tensor. S3, based on the sparse vector, identify whether a valid contact event with controllable significance has occurred within the current time window, specifically by judging the activation mode of the sparse vector, including judging whether its L1 norm exceeds a first threshold. S4. When a valid contact event is detected, the robot controller invokes a fine-tuning control strategy predefined for that event type to execute an action response. The fine-tuning control strategy includes adjusting the insertion speed, applying a fine-tuning angle, or applying a constant micro-thrust.
[0025] In some embodiments, for step S1, according to the present invention, a 6-7 axis wire harness assembly robot is used as the core execution entity to carry out real-time acquisition of multi-dimensional sensor data during the wire harness insertion process. The acquisition frequency needs to be compatible with the sampling capability of the robot controller. To ensure accurate capture of short-term weak contact signals, the basic sampling rate is set to no less than 200Hz, preferably 500-1000Hz. This sampling frequency range can effectively cover the signal characteristics of typical short-term contact events such as terminal scratching and slot damping, avoiding signal distortion or feature loss due to insufficient sampling frequency. The specific sensor data types and implementation methods are as follows: Joint torque data is collected, for example, by combining joint current inversion with dedicated torque sensors to obtain the torque data of each joint of the robot. For instance, for each motion joint of the robot (e.g., a 6-axis robot has 6 joints), the motor operating current is collected in real time by a current sensor installed at the joint motor. Based on the linear correspondence between the motor torque constant and the current, the torque data is obtained using a formula... (in For joint torque, The torque constant of the motor is... The initial torque value is obtained by inversion from the motor current. Meanwhile, a high-precision torque sensor (with a measurement accuracy of no less than 0.01 N·m) is installed at the end of the joint output shaft to directly collect the actual output torque of the joint. The inverted value and the direct measurement value are then weighted and fused (the weight is dynamically adjusted according to the sensor accuracy, such as a weight of 0.7 for the torque sensor and a weight of 0.3 for the current inversion value). Finally, one scalar torque value is output for each axis to ensure the accuracy and reliability of the torque data. This data directly reflects the change in force on the joint during the wiring harness insertion process and is one of the core data for judging contact events.
[0026] Collecting end effector position error data, which is, for example, combined with a visual positioning system and a robot kinematics inverse solution algorithm to achieve accurate calculation of position error. Specifically, an industrial camera (resolution not less than 1920x1080, frame rate not less than 30 fps) is fixedly installed in the robot working area, and a positioning marker point (such as a circular reflective marker) on the end effector is identified in real time by a machine vision algorithm to obtain the actual coordinates of the marker point in the world coordinate system ; at the same time, based on the robot kinematics model, the theoretical coordinates of the end effector are calculated by inverse solution based on the joint angle values collected by the joint encoders ; the position error vector is obtained by subtracting the actual coordinates from the theoretical coordinates, that is , , , the unit is mm, which represents the deviation of the actual position of the end effector from the theoretical insertion path, and provides data support for adjusting the insertion posture; Collecting end effector posture change data, which is, for example, collected by combining a gyroscope and a posture sensor to output three-dimensional angle error . Specifically, a MEMS gyroscope (range ±2000° / s, zero drift ≤0.01° / h) and a three-axis acceleration sensor are integrated inside the end effector to collect angular velocity and acceleration data of the end effector in real time, and the collected data is fused and processed based on a Kalman filter algorithm to obtain the actual attitude angle of the end effector in the world coordinate system ; combined with the theoretical attitude angle obtained by inverse solution of the robot kinematics , the attitude error , , is calculated, the unit is ° or rad, which reflects the attitude deviation of the end effector during the insertion process and is a key parameter to ensure insertion accuracy.
[0027] According to the embodiment of the application, the collected multi-dimensional sensor data is divided in time dimension by using a sliding time window, and the setting of the window parameter needs to balance between "signal integrity" and "real-time", to ensure that it can completely cover a single short-time contact event, and will not cause too high control delay due to too large window. The specific parameter setting and preprocessing operation are as follows: A sliding time window parameter is set, for example, the window length is set to 40-120 ms, exemplarily, when the sampling rate is 500 Hz, the window length of 80 ms corresponds to the number of data frames in the window of 40 frames, and the length can completely cover the common short-time contact events such as terminal scraping and jack inlet (typical duration 30-80 ms); the window shift step is set to 10-25 ms, and the window overlap rate is 50-80%, for example, the window length is 80 ms, the step is 20 ms, and the overlap rate is 75%, by setting a high overlap rate, the missed detection of contact events caused by too large window interval is avoided, and it is ensured that all short-time contact events can be captured by at least one window; Data preprocessing, pre-processing operation is performed on the multi-channel data in each window, and the data noise and dimension difference are eliminated, which lays a foundation for subsequent tensor construction and model training, specifically including: DC offset removal processing, for example, the mean value of all data points of each channel in the current window is calculated, and each data point of the channel is subtracted by the corresponding mean value, that is, (wherein is the original data point, is the mean value in the window, is the data point after removing the offset), through the operation, the baseline interference caused by sensor zero drift or static bias is removed, the data is concentrated around zero value, and the dynamic change caused by the contact event is more clearly reflected.
[0028] Standardization processing, the data after removing the offset is normalized by using the running sliding mean and standard deviation, specifically, based on the data set of 10-20 historical windows, the sliding mean of each channel is calculated in real time And the sliding standard deviation , through the formula , the data is mapped to the standard normal distribution interval (mean 0, standard deviation 1); optionally, upper and lower limit thresholds calibrated offline can also be used for normalization, and the data is mapped to the interval [0, 1] or [-1, 1], that is, (wherein , is the minimum and maximum value of the channel data obtained by offline calibration), the standardization processing ensures that the data magnitudes of different channels (such as torque, position error and attitude angle) are uniform, and avoids that a channel data magnitude is too large to dominate the model training.
[0029] Preferably, it further includes a differential feature enhancement processing, for example, the first-order difference and the second-order difference of the pre-processed original data are calculated, and they are supplemented to the data set as additional channels to enhance the representation ability of short-time mutation signals. Specifically, the first-order difference is calculated by the difference value of adjacent frame data, that is, (wherein is the first-order difference data of the i-th frame, , is the first-order difference, reflecting the rate of change of the data; the second-order difference is calculated by the adjacent frame difference of the first-order difference, i.e. , reflecting the acceleration of data change; by adding a difference feature channel, sudden terminal scratches, instantaneous slot damping changes and other short-time mutation events can be more sensitively captured, and the recognition accuracy of the model for effective contact events is improved.
[0030] In an embodiment, for contact event tensor construction, specifically, the multi-channel data after preprocessing and intercepted by a sliding time window is converted into a unified tensor structure, i.e., a contact event tensor, which integrates the time dimension time sequence information and the spatial dimension multi-sensor correlation information, and provides a structured input for the sparse autoencoder to extract contact pattern features.
[0031] According to an embodiment of the present application, the dimension of the contact event tensor is defined as , wherein: is the number of channels, which is composed of original sensor channels and optional difference channels. Exemplarily, for a 6-axis robot, the original sensor channels include 6 channels of joint torque, 3 channels of position error, and 3 channels of attitude change, totaling 12 channels; if first-order difference and second-order difference features are added, each original channel corresponds to 1 first-order difference channel and 1 second-order difference channel, and 24 new channels are added, so the total number of channels ; if only first-order difference features are added, the total number of channels is , and the number of channels can be flexibly adjusted according to actual training effect and hardware computing power; is the number of data frames within the window, which is determined by the window length and the sampling rate. For example, when the window length is set to 80 ms and the sampling rate is 500 Hz, frames; when the window length is set to 120 ms and the sampling rate is 1000 Hz, frames. In the tensor construction process, the multi-channel data within each window is arranged in the dimension of "channel x frame" to form a two-dimensional tensor structure, wherein each channel corresponds to a row of data and each frame corresponds to a column of data, ensuring the time sequence continuity and channel independence of the data.
[0032] In some embodiments, for step S2, according to an embodiment of the present application, referring to Figure 2 , a sparse autoencoder model architecture diagram, the sparse autoencoder model adopts an end-to-end network structure of "time sequence encoding-sparse bottleneck-symmetrical decoding", which adapts to the time sequence characteristics and multi-channel features of the contact event tensor, ensures that the model can effectively extract sparse key features of the contact event, and realizes accurate reconstruction of the input tensor. The specific model structure design is as follows: wherein the encoder structure specifically comprises that the encoder adopts a form of combination of a 1D convolution layer and a fully connected layer, which not only preserves the time sequence structure characteristics of data, but also realizes dimension compression and key feature extraction. Specifically, the first layer of the encoder adopts a 1D convolution layer (Conv1D), the convolution kernel size is set to 5, the step is 1, the output channel number is 64, and the activation function is ReLU. The layer extracts short-time local time sequence characteristics of the contact event, such as short pulse changes of the moment signal and local fluctuation patterns of the position error, through local convolution operation of 5 consecutive frames. The convolution layer output feature map dimension is (wherein is the time sequence length of the input tensor); The second layer of the encoder also adopts a 1D convolution layer, the convolution kernel size is set to 3, the step is 1, the output channel number is 32, the activation function is ReLU, the channel dimension is further compressed, and the representation ability of key features is further strengthened. The output feature map dimension is . A flattening operation (Flatten) is connected after the convolution layer, which converts the two-dimensional feature map into a one-dimensional vector, and the vector length is , which is prepared for the subsequent fully connected layer input; The fully connected layer as the output layer of the encoder maps the flattened vector to a sparse bottleneck dimension, the bottleneck dimension is preferably 16-32 dimensions (balance feature expression ability and calculation efficiency), the activation function adopts softplus, and the softplus function can avoid the neuron death problem, which is more suitable for processing weak feature signals. The sparse vector is output through the layer, realizing dimension compression and sparse feature extraction of the input tensor.
[0033] The decoder structure specifically comprises that the decoder adopts a symmetric structure design as the encoder, realizing reconstruction from the sparse vector to the original contact event tensor dimension. Specifically, the first layer of the decoder is a fully connected layer, which maps the sparse vector back to the same length as the flattened output of the second layer of the encoder, the activation function is ReLU, and the non-linear mapping ability of the reconstruction process is ensured; the fully connected layer output vector is converted into a feature map with the same dimension as the output of the second layer of the encoder (N2*64) through reshape operation; The second layer of the decoder is a transpose 1D convolution layer (Conv1DTranspose), the convolution kernel size is set to 3, the step is 1, the input channel number is 32, the output channel number is 64, and the activation function is ReLU. The time sequence dimension is gradually recovered through the transpose convolution operation, and the dimension compression loss in the encoding process is compensated. The layer output feature map dimension is ; The output layer of the decoder is a transpose 1D convolutional layer with a kernel size of 5, a stride of 1, an input channel number of 64, and an output channel number of (the same as the number of channels of the input contact event tensor), without an activation function, and directly outputs the reconstructed tensor , which has the same dimension as the input tensor , ensuring complete reconstruction of the input signal and providing a foundation for subsequent signal decomposition; , ensuring complete reconstruction of the input signal and providing a foundation for subsequent signal decomposition; The sparse bottleneck is a sparse vector output by the fully connected layer of the encoder , which has a dimension of 16-32 and is the core component of the model for realizing signal semantic decoupling. By applying sparse regularization constraints during training, only a small number of elements in the sparse vector are significantly activated (i.e., the proportion of non-zero value elements is extremely low), and these activated elements correspond to key features of effective contact events (such as specific torque patterns of terminal scratching, position change characteristics of slot damping), while the unactivated elements correspond to irrelevant interference information, thereby achieving preliminary separation of effective contact features and interference noise.
[0034] In one embodiment, preferably, the sparse autoencoder model adopts an unsupervised training method, optimizes model parameters using unlabeled multi-scenario assembly data, and does not require manual labeling of effective contact events and interference noise, thereby reducing data labeling costs and improving the adaptability of the model to different scenarios. Referring to Figure 3 , the specific training process is as follows: S301, construct a training data set, collect multi-scenario unlabeled sensor data during the robot wire harness assembly process, covering typical scenarios such as normal insertion, terminal scratching, slot damping, bottom unlocking, wire harness swinging, structure resonance, and static stability, to ensure that the data set can cover all possible contact events and interference types. The collected raw sensor data is preprocessed, windowed, and contact event tensor constructed according to the methods of steps 1.1-1.3, to finally form the training data set. Exemplarily, the number of data set samples is not less than 10,000, ensuring the sufficiency and generalization ability of model training; S302, design a loss function, specifically including minimizing a custom loss function, which includes a reconstruction error term, a sparse regularization term, and a weight decay term, with the specific expression being: , where the reconstruction error term is calculated using the mean square error (MSE) of the difference between the input tensor and the reconstructed tensor , forcing the decoder to learn the basic distribution and timing characteristics of the data and ensuring that the model can accurately reconstruct the input signal to provide a reliable reconstruction basis for subsequent signal decomposition; Sparse regularization term (L1 norm penalty) , which is the core of the model sparsity, by imposing a penalty constraint on the sparse vector , only a small number of elements in the sparse vector are significantly activated. Optionally, the sparse regularization term can take two forms: one is the L1 norm penalty, i.e. , by imposing a penalty on the L1 norm of the sparse vector , directly limiting the number of non-zero elements, , the initial value is set to 1e-3; the second is the KL divergence constraint, i.e. , where is the target activation rate (preferably 0.05, i.e. only 5% of the elements are activated), is the average activation rate of the sparse vector , which forces the actual activation rate to approach the target value through KL divergence, ensuring the effectiveness of the sparsity constraint; Weight decay term (L2 regularization) , L2 regularization is applied to all trainable parameters of the model, is set to 1e-5 to prevent overfitting due to excessive parameters and improve the generalization ability of the model.
[0035] S303, set the training parameters and execute the training process, which includes batch training, batch size is set to 32-256, the specific value can be adjusted according to the data volume and hardware computing power, for example, when the GPU computing power is sufficient and the data volume is large (100,000+ samples), the batch size is set to 256 to improve the training efficiency; in the CPU or low-power GPU environment, the batch size is set to 32 to avoid memory overflow. The optimizer is Adam, the initial learning rate is set to 1e-3, and the adaptive learning rate adjustment strategy is used, when the validation set reconstruction error does not decrease for 5 consecutive rounds, the learning rate is reduced to 0.5 times of the original, to ensure fast convergence and avoid local optimal solution.
[0036] The training rounds are set to 50-200 rounds, and the early stopping strategy is adopted, when the validation set reconstruction error does not decrease for 10 consecutive rounds, the training is stopped, and the current optimal model parameters (including the convolution kernel weights, fully connected layer weights, bias terms, etc. of the encoder and decoder) are saved to avoid overfitting of the model.
[0037] In one embodiment, preferably, signal decomposition is achieved by comparing the input tensor and the reconstructed tensor, specifically, the real-time constructed contact event tensor is input into the pre-trained sparse autoencoder, through the cooperation of the encoder and the decoder, the automatic semantic decoupling of "effective contact event-interference noise" is realized, which breaks through the limitation of traditional "threshold + filtering" that can only distinguish signals based on amplitude, and the specific implementation is as follows: Sparse component extraction, specifically, the encoder receives the real-time contact event tensor After feature extraction and dimension compression by convolutional layers and fully connected layers, the sparse vector is output If the input tensor contains valid contact events (such as terminal scratching, slot damping), the sparse vector The elements corresponding to the valid contact features are significantly activated (output value is larger), while the elements corresponding to the interference noise are in a suppressed state (output value is close to 0); the decoder performs inverse mapping and reconstruction operations based on the sparse vector The sparse reconstruction tensor is output , which only retains feature information related to valid contact events, such as the moment monopole short pulse corresponding to terminal scratching and the position error stable offset pattern corresponding to slot damping, i.e. the sparse component, whose semantics correspond to "valid contact events"; Residual component calculation, specifically, the residual tensor is obtained by subtracting the reconstruction tensor from the input tensor , which contains chaotic signals that the model cannot represent with sparse features, these signals have the characteristics of dense distribution, no obvious directionality and regularity, corresponding to the moment fluctuations of the wire bundle swinging irregularly, the low-amplitude periodic noise of the structure resonance, and the interference information such as the working platform vibration, i.e. the residual component, whose semantics correspond to "interference noise"; For example, when the input tensor contains terminal scratching events, the sparse vector output by the encoder The 3rd and 7th dimension elements are significantly activated (output value is greater than 0.8), and the output values of other dimension elements are close to 0; based on The reconstruction tensor , only the moment change waveform and position error offset features corresponding to terminal scratching are retained; the residual tensor contains the wire bundle swinging noise collected at the same time, which is characterized by irregular small-amplitude moment fluctuations, through this decomposition process, the accurate semantic separation of valid contact events and interference noise is realized, providing reliable input for subsequent controller decision-making.
[0038] In some embodiments, preferably, after the sparse autoencoder model training is completed, a model semantic operation is required to realize the semantic mapping of the sparse vector to "valid / invalid events" and provide a basis for real-time valid contact event recognition. Referring to Figure 4 , the semantic diagram, the specific implementation is as follows: S401, sparse vector clustering, including extracting the sparse vectors corresponding to all samples (not less than 10000) in the training process , using K-means clustering algorithm or Gaussian mixture model (GMM) for unsupervised clustering, the number of clusters Set to 6-12, covering all typical scenarios in the wire harness assembly process (entry scratch, slot damping, bottom unlocking, wire harness swing noise, structure resonance, static stability, etc.); Specifically, the implementation steps of the K-means clustering algorithm are as follows: first, randomly select vectors from the sparse vector sample set as initial cluster centers; calculate the Euclidean distance between each sample vector and each cluster center, and assign the sample to the cluster with the nearest cluster center; calculate the mean of all samples in each cluster and update the cluster center; repeat the above assignment and update steps until the cluster center no longer changes or the maximum number of iterations (set to 100) is reached; To determine the optimal number of clusters , the elbow rule is used to analyze the clustering error, that is, the sum of squares within clusters (SSE) corresponding to different values is calculated; when increases from 1 to 6, the SSE decreases significantly; when exceeds 6, the SSE decreases significantly, forming an "elbow", at which point the optimal number of clusters , corresponding to the six typical scenarios of entry scratch, slot damping, bottom unlocking, wire harness swing noise, structure resonance, and static stability; In one embodiment, S402, semantic labeling, includes randomly extracting 50 to 100 corresponding original contact event tensors and sensor data for each cluster, and performing semantic labeling by manual review. Specifically, the human being reviews the original sensor data waveform (such as torque-time curve, position error-time curve) and contact event tensor features corresponding to each sample, and classifies and labels according to the feature pattern, Specifically, if the sensor data waveform corresponding to the sample has obvious directionality and regularity, which conforms to the feature pattern of the effective contact event, it is labeled as a specific contact event type. For example, if the torque waveform corresponding to the sample presents a single-peak short pulse (duration 30-50 ms, peak value 0.5-1.2 N·m), and the position error presents a transient offset followed by a rapid recovery feature, it is labeled as an "entry scratch event"; if the torque waveform corresponding to the sample presents a stable small rise followed by a constant value (duration 100-200 ms, stable value 0.3-0.8 N·m), and the position error presents a continuous and stable offset, it is labeled as a "slot damping event"; if the torque waveform corresponding to the sample presents a transient drop (drop amplitude 0.4-0.6 N·m, duration 10-20 ms), and the position error does not change significantly, it is labeled as a "bottom unlocking event"; If the sensor data waveform corresponding to the sample has no obvious directionality and regularity, and presents chaotic fluctuation characteristics, it is labeled as interference type. For example, the torque waveform corresponding to the sample presents irregular small fluctuations (amplitude 0.05-0.2 N·m, no fixed period), which is labeled as "wire bundle swing noise"; the torque waveform corresponding to the sample presents low-amplitude periodic fluctuations (period 20-50 ms, amplitude 0.03-0.15 N·m), which is labeled as "structure resonance noise"; all sensor data waveforms corresponding to the sample are close to zero and have no obvious fluctuations, which is labeled as "static stability" (no event); In an embodiment, preferably, if a small amount of artificially labeled "effective event samples" (such as artificially triggering terminal scratching, slot damping events, and recording the time, with a sample size of not less than 500) can be obtained, a light classifier (such as logistic regression, 3-layer small MLP) can be trained based on the sparse vectors of these samples to fine-tune the cluster boundaries after clustering and improve the accuracy of semantic labeling. Specifically, the labeled effective event sample sparse vector is used as a positive sample, and the invalid noise sample sparse vector is used as a negative sample, which is input into the logistic regression classifier for training. By adjusting the weights of the classifier, the boundary between clusters is optimized, and the discrimination accuracy of similar events (such as slight scratching and wire bundle swing noise) is improved by 10-15%, further ensuring the accuracy of subsequent real-time identification.
[0039] In some embodiments, for step S3, the sparse vector obtained by real-time inference is used to identify effective contact events using a "threshold judgment + cluster matching" dual logic to ensure the accuracy and reliability of the identification results, and the specific implementation is as follows: First, the L1 norm of the sparse vector is calculated (||x||1) , where is the dimension of the sparse vector), which is compared with a preset first threshold . The first threshold is obtained by statistical analysis of the L1 norm of the sparse vector of the effective event cluster in the training set, specifically, it is 0.8 times the minimum value of the L1 norm of all sample sparse vectors in the effective event cluster, for example, if the minimum value of the L1 norm of the sample in the effective event cluster is 2.5, then ; if , it is preliminarily determined that there is an effective contact event in the current window; At the same time, the activation value of a specific dimension (corresponding to the dimension of a key contact feature, such as the dimension reflecting the change of axial torque or the dimension of position error offset) in the sparse vector is extracted and compared with a preset second threshold . The second threshold is 0.7 times the average activation value of this dimension in the effective event cluster, for example, if the average activation value of a certain key dimension in the effective event cluster is 0.6, then If the activation value of this dimension exceeds This further verifies the possibility of a valid contact event. Next, Euclidean distance is used to calculate the real-time sparse vector. The distance to the center of each semantic annotation cluster is calculated using the following formula: (in for With the j-th cluster center distance, for The element of the i-th dimension (for the j-th cluster center element of dimension i); Assign to the nearest cluster (i.e.) ); If the cluster is a "valid event cluster," the current event type is determined (e.g., "entrance scratch event," "slot damping event"); if it is an "invalid noise cluster," it is identified as interference noise, and subsequent control responses are terminated. For example, real-time sparse vectors... Its L1 norm is 2.3 (greater than 2.3). Its third dimension activation value is 0.5 (greater than 0.5). The distance to the center of the "entrance scratch event" cluster is 0.8 (the smallest among all clusters), therefore it is determined that an "entrance scratch event" has occurred; if The L1 norm is 1.2 (less than 1). If the noise source is closest to the center of the "wire harness oscillation noise" cluster, it is determined to be interference noise. Thus, this dual-judgment logic not only ensures recognition speed (threshold judgment can quickly filter) but also improves the accuracy of event type recognition through cluster matching, realizing semantic recognition of weak contact signals.
[0040] In some embodiments, for step S4, a mapping table of "event type - fine-tuning strategy" is constructed based on the semantically labeled valid event types. The table defines the specific control parameters and action instructions corresponding to each type of valid event. The control parameters are determined through offline debugging and actual assembly verification to ensure the effectiveness and security of the fine-tuning strategy. The specific construction method is as follows: For example, the strategy mapping table is stored in a structured data format (such as JSON or XML) in a local or cloud database on the robot controller, facilitating real-time access and parameter updates. The core fields of the mapping table include "Event Type ID," "Event Type Name," "Fine-tuning Strategy Type," "Control Parameter Set," "Execution Duration," and "Applicable Scenarios." The "Control Parameter Set" contains all the specific parameters required to implement the fine-tuning strategy, and its content is dynamically adjusted based on the event type. For example, the specific mapping relationship and parameter configuration may include, for the entry scratch event, for example, the event type ID is 001, and the fine-tuning strategy type is "deceleration + angle fine-tuning". The set of control parameters includes "speed adjustment coefficient" (set to 0.5-0.7, that is, the insertion speed is reduced by 30-50%, dynamically adjusted based on the current speed, for example, the current speed is 10mm / s, and the adjusted speed is 5-7mm / s), "micro-rotation angle" (micro-rotation of 0.3-0.5° along the normal direction of the end effector, the specific angle is dynamically adjusted according to the scratching force, the greater the scratching torque, the larger the rotation angle), and "rotation direction" (opposite to the scratching direction, determined by the sparse vector activation dimension, such as the 3rd dimension activation corresponding to clockwise scratching, then the rotation direction is counterclockwise); the execution duration is set to 30-80ms, matching the typical duration of the entry scratch event; the applicable scenario is "in the initial stage of terminal insertion into the socket, the terminal and the socket entry experience slight scratching"; For slot damping events, for example, event type ID is 002, and fine-tuning strategy type is "constant micro-thrust + attitude correction". The set of control parameters includes "micro-thrust magnitude" (applying a constant axial micro-thrust of 0.2-0.5 N·m, achieved through joint torque control; the thrust magnitude is dynamically adjusted according to the damping torque; the larger the damping torque, the greater the thrust) and "attitude correction amplitude" (correcting attitude error). Compensation is performed, with the correction range being 80-90% of the current error to ensure accurate end effector posture; the correction rate is set to 0.1° / ms to avoid damage to the wiring harness caused by excessively rapid posture adjustment; the execution duration is set to 100-200ms to match the typical duration of the slot damping event; the applicable scenario is "the insertion resistance increases due to slot damping during the terminal insertion process"; For bottom unlocking events, for example, event type ID is 003, and the fine-tuning strategy type is "short pulse advance + retraction confirmation". The set of control parameters includes "pulse advance displacement" (set to 0.1-0.2mm to ensure the terminal fully reaches the bottom of the plug and triggers the unlocking mechanism), "pulse duration" (set to 10-20ms to avoid excessive advance causing damage to the wiring harness), "retraction displacement" (set to 0.05-0.1mm, retraction duration 10ms), and "confirmation condition" (monitoring torque changes during retraction; if the torque value recovers to the stable threshold range (0.05-0.1N·m), unlocking is considered successful; otherwise, the advance operation is re-executed, with a maximum of 3 retries); the execution duration is set to 30-50ms; the applicable scenario is "the terminal is close to the bottom of the plug and the unlocking mechanism needs to be triggered to complete the final fixation".
[0041] In one embodiment, preferably, to avoid control instruction mutation causing harness damage or assembly deviation, the fine-tuning strategy is implemented in the manner of "additive superposition + smooth transition", the fine-tuning micro-motion is organically fused with the robot nominal trajectory, and the smoothness and precision of motion execution are ensured. The specific implementation is as follows: The robot main controller generates a theoretical nominal trajectory according to the preset plug-in path (including position and attitude information), which serves as the basic motion instruction; the micro-motion trajectory generated by the fine-tuning strategy (such as position offset corresponding to angle fine-tuning, displacement compensation corresponding to micro-thrust) is superposed to the nominal trajectory in an additive manner to obtain the final control instruction , wherein is a smoothing coefficient (value range 0-1) for realizing gradual mixing of micro-motion; Specifically, the change curve of the smoothing coefficient is designed using an S-shaped function (such as Sigmoid function), the initial value is 0, and it is smoothly raised from 0 to 1 within 5-10 ms, and then remains 1 until the end of the micro-motion, and after the end of the micro-motion, it is smoothly lowered from 1 to 0 within 5-10 ms to restore the nominal trajectory control. Through this design, the micro-motion and the nominal trajectory are seamlessly connected, the harness vibration or plug-in deviation caused by motion mutation is avoided, and the smoothness of the assembly process is ensured; Exemplarily, when an entry scraping event occurs, the nominal trajectory is a uniform linear motion along the positive direction of the X-axis (speed 10 mm / s); the micro-motion trajectory is a rotation motion along the Z-axis (angle 0.5°) and a deceleration motion along the X-axis (speed from 10 mm / s to 5 mm / s); the smoothing coefficient is raised from 0 to 1 within 8 ms, and the final control instruction realizes the robot to rotate slowly along the Z-axis while decelerating along the X-axis, and smoothly completes the scraping avoidance action.
[0042] Preferably, the robot controller converts the final control instruction into motion instructions (such as joint angle, angular velocity) of each joint, controls the joint motor to execute the action through the servo driver; at the same time, the joint torque, end position and attitude data are collected in real time, the deviation between the actual micro-motion execution effect and the theoretical effect is calculated, and if the deviation exceeds the preset threshold (such as position deviation 0.1 mm, attitude deviation 0.1°), the smoothing coefficient and the micro-motion parameters (such as increasing the rotation angle, prolonging the deceleration time) are dynamically adjusted to ensure that the fine-tuning strategy achieves the expected effect.
[0043] In some embodiments, preferably, during the micro-motion execution process, a safety mechanism needs to be established to monitor the joint torque data in real time. If the torque exceeds the limit, an interruption and safety shutdown logic is triggered immediately to ensure the safety of the assembly process and avoid damage to the wiring harness or the robot equipment.
[0044] For example, the robot controller collects torque data of each joint in real time through the joint torque sensor, with a sampling interval of no more than 1 ms, ensuring that torque abnormalities can be detected in time. Based on the material of the wiring harness, the strength of the terminal, and the load capacity of the robot, the torque safety threshold of each joint is preset, which is set to 80% of the maximum allowable torque of the joint (for example, if the maximum allowable torque of the joint is 2 N·m, the safety threshold is set to 1.6 N·m), leaving a certain safety margin to avoid false triggering due to instantaneous torque fluctuations. If the torque value of a joint exceeds the safety threshold, an interruption signal is sent to the micro-motion execution module to stop the current micro-motion execution, and the smoothing coefficient is set to 0 immediately to prevent the micro-motion from continuing to execute and causing the torque to further increase. The safety shutdown logic is started to control the end effector to stop the feeding motion and retreat to the nearest safe position (which is pre-stored in the controller, 5-10 mm away from the current position, ensuring no contact with the workpiece) at a preset safety speed (such as 2 mm / s).
[0045] Figure 5 A wiring harness assembly robot control system 500 based on machine learning is shown. The system embodiment corresponds to the method embodiment shown Figure 1 and specifically includes: The data acquisition and construction module 501 is used to collect multi-dimensional sensor data of the robot in real time when performing the wiring harness insertion task, and the multi-dimensional sensor data is intercepted by a sliding time window to construct a multi-channel contact event tensor, wherein the multi-dimensional sensor data at least includes joint torque, end effector position error, and end effector attitude change. The signal decomposition module 502 is used to input the contact event tensor into a pre-trained sparse autoencoder model. The encoder of the sparse autoencoder model encodes the contact event tensor into a sparse vector, and the decoder reconstructs a reconstructed tensor based on the sparse vector. By comparing the contact event tensor and the reconstructed tensor, the input signal is decomposed into a sparse component corresponding to an effective contact event and a residual component corresponding to interference noise. The event recognition module 503 is used to identify whether an effective contact event with control significance occurs in the current time window based on the sparse vector. Specifically, it is realized by judging the activation pattern of the sparse vector, including judging whether its L1 norm exceeds a first threshold. The control execution module 504 is configured to, when an effective contact event is identified, the robot controller invokes a fine control strategy predefined for the event type to perform an action response, the fine control strategy including adjusting the insertion speed, applying a fine angle, or applying a constant fine thrust.
[0046] The above description is merely a specific implementation of the application. The scope of protection of the application is not limited in this manner. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the application, and all such changes or replacements should be covered within the scope of protection of the application. Therefore, the scope of protection of the application should be subject to the scope of protection of the claims.
Claims
1. A machine learning-based control method for a wire harness assembly robot, characterized in that, Includes the following steps: S1. Real-time acquisition of multi-dimensional sensor data of the robot when performing wire harness insertion task, and the multi-dimensional sensor data is truncated by sliding time window to construct multi-channel contact event tensor, wherein the multi-dimensional sensor data includes at least joint torque, end effector position error and end effector posture change. S2, the contact event tensor is input into a pre-trained sparse autoencoder model. The encoder of the sparse autoencoder model encodes the contact event tensor into a sparse vector. Its decoder reconstructs the reconstructed tensor based on the sparse vector. By comparing the contact event tensor with the reconstructed tensor, the input signal is decomposed into sparse components corresponding to valid contact events and residual components corresponding to interference noise. S3, based on the sparse vector, identify whether a valid contact event with controllable significance has occurred within the current time window, specifically by judging the activation mode of the sparse vector, including judging whether its L1 norm exceeds a first threshold. S4. When a valid contact event is detected, the robot controller invokes a fine-tuning control strategy predefined for that event type to execute an action response. The fine-tuning control strategy includes adjusting the insertion speed, applying a fine-tuning angle, or applying a constant micro-thrust.
2. The machine learning-based control method for a wire harness assembly robot according to claim 1, characterized in that, It also includes, In step S1, the multidimensional sensor data is preprocessed. The preprocessing includes removing the DC offset from the data of each channel and performing normalization processing. And / or, the contact event tensor also includes first-order difference features and / or second-order difference features calculated from the original sensor data.
3. The machine learning-based control method for a wire harness assembly robot according to claim 1, characterized in that, The sparse autoencoder model is trained in an unsupervised manner, and its loss function includes a reconstruction error term and a sparse regularization term, wherein the sparse regularization term is an L1 norm penalty on the sparse vector.
4. The machine learning-based control method for a wire harness assembly robot according to claim 1, characterized in that, include: Before step S3, the model semanticization is also included, which includes, after the sparse autoencoder model is trained, using a clustering algorithm to cluster the sparse vectors obtained from the training dataset to form multiple clusters, semantically labeling these clusters, marking the clusters corresponding to valid contact events as valid event types, and marking the clusters corresponding to various types of interference as noise.
5. The machine learning-based control method for a wire harness assembly robot according to claim 4, characterized in that, Step S4 includes: dividing the sparse vectors obtained in real time into semantically labeled clusters to directly identify specific types of valid contact events, including entry scratch events, slot damping events, or bottom unlocking events.
6. The machine learning-based control method for a wire harness assembly robot according to claim 5, characterized in that, Step S5 includes a preset strategy mapping table, which defines the mapping relationship between different effective contact event types and specific fine-tuning control strategies. When an inlet scraping event is detected, a strategy of deceleration and fine-tuning the angle is executed. When a slot damping event is detected, a strategy of applying constant micro-thrust and correcting the attitude is executed.
7. The machine learning-based control method for a wire harness assembly robot according to claim 1, characterized in that, The signal decomposition specifically includes the following: the sparse component is characterized by the sparse reconstruction tensor obtained by the decoder decoding the sparse vector; and the residual component is obtained by subtracting the sparse reconstruction tensor from the contact event tensor.
8. The machine learning-based control method for a wire harness assembly robot according to claim 1, characterized in that, In step S5, the micro-motions generated by the fine-tuning control strategy are additively superimposed onto the robot's nominal trajectory and progressively mixed using a smoothing coefficient.
9. The machine learning-based control method for a wire harness assembly robot according to claim 1, characterized in that, If a joint torque exceeds a safety threshold is detected during micro-motion execution, the current micro-motion is immediately interrupted and a safety stop is triggered.
10. A machine learning-based control system for a wire harness assembly robot, characterized in that, include: The data acquisition and construction module is used to acquire multi-dimensional sensor data of the robot in real time when performing the wire harness splicing task, and to extract the multi-dimensional sensor data by using a sliding time window to construct a multi-channel contact event tensor. The multi-dimensional sensor data includes at least joint torque, end effector position error and end effector posture change. The signal decomposition module is used to input the contact event tensor into a pre-trained sparse autoencoder model. The encoder of the sparse autoencoder model encodes the contact event tensor into a sparse vector, and its decoder reconstructs the reconstructed tensor based on the sparse vector. By comparing the contact event tensor with the reconstructed tensor, the input signal is decomposed into sparse components corresponding to valid contact events and residual components corresponding to interference noise. The event recognition module is used to identify whether a valid contact event with controllable significance has occurred within the current time window based on the sparse vector. Specifically, it is achieved by judging the activation mode of the sparse vector, including judging whether its L1 norm exceeds a first threshold. The control execution module is used to execute an action response by calling a predefined fine-tuning control strategy for the event type when a valid contact event is detected. The fine-tuning control strategy includes adjusting the insertion speed, applying a fine-tuning angle, or applying a constant micro-thrust.
Citation Information
Patent Citations
Igniter robot production line control method and system based on deep reinforcement learning
CN119579120A
Self-adaptive adjustment system for behavior mode of intelligent robot with body
CN120773064A
Multi-point touch intelligent coding and decoding device, system and method
CN120786067A
Systems and methods for sparse convolution of unstructured data
US20220381914A1
Gated sparse encoder neural networks
WO2025226963A1
Cited By
Robot attitude and emergency synchronous identification method
CN121542927A
A robot posture and emergency synchronous identification method
CN121542927B