Deep learning-based assembly production line digital twinborn cooperative regulation and control method and system

Through deep learning methods, the spatial and temporal reference alignment and feature fusion of multi-source heterogeneous data are achieved, and combined with the prediction model of physical law constraints, the problem of insufficient prediction reliability of digital twin systems in complex assembly scenarios is solved, and the control accuracy of the assembly process and the stress accumulation defect relief effect are improved.

CN120469368AInactive Publication Date: 2025-08-12YANGZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510602571.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In complex assembly scenarios, the fusion problem of multi-source heterogeneous data spatiotemporal reference fusion of existing digital twin systems and the lack of physical constraints in data-driven models lead to insufficient prediction reliability, affecting the control accuracy of the assembly process.

Method used

Using a deep learning-based method, stress distribution data is predicted through space-time reference alignment, multimodal feature fusion and sparse space-time attention network, and a multi-objective optimization algorithm is used to dynamically adjust the robot's motion trajectory and conveying line speed to generate joint control instructions.

Benefits of technology

It significantly improves the control accuracy and stress accumulation defect relief effect of the assembly process, realizes efficient fusion of multi-source data and embeds physical laws, and supports real-time decision-making and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469368A_ABST
    Figure CN120469368A_ABST
Patent Text Reader

Abstract

The invention relates to an assembly production line digital twinborn cooperative regulation and control method and system based on deep learning. The method comprises the following steps: acquiring multi-source heterogeneous data; eliminating a space-time reference difference in the multi-source heterogeneous data through space-time reference alignment, fusing the multi-source heterogeneous data after the space-time reference difference is eliminated by adopting a multi-modal feature fusion technology to obtain a fused feature vector, and transmitting the feature vector to a digital twinborn model; a model prediction layer in the digital twin model receives the feature vector through a sliding window mechanism, and predicts stress distribution data through time sequence convolution and a sparse space-time attention network; and based on the stress distribution data, combining with the real-time state of the equipment, dynamically correcting the motion trail of the robot through a multi-target optimization algorithm, adjusting the running speed of the conveying line, generating a combined control instruction and issuing the combined control instruction to an execution end. According to the method disclosed by the invention, the prediction reliability of the digital twin system in a complex assembly scene and the control precision of the assembly process can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligent manufacturing technology, and in particular to a deep learning-based digital twin collaborative control method and system for an assembly production line. Background Art

[0002] In high-end equipment manufacturing, assembly accuracy and stress accumulation are core metrics for measuring production line performance. For example, in aerospace engine rotor assembly, an axial positioning deviation exceeding ±8μm in the interference fit between the blade and the journal can cause excessive vibration of the entire machine. During repetitive assembly operations by industrial robots, the accumulation of contact stress on titanium alloy fasteners can cause localized lattice slip, resulting in a 23% decrease in fatigue strength after 104 cycles. These precision and stress issues are equally significant in scenarios such as multi-robot collaborative riveting of automotive bodies in white and the precision assembly of satellite reflectors. Their essence stems from the strong coupling of multiple physical fields—geometry, mechanics, and kinematics—during the assembly process.

[0003] Currently, digital twin technology, as a core means of achieving virtual-reality feedback in the assembly process, directly impacts the effectiveness of mitigating the aforementioned issues through its mapping accuracy and real-time decision-making. However, in complex, multi-device collaborative assembly scenarios, existing technology systems still face the following key bottlenecks: 1) The challenge of integrating spatiotemporal data from multiple sources of heterogeneous data. Assembly lines involve multimodal data such as process symbols, robot six-dimensional poses, and component stress tensors. Due to differences in sampling frequencies and non-uniform coordinate systems among different sensors, data fusion suffers from millisecond-level spatiotemporal misalignment. Traditional synchronization methods, such as Kalman filtering, struggle to dynamically calibrate multi-source streaming data, resulting in reduced virtual-reality mapping accuracy for digital twin models. 2) The lack of physical constraints in data-driven models. Existing research often uses purely data-driven models, such as LSTM and Transformer, to predict robot end-point contact forces. While these models can capture temporal characteristics, they ignore rigid body dynamics and contact geometry constraints. For example, in shaft-hole interference fit assembly scenarios, the mean squared error between the predicted force and the true value of traditional models is too high, easily leading to plastic deformation of the assembly part or emergency braking of the robot. The above bottlenecks lead to insufficient predictive reliability of current digital twin systems in complex assembly scenarios, and there is an urgent need to establish a unified spatiotemporal modeling method that integrates physical laws. Summary of the Invention

[0004] To address the problems of insufficient prediction reliability of traditional digital twin systems in complex assembly scenarios, which results in reduced assembly process control accuracy, this paper proposes a deep learning-based collaborative control method for digital twins of assembly production lines to address the above issues.

[0005] According to one aspect of the present disclosure, a method for collaborative control of a digital twin of an assembly line based on deep learning is provided, comprising:

[0006] S10, acquiring multi-source heterogeneous data, wherein the multi-source heterogeneous data is obtained through real-time monitoring of the assembly line and various status detection sensors deployed thereon;

[0007] S20, eliminating the spatiotemporal reference differences in the multi-source heterogeneous data through spatiotemporal reference alignment, fusing the multi-source heterogeneous data after eliminating the spatiotemporal reference differences using a multimodal feature fusion technique to obtain a fused feature vector, and transmitting the feature vector to the digital twin model;

[0008] S30, the model prediction layer in the digital twin model receives the feature vector through a sliding window mechanism, and predicts the stress distribution data through temporal convolution and sparse spatiotemporal attention network;

[0009] S40. Based on stress distribution data and combined with the real-time status of the equipment, the robot motion trajectory is dynamically corrected and the conveyor line speed is adjusted through a multi-objective optimization algorithm. A joint control instruction is generated and sent to the execution end.

[0010] Preferably, eliminating the spatiotemporal reference differences in the multi-source heterogeneous data by spatiotemporal reference alignment includes:

[0011] A global clock reference is established through a time-sensitive network switch, and a cubic spline interpolation algorithm is used to align timestamps of multi-source heterogeneous data with different sampling rates.

[0012] By calibrating the relative pose transformation matrix between the robot base coordinate system and the stress sensor coordinate system, rigid body transformation compensation is applied to the robot pose data to achieve spatial alignment of multi-source device data.

[0013] Preferably, a multimodal feature fusion technology is used to fuse the multi-source heterogeneous data after eliminating the temporal and spatial reference differences to obtain a fused feature vector, including:

[0014] Use the attention mechanism to parse the extensible markup language process tree and output the process feature vector;

[0015] Mapping pose data to feature space through Lie group manifold encoder;

[0016] Extract the core features of the stress field based on the Tucker tensor decomposition model;

[0017] Each modal data is mapped to a unified manifold space through a Lie group manifold encoder, multimodal feature fusion is performed in the manifold space, and a standardized feature vector is output.

[0018] Preferably, multimodal feature fusion is performed in the manifold space, which is expressed as:

[0019] F fused =Log(exp(F symbol)⊙exp(F geo )⊙exp(F tensor )),

[0020] Where ⊙ represents the Lie group multiplication operation, F symbol represents the 64-dimensional process feature vector, F geo is the robot pose feature vector, F tensor The core tensor is flattened to output a 32768-dimensional vector.

[0021] Preferably, predicting stress distribution data by temporal convolution and sparse spatiotemporal attention network includes:

[0022] A five-layer dilated causal convolutional network is used to extract multi-scale temporal features;

[0023] Introducing a physical constraint embedding mechanism into model training, and forcing the prediction results to conform to the laws of material mechanics and kinematic constraints through a regularized loss function;

[0024] A sparse spatiotemporal attention network is constructed based on the assembly CAD model, key nodes are dynamically screened based on the spatial adjacency relationship of components, and stress distribution data is output.

[0025] Preferably, the sparse spatiotemporal attention network is implemented as follows:

[0026] Generate an adjacency matrix based on the spatial topological relationship of the assembly CAD model;

[0027] A multi-head masked attention mechanism is used to limit the attention weight to be transferred only between adjacent nodes;

[0028] Enhance network stability through residual connection and layer normalization techniques.

[0029] Preferably, the robot motion trajectory is dynamically corrected and the conveyor line speed is adjusted through a multi-objective optimization algorithm, and a joint control instruction is generated and sent to the execution end, including:

[0030] Dynamically correct the robot's motion trajectory and conveyor line speed based on stress distribution data;

[0031] A three-level stress threshold monitoring mechanism is established, and when the predicted stress exceeds the allowable value of the material, a graded response instruction is triggered.

[0032] According to one aspect of the present disclosure, a deep learning-based digital twin collaborative control system for an assembly line is provided, comprising:

[0033] A multi-source heterogeneous data acquisition module, which acquires multi-source heterogeneous data, wherein the multi-source heterogeneous data is obtained through real-time monitoring of the assembly line and various status detection sensors deployed thereon;

[0034] A feature vector alignment and fusion module eliminates the spatiotemporal reference differences in the multi-source heterogeneous data through spatiotemporal reference alignment, fuses the multi-source heterogeneous data after eliminating the spatiotemporal reference differences using multimodal feature fusion technology, obtains a fused feature vector, and transmits the feature vector to the digital twin model;

[0035] A stress distribution data prediction module, in which the model prediction layer in the digital twin model receives the feature vector through a sliding window mechanism and predicts the stress distribution data through temporal convolution and a sparse spatiotemporal attention network;

[0036] The optimization scheduling module, based on stress distribution data and combined with the real-time status of the equipment, dynamically corrects the robot's motion trajectory and adjusts the conveyor line speed through a multi-objective optimization algorithm, and generates joint control instructions and sends them to the execution end.

[0037] According to one aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the above-mentioned deep learning-based digital twin collaborative control method for assembly lines.

[0038] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above-mentioned deep learning-based digital twin collaborative control method of the assembly production line is implemented.

[0039] Compared with the prior art, the beneficial effects of the present disclosure are:

[0040] 1) This paper constructs a three-tiered collaborative control architecture for digital twins of assembly lines: spatiotemporal alignment, physical embedding, and closed-loop control. By building a multi-layered collaborative architecture, integrating real-time data synchronization, physical law modeling, and dynamic control modules, it establishes an end-to-end closed-loop control system. Dynamically adjusting execution instructions based on equipment status and predicted data significantly improves assembly process control accuracy and mitigates stress accumulation defects.

[0041] 2) This paper uses a spatiotemporal alignment method to unify multi-source data benchmarks, eliminates spatial deviations through coordinate system compensation, and integrates process knowledge, equipment status, and physical measurement information. This enables efficient compression and feature extraction of heterogeneous data, effectively improving data quality and processing efficiency.

[0042] 3) This paper embeds physical constraints into a time-series prediction model, iteratively updates parameters based on real-time feedback data, and uses adaptive algorithms to balance computational efficiency and prediction accuracy. This establishes a multi-dimensional state joint prediction mechanism, significantly enhancing the accuracy of predictions for key physical quantities and supporting real-time decision-making and control.

[0043] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure.

[0044] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0046] Figure 1 A flowchart of the collaborative control method of digital twins of assembly lines based on deep learning is shown;

[0047] Figure 2 A schematic diagram of the perception fusion layer of the digital twin system is shown;

[0048] Figure 3 A schematic diagram of the model prediction layer of the digital twin system is shown;

[0049] Figure 4 A structural block diagram of the assembly line digital twin collaborative control system based on deep learning in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0050] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0051] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0052] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent the existence of three situations: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.

[0053] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0054] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0055] Example 1

[0056] Based on the above ideas, the present invention proposes a digital twin collaborative control method for assembly production lines based on deep learning. Figure 1 A flowchart of a deep learning-based collaborative control method for digital twins of assembly lines is shown. The method includes:

[0057] S10, acquiring multi-source heterogeneous data, wherein the multi-source heterogeneous data is obtained through real-time monitoring of the assembly line and various status detection sensors deployed thereon;

[0058] S20, eliminating the spatiotemporal reference differences in the multi-source heterogeneous data through spatiotemporal reference alignment, fusing the multi-source heterogeneous data after eliminating the spatiotemporal reference differences using a multimodal feature fusion technique to obtain a fused feature vector, and transmitting the feature vector to the digital twin model;

[0059] S30, the model prediction layer in the digital twin model receives the feature vector through a sliding window mechanism, and predicts the stress distribution data through temporal convolution and sparse spatiotemporal attention network;

[0060] S40. Based on stress distribution data and combined with the real-time status of the equipment, the robot motion trajectory is dynamically corrected and the conveyor line speed is adjusted through a multi-objective optimization algorithm. A joint control instruction is generated and sent to the execution end.

[0061] The present disclosure provides a method for collaborative control of a digital twin of an assembly line based on deep learning, which specifically includes the following steps:

[0062] S10. Acquire multi-source heterogeneous data, wherein the multi-source heterogeneous data is obtained through real-time monitoring of the assembly line and various status detection sensors deployed thereon.

[0063] In this embodiment, the interference fit process between the input shaft and gear of an automobile transmission is used as an example. This process requires that the interference fit between the inner hole of the gear and the journal (0.02-0.05mm) be precisely controlled. Traditional assembly relies on manual experience to adjust the press-fit speed, which can easily lead to micro-deformation of the gear inner ring or failure of the bearing preload due to local stress concentration. In this embodiment, a digital twin is used to monitor the stress distribution of the gear inner ring in real time, dynamically correcting the robot's press-fit trajectory and speed to ensure that the assembly stress is uniform and meets the material allowable value.

[0064] Full-dimensional monitoring of the assembly process is achieved through the coordinated configuration of multiple types of high-precision sensors: High-density stress sensor arrays (such as piezoelectric film sensors or fiber Bragg grating sensors, with a sampling rate of 1-2kHz and a spatial resolution of 0.1-0.5mm) are deployed circumferentially at key assembly interfaces, such as the inner bore of the gear, to capture contact stress distribution in real time. A six-dimensional force / torque sensor (accuracy of ±0.1%-0.5%) and a laser rangefinder (accuracy of ±5-10μm) are integrated at the end of the robot to synchronously collect press force and displacement trajectory. An infrared thermal imager (frame rate of 30Hz) or distributed temperature sensor is also equipped to dynamically compensate for interference deviations caused by thermal expansion. All sensor data is transmitted in real time to edge computing nodes via industrial Ethernet (such as EtherCAT or Profinet) or OPC UA protocols, generating a continuous spatiotemporal dataset at a press speed of 0.05-0.1mm / s, providing high-fidelity, low-latency physical field observation input for subsequent digital twin modeling.

[0065] In this embodiment, a standardized processing system for multi-source heterogeneous data is constructed, such as Figure 2 As shown in the figure, the input data all comes from the physical production line layer, including: the assembly process tree (XML structured description), the robot's six-dimensional pose stream (time series data), and high-frequency stress tensors. First, the spatiotemporal reference alignment module eliminates spatiotemporal reference differences between devices. Subsequently, the multimodal feature fusion module extracts the physical features of process constraints, geometric motion, and stress fields. Finally, multimodal feature fusion is performed in manifold space. The processing flow is as follows: spatiotemporal alignment → process feature extraction → pose encoding → stress field decomposition → Lie group space alignment and fusion. The output: a 128-dimensional normalized feature vector, which is transmitted to the model prediction layer via the OPC UA protocol. This effectively addresses fusion failures caused by data heterogeneity while preserving the physical nature and spatiotemporal correlation characteristics of the assembly process.

[0066] S20. Eliminate the spatiotemporal reference differences in the multi-source heterogeneous data through spatiotemporal reference alignment, use multimodal feature fusion technology to fuse the multi-source heterogeneous data after eliminating the spatiotemporal reference differences, obtain a fused feature vector, and transmit the feature vector to the digital twin model.

[0067] In this embodiment, the spatiotemporal reference differences in the multi-source heterogeneous data are eliminated through spatiotemporal reference alignment, including: establishing a global clock reference through a time-sensitive network switch, and using a cubic spline interpolation algorithm to perform timestamp alignment on the multi-source heterogeneous data with different sampling rates; by calibrating the relative posture transformation matrix between the robot base coordinate system and the stress sensor coordinate system, applying rigid body transformation compensation to the robot posture data, and realizing spatial alignment of multi-source device data.

[0068] The time-space reference alignment eliminates the time-space coordinate differences between multiple source devices through a time synchronization mechanism and a space coordinate system technology, and realizes millisecond-level alignment and fusion of heterogeneous data streams under a unified time-space reference.

[0069] The time synchronization mechanism first deploys the IEEE 802.1AS precise clock protocol through a TSN (Time Sensitive Network) switch to establish a global clock reference. It then uses a cubic spline interpolation algorithm to align timestamps for data at different sampling rates. The interpolation formula is:

[0070] Satisfy x(t k )=x k , x′(t k )=v k ,

[0071] Where t is time, a k is the custom coefficient, t k is the original sampling point, x k is the measured value, v k is the derivative boundary condition.

[0072] The spatial coordinate system first calibrates the relative pose transformation matrix T between the robot base coordinate system {B} and the stress sensor coordinate system {S} B→S ∈SE(3); then rigid body transformation compensation is applied to the robot pose data, expressed as:

[0073] p S =T B→S ·p B ,

[0074]

[0075] Where p B and q B They represent the robot’s three-dimensional position coordinates and its attitude quaternion in the base coordinate system {B}, p S and q S They represent the three-dimensional position coordinates of the robot in the sensor coordinate system {S} and its attitude quaternion after coordinate transformation, q B→SRepresents the relative rotation quaternion from the base coordinate system {B} to the sensor coordinate system {S}, measured by the calibration equipment, Represents quaternion multiplication, which maps all data to the unified sensor coordinate system {S}. Finally, a time-space aligned quintuple sequence is generated, which is expressed as:

[0076] [[t,p S ,q S ,σ,S symbol ]],

[0077] Where σ is the stress tensor, S symbol is a symbolic state identifier.

[0078] The multimodal feature fusion module integrates multimodal data such as process symbols, robot posture and stress tensor through attention mechanism, Lie group encoding and tensor decomposition technology, realizes feature expression that maintains physical constraints in a unified Lie group space, and significantly improves fusion accuracy.

[0079] Furthermore, multimodal feature fusion technology is used to fuse multi-source heterogeneous data after eliminating the temporal and spatial reference differences to obtain a fused feature vector, including: using the attention mechanism to parse the extensible markup language process tree and outputting the process feature vector; mapping the posture data to the feature space through the Lie group manifold encoder; extracting the core features of the stress field based on the Tucker tensor decomposition model; mapping each modal data to a unified manifold space through the Lie group manifold encoder, performing multimodal feature fusion in the manifold space, and outputting a standardized feature vector.

[0080] The specific steps are as follows:

[0081] First, the attention mechanism is used to parse the XML process tree, calculate the weight distribution of assembly constraints and output a 64-dimensional process feature vector F symbol Encode assembly priorities and physical constraints, expressed as:

[0082]

[0083] Where, α i is the attention weight of the i-th process symbol node, indicating the importance of the node in the current assembly context, Q, K i and V i are query, key, and value vectors respectively, and d is the key vector K i dimension, softmax(·) is a normalized exponential function that converts the input vector into a probability distribution, N node The number of nodes for inputting process symbol data.

[0084] Then, the pose data is mapped to a 64-dimensional feature space through the Lie group SE(3) manifold encoder, which is expressed as:

[0085] F geo =Log SE(3) (T(q S ))·W geo ,

[0086] In the formula, T(q s ) is the rigid body transformation matrix corresponding to the posture, Log SE(3) is the Lie group logarithmic mapping, W geo is the trainable parameter matrix.

[0087] Secondly, the core features of the axial stress field of the gear inner hole are extracted based on the Tucker tensor decomposition model, which is expressed as:

[0088]

[0089] Where, is the core tensor, ×1, ×2 and ×3 represent the product operation of the tensor and the matrix in the 1st, 2nd and 3rd dimensions respectively, U (1) 、U (2) and U (3) are the basis vector sets of the 1st, 2nd and 3rd modes respectively. After flattening, the output is a 32768-dimensional vector F tensor .

[0090] Multimodal feature fusion is performed in the manifold space, which is expressed as:

[0091] F fused =Log(exp(F symbol )⊙exp(F geo )⊙exp(F tensor )),

[0092] In the formula, ⊙ represents the Lie group multiplication operation. After fusion, the dimension is reduced to 128 through the fully connected layer. symbol represents the 64-dimensional process feature vector, F geo is the robot pose feature vector. This feature vector satisfies the memory and computing power constraints of the edge computing node and is transmitted to the model prediction layer via the OPCUA protocol, providing standardized input for subsequent physical constraint modeling.

[0093] S30. The model prediction layer in the digital twin model receives the feature vector through a sliding window mechanism, and predicts the stress distribution data through temporal convolution and sparse spatiotemporal attention network.

[0094] See also Figure 3In this embodiment, the model prediction layer receives the 128-dimensional fusion feature vector sequence output by the perception layer through a sliding window mechanism, first uses the edge temporal convolutional network to extract multi-scale temporal features, and outputs the feature H conv The algorithm incorporates multi-scale dynamic features from 7 to 187 frames. Furthermore, a composite loss function L is designed by embedding material mechanics and physical constraints into model training. This regularized loss function ensures that predictions conform to physical laws. Finally, a sparse spatiotemporal attention prediction network is constructed based on the CAD model of the gear-shaft assembly. This network dynamically selects key nodes based on the spatial adjacency of components and outputs a prediction of the spatiotemporal distribution of the gear-shaft stress field for the next 5 seconds. This process, by integrating data-driven and physical modeling, reduces stress prediction errors and minimizes prediction latency while ensuring mechanical rationality.

[0095] Stress distribution data is predicted through temporal convolution and sparse spatiotemporal attention networks, including: using a five-layer dilated causal convolutional network to realize multi-scale temporal feature extraction; introducing a physical constraint embedding mechanism in model training, and forcing the prediction results to conform to the laws of material mechanics and kinematic constraints through a regularized loss function; constructing a sparse spatiotemporal attention network based on the assembly CAD model, dynamically screening key nodes based on the spatial adjacency relationship of components, and outputting stress distribution data.

[0096] The edge temporal convolutional network uses a five-layer dilated causal convolutional network to realize multi-scale temporal feature extraction. The specific structure is as follows:

[0097] (1) Basic causal convolution operator

[0098] The single-layer causal convolution operation is defined as:

[0099]

[0100] Where l represents the lth layer in the network, ReLU(·) represents the linear rectification function, which is a commonly used activation function in deep learning, K is the kernel size, and W is the kernel size. (l) is the convolution kernel parameter of the lth layer, X (l-1) is the input feature of the l-1 layer, d (l) is the expansion factor of the lth layer, and k represents the size of the convolution kernel (i.e., the number of time steps covered by the convolution kernel).

[0101] (2) Hierarchical parameter configuration

[0102] The structural parameters of each layer are shown in the following table.

[0103]

[0104]

[0105] (3) Multi-scale feature extraction

[0106] The formula for calculating the total receptive field of the network is:

[0107]

[0108] Substituting the parameters into the total receptive field coverage is:

[0109]

[0110] The final output feature H conv Contains multi-scale dynamic features from 7 to 187 frames, satisfying:

[0111]

[0112] Where, represents the channel splicing operation, Δt (l) is the delay offset corresponding to each layer.

[0113] The physical constraint embedding mechanism ensures that the prediction results conform to physical laws by embedding material mechanics laws and kinematic constraints in model training. The specific implementation includes the following core modules.

[0114] (1) Physical regularization loss function design

[0115] The composite loss function is used and expressed as:

[0116] L=α·L MSE +β·R Hooke +γ·R Geometry ,

[0117] Where α, β, and γ are regularization coefficients, and L MSE is the basic prediction error term, R Hooke is the regularization term of Hooke’s law, R Geometry is a geometric constraint.

[0118] The basic forecast error term is expressed as:

[0119]

[0120] Where N is the number of samples, To predict the stress tensor, is the actual measured value.

[0121] The regularization term of Hooke's law is expressed as:

[0122]

[0123] Where, It is a material database query function, based on the current temperature T and strain rate Returns the elastic modulus, The data obtained by strain gauge measurement is preprocessed.

[0124] The geometric constraints are expressed as:

[0125]

[0126] Where, is the pseudo-inverse matrix, J is the Jacobian matrix of the robot, Δx (i) is the change in the end effector posture, is the change in robot joint angle predicted by the model.

[0127] (2) Constraint implementation process

[0128] a. Material parameter embedding:

[0129] The Poisson's ratio μ = 0.3 and the elastic modulus E = 210 GPa were obtained by performing a real-time table lookup in the material database.

[0130] b. Adaptive weight adjustment:

[0131] Regularization coefficients β and γ are dynamically adjusted according to the model convergence:

[0132]

[0133] Where, β t is the regularization coefficient at the tth time step, L val is the validation set loss, when the validation set L val Losses exceed historical best 5%, gradually increase the strength of physical constraints.

[0134] c. Gradient calculation strategy:

[0135] The gradient of the strain prediction value is achieved through automatic differentiation and is expressed as:

[0136]

[0137] Where,∈ pred To predict the strain value, ∈ true is the true strain value, X is the true strain value, and the second-order differential calculation is used to ensure accuracy.

[0138] The implementation steps of the sparse spatiotemporal attention network are as follows: generate an adjacency matrix based on the spatial topological relationship of the assembly CAD model; adopt a multi-head masked attention mechanism to limit the attention weight to be transferred only between adjacent nodes; and enhance network stability through residual connections and layer normalization techniques.

[0139] The proposed sparse spatiotemporal attention prediction network builds a sparse adjacency matrix based on the mesh topology of the gear inner hole in the gear-shaft CAD model, restricting attention weights to be transferred only between adjacent mesh nodes. The core implementation steps are as follows.

[0140] (1) Adjacency matrix generation

[0141] According to the spatial topological relationship of the gear-shaft assembly CAD model, the adjacency matrix generation rule is defined as follows:

[0142]

[0143] (2) Sparse Attention Computation

[0144] The multi-head masked attention mechanism is adopted, and the specific implementation process is as follows.

[0145] a. Linear projection generates query, key, and value matrices

[0146] Q=HW Q ,

[0147] K=HW K ,

[0148] V=HW V ,

[0149] Where H is the input feature matrix, W Q 、W K 、W V are learnable parameters.

[0150] b. Sparse attention weight calculation

[0151]

[0152] Where, d k is the key vector K k Dimension, A ij are the elements of the adjacency matrix A, which serves as a mask matrix to force the attention weights of non-adjacent nodes to zero.

[0153] c. Feature aggregation

[0154] Z=Concat(head1,...,head h )W O ,

[0155] In the formula, head i =AttnV,W O represents the output projection matrix, and Concat(·) represents the concatenation operation of the feature vectors.

[0156] (3) Network architecture configuration

[0157] The network consists of the following core components.

[0158] a. Spatiotemporal Position Coding

[0159] A learnable parameter matrix is used to encode the spatiotemporal position information, which can be expressed as:

[0160] E pos =[e1,...,e T ] T ,

[0161] Where T is the number of time steps, e t Represents the spatiotemporal position encoding vector at the t-th time step.

[0162] The input features are injected into the position information through the residual connection, which is expressed as:

[0163] H=F+E pos ,

[0164] Where F is the temporal convolution output.

[0165] b. Encoder stacking

[0166] It consists of 4 cascaded sparse attention layers, each layer contains: 8-head attention mechanism; feedforward network:

[0167] FFN(x)=ReLU(xW1+b1)W2+b2,

[0168] Where b1 and b2 are the learnable bias terms in the feedforward neural network (FFN), and W1 and W2 are learnable parameters.

[0169] Layer Normalization:

[0170] LayerNorm(x+Dropout(Sublayer(x))).

[0171] c. Stress field decoding

[0172] Mapped to the stress tensor through a fully connected network, expressed as:

[0173] σ pred =W5·GELU(W4·LayerNorm(W3Z)),

[0174] Where W3 and W4 are learnable parameters that ultimately output a 3×3 stress tensor, and W5 is the final linear mapping weight matrix in the stress field decoding stage.

[0175] The final prediction output is the spatiotemporal prediction field of the circumferential stress distribution of the gear inner hole within the next 5 seconds.

[0176] S40. Based on stress distribution data and combined with the real-time status of the equipment, the robot motion trajectory is dynamically corrected and the conveyor line speed is adjusted through a multi-objective optimization algorithm. A joint control instruction is generated and sent to the execution end.

[0177] In this embodiment, a multi-objective optimization algorithm is used to dynamically correct the robot's motion trajectory, adjust the conveyor line's operating speed, and generate joint control instructions and send them to the execution end, including: dynamically correcting the robot's motion trajectory and conveyor line's operating speed based on stress distribution data; establishing a three-level stress threshold monitoring mechanism, and triggering a graded response instruction when the predicted stress exceeds the material's allowable value.

[0178] The optimization control layer realizes real-time control of the assembly process through a multi-objective dynamic optimization algorithm based on the stress spatiotemporal distribution prediction results of the model prediction layer: when the predicted stress value is close to the allowable stress of the bearing material, the SE (3) kinematic inverse solution algorithm is called to generate a robot candidate trajectory cluster that meets the stress constraint, and the trajectory sensitivity is calculated in combination with the stress gradient field, and the optimal trajectory with a maximum stress reduction of more than 15% and a trajectory deviation of less than 0.1mm is screened; the control instructions are sent through the real-time Ethernet to synchronously adjust the robot end feed speed, compensate the journal deflection angle and the conveyor line beat, and implement a graded response based on a three-level threshold mechanism - when the predicted stress reaches 800MPa, the process parameter self-check is triggered, when it reaches 850MPa, the reverse micro-retreat is started (withdraw 0.01mm to release local stress), and when it reaches 900MPa, the emergency shutdown is carried out and the acoustic emission detection system is activated, so as to realize full closed-loop control with small end-to-end delay and high stress control accuracy from prediction to control, and finally significantly reduce the stress exceeding rate of the automobile gearbox input shaft and gear assembly.

[0179] Through its application in the interference fit of the input shaft and gear of an automobile transmission, the digital twin collaborative control method of the assembly production line for assembly quality optimization in this embodiment effectively realizes the precise integration of multi-source heterogeneous data, the deep embedding of physical laws, and the real-time optimization and control of the assembly process, significantly improving the assembly quality and efficiency, and verifying the effectiveness and practicality of the method.

[0180] Example 2

[0181] As another aspect of the embodiment of the present disclosure, a deep learning-based assembly line digital twin collaborative control system 100 is also provided. Figure 4 As shown, including:

[0182] Multi-source heterogeneous data acquisition module 1, which acquires multi-source heterogeneous data, wherein the multi-source heterogeneous data is obtained through real-time monitoring of the assembly line and various status detection sensors deployed thereon;

[0183] Feature vector alignment and fusion module 2, which eliminates the spatiotemporal reference differences in the multi-source heterogeneous data through spatiotemporal reference alignment, fuses the multi-source heterogeneous data after eliminating the spatiotemporal reference differences using multimodal feature fusion technology, obtains a fused feature vector, and transmits the feature vector to the digital twin model;

[0184] Stress distribution data prediction module 3, the model prediction layer in the digital twin model receives the feature vector through a sliding window mechanism, and predicts the stress distribution data through temporal convolution and sparse spatiotemporal attention network;

[0185] The optimization scheduling module 4, based on the stress distribution data and combined with the real-time status of the equipment, dynamically corrects the robot's motion trajectory and adjusts the conveyor line running speed through a multi-objective optimization algorithm, and generates joint control instructions and sends them to the execution end.

[0186] In the absence of any contradiction, the above modules in the system of the embodiment of the present disclosure can implement any implementation of the above method.

[0187] Based on the description of the above embodiments, it can be seen that the embodiments of the present disclosure can achieve the following technical effects:

[0188] 1) This paper constructs a three-tiered collaborative control architecture for digital twins of assembly lines: spatiotemporal alignment, physical embedding, and closed-loop control. By building a multi-layered collaborative architecture, integrating real-time data synchronization, physical law modeling, and dynamic control modules, it establishes an end-to-end closed-loop control system. Dynamically adjusting execution instructions based on equipment status and predicted data significantly improves assembly process control accuracy and mitigates stress accumulation defects.

[0189] 2) This paper uses a spatiotemporal alignment method to unify multi-source data benchmarks, eliminates spatial deviations through coordinate system compensation, and integrates process knowledge, equipment status, and physical measurement information. This enables efficient compression and feature extraction of heterogeneous data, effectively improving data quality and processing efficiency.

[0190] 3) This paper embeds physical constraints into a time-series prediction model, iteratively updates parameters based on real-time feedback data, and uses adaptive algorithms to balance computational efficiency and prediction accuracy. This establishes a multi-dimensional state joint prediction mechanism, significantly enhancing the accuracy of predictions for key physical quantities and supporting real-time decision-making and control.

[0191] The present disclosure also provides an electronic device comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the aforementioned deep learning-based collaborative control method for a digital twin of an assembly line. The electronic device may be provided as a terminal, server, or other device.

[0192] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon. When executed by a processor, the computer program instructions implement the aforementioned deep learning-based collaborative control method for a digital twin of an assembly line. The computer-readable storage medium may be a non-volatile computer-readable storage medium.

[0193] Those skilled in the art will understand that in the above-mentioned deep learning-based assembly production line digital twin collaborative control method and system in the specific implementation method, the writing order of each step does not mean a strict execution order and constitutes any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0194] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0195] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technical improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A collaborative control method for digital twins of assembly lines based on deep learning, characterized by: The steps include: S10, acquiring multi-source heterogeneous data, wherein the multi-source heterogeneous data is obtained through real-time monitoring of the assembly line and various status detection sensors deployed thereon; S20, eliminating the spatiotemporal reference differences in the multi-source heterogeneous data through spatiotemporal reference alignment, fusing the multi-source heterogeneous data after eliminating the spatiotemporal reference differences using a multimodal feature fusion technique to obtain a fused feature vector, and transmitting the feature vector to the digital twin model; S30, the model prediction layer in the digital twin model receives the feature vector through a sliding window mechanism, and predicts the stress distribution data through temporal convolution and sparse spatiotemporal attention network; S40. Based on stress distribution data and combined with the real-time status of the equipment, the robot motion trajectory is dynamically corrected and the conveyor line speed is adjusted through a multi-objective optimization algorithm. A joint control instruction is generated and sent to the execution end.

2. The method according to claim 1, characterized in that Eliminating the spatiotemporal reference differences in the multi-source heterogeneous data through spatiotemporal reference alignment includes: A global clock reference is established through a time-sensitive network switch, and a cubic spline interpolation algorithm is used to align timestamps of multi-source heterogeneous data with different sampling rates. By calibrating the relative pose transformation matrix between the robot base coordinate system and the stress sensor coordinate system, rigid body transformation compensation is applied to the robot pose data to achieve spatial alignment of multi-source device data.

3. The method according to claim 1, characterized in that Multimodal feature fusion technology is used to fuse multi-source heterogeneous data after eliminating the temporal and spatial reference differences, and the fused feature vector is obtained, including: Use the attention mechanism to parse the extensible markup language process tree and output the process feature vector; Mapping pose data to feature space through Lie group manifold encoder; Extract the core features of the stress field based on the Tucker tensor decomposition model; Each modal data is mapped to a unified manifold space through a Lie group manifold encoder, multimodal feature fusion is performed in the manifold space, and a standardized feature vector is output.

4. The method according to claim 3, characterized in that Multimodal feature fusion is performed in the manifold space, which is expressed as: F fused =Log(exp(F symbol )⊙exp(F geo )⊙exp(F tensor )), Where ⊙ represents the Lie group multiplication operation, F symbol represents the 64-dimensional process feature vector, F geo is the robot pose feature vector, F tensor The core tensor is flattened to output a 32768-dimensional vector.

5. The method according to claim 1, wherein Predict stress distribution data through temporal convolution and sparse spatiotemporal attention network, including: A five-layer dilated causal convolutional network is used to extract multi-scale temporal features; Introducing a physical constraint embedding mechanism into model training, and forcing the prediction results to conform to the laws of material mechanics and kinematic constraints through a regularized loss function; A sparse spatiotemporal attention network is constructed based on the assembly CAD model, key nodes are dynamically screened based on the spatial adjacency relationship of components, and stress distribution data is output.

6. The method according to any one of claims 1 or 5, characterized in that The steps to implement the sparse spatiotemporal attention network are as follows: Generate an adjacency matrix based on the spatial topological relationship of the assembly CAD model; A multi-head masked attention mechanism is used to limit the attention weight to be transferred only between adjacent nodes; Enhance network stability through residual connection and layer normalization techniques.

7. The method according to claim 1, characterized in that The robot's motion trajectory is dynamically corrected through a multi-objective optimization algorithm, the conveyor line speed is adjusted, and joint control instructions are generated and sent to the execution end, including: Dynamically correct the robot's motion trajectory and conveyor line speed based on stress distribution data; A three-level stress threshold monitoring mechanism is established, and when the predicted stress exceeds the allowable value of the material, a graded response instruction is triggered.

8. The assembly line digital twin collaborative control system based on deep learning is characterized by: include: A multi-source heterogeneous data acquisition module, which acquires multi-source heterogeneous data, wherein the multi-source heterogeneous data is obtained through real-time monitoring of the assembly line and various status detection sensors deployed thereon; A feature vector alignment and fusion module eliminates the spatiotemporal reference differences in the multi-source heterogeneous data through spatiotemporal reference alignment, fuses the multi-source heterogeneous data after eliminating the spatiotemporal reference differences using multimodal feature fusion technology, obtains a fused feature vector, and transmits the feature vector to the digital twin model; A stress distribution data prediction module, in which the model prediction layer in the digital twin model receives the feature vector through a sliding window mechanism and predicts the stress distribution data through temporal convolution and a sparse spatiotemporal attention network; The optimization scheduling module, based on stress distribution data and combined with the real-time status of the equipment, dynamically corrects the robot's motion trajectory and adjusts the conveyor line speed through a multi-objective optimization algorithm, and generates joint control instructions and sends them to the execution end.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the deep learning-based digital twin collaborative control method for assembly production lines described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the deep learning-based digital twin collaborative control method for assembly production lines described in any one of claims 1 to 7.

Citation Information

Cited By

  • Pantograph current collection quality closed-loop optimization method and system

    CN121069740A