An adaptive assembly device based on multi-modal perception and deep reinforcement learning

CN121339908BActive Publication Date: 2026-09-25KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511452631.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-12
Publication Date
2026-09-25
Estimated Expiration
2045-10-12

AI Technical Summary

Technical Problem

传统装配模式存在显著局限性:人工操作受限于体力与技能差异,装配精度低、一致性差;固定程序自动化设备虽提升效率,但依赖预设参数,无法应对零件公差、环境波动等动态变化,换型时需重新编程调试,严重制约生产效率

Benefits of technology

[0013]本发明采用了部分标准件,有效降低了整机的制作成本,同时标准化零件的采购与替换更为便捷,进一步减少了生产及后期维护的经济投入。采用模块化设计,整机包含多个独立模块,各模块可单独安装后再汇总拼装。这种设计不仅简化了加工生产流程,有利于批量制造,还极大地方便了后期的维修与部件更换,降低了维护难度和时间成本。通过传感与控制模块的协同工作,遵循“感知-决策-执行-反馈”的闭环控制逻辑,实现了装配过程数据的实时采集与策略的动态优化,使操作人员能够实时掌握装配状态,并及时预警与处理异常。装配执行模块设计精巧,在驱动滑台的带动下,电动扳手能够精准执行装配操作,配合压紧弹簧缓冲装配冲击,在保证结构稳定性的同时,显著提升了装配效率,可快速完成对多种规格零部件的装配作业。物料输送模块能够适应不同待装配配件及连接件的推送需求,结合装配执行模块的灵活动作和传感与控制模块的自适应调节,使装置能够适配多种不同形状和规格的零部件装配任务,增强了设备的适用性,满足了多样化的装配需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121339908B_ABST
    Figure CN121339908B_ABST
Patent Text Reader

Abstract

The application discloses a kind of adaptive assembly device based on multi-modal perception and depth reinforcement learning, belong to intelligent manufacturing field.The device is through the collaborative work of body, material conveying module, assembly execution module and sensing and control module, realize the automatic assembly of multiple specifications parts, and utilize depth reinforcement learning algorithm combined with multi-modal perception data continuously optimize assembly strategy, to adapt to dynamic change assembly demand.Device includes body, material conveying module, assembly execution module, sensing and control module.The application gives device autonomous perception, dynamic decision and self-optimization capability through multi-modal perception and depth reinforcement learning, so that it can adapt to different parts assembly task, significantly improve assembly precision, efficiency and flexibility, with wide industrial application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an adaptive assembly device that integrates multimodal perception technology and deep reinforcement learning algorithms, belonging to the field of intelligent manufacturing and automated assembly technology. Background Technology

[0002] In modern manufacturing, the quality and efficiency of the assembly process directly determine product competitiveness. Traditional assembly methods have significant limitations: manual operation is limited by differences in physical strength and skills, resulting in low assembly accuracy and poor consistency; while automated equipment with fixed programs improves efficiency, it relies on preset parameters and cannot cope with dynamic changes such as part tolerances and environmental fluctuations. When changing models, reprogramming and debugging are required, which seriously restricts production efficiency.

[0003] Therefore, developing an adaptive assembly device based on multimodal perception and deep reinforcement learning, constructing a complete assembly "cognitive picture" through multi-dimensional state perception, and achieving autonomous evolution of assembly strategies by combining deep reinforcement learning to adapt to various parts of different shapes and specifications has become a key technological direction for solving the pain points of traditional assembly and promoting the intelligent upgrading of the manufacturing industry. Summary of the Invention

[0004] This invention provides an adaptive assembly device based on multimodal perception and deep reinforcement learning. The device is ingeniously designed, simple in structure and easy to operate, and solves the systematic defects of traditional assembly in terms of dynamic adaptability, perception integrity and strategy intelligence.

[0005] The technical solution of this invention is:

[0006] An adaptive assembly device based on multimodal perception and deep reinforcement learning includes a body 1, a material conveying module 2, an assembly execution module 3, and a sensing and control module 4. The material conveying module 2 is installed on the guide rails I1-8 of the body 1 to achieve rapid material loading. The multimodal sensor group in the sensing and control module 4 collects visual images, vibration signals, torque data, and displacement information of the assembly process in real time. The controller 4-5 in the sensing and control module 4 preprocesses and fuses the collected multimodal data, and inputs it into the deep reinforcement learning model for state recognition and decision analysis. Based on the output of the deep reinforcement learning model, an assembly strategy is dynamically generated to control the electric wrench 3-2 and the drive slide 3-11 in the assembly execution module 3 to perform assembly actions. Based on the real-time feedback data of the multimodal sensor group, a closed-loop adaptive control is formed to complete the assembly task.

[0007] Specifically, the machine body 1 includes casters 1-1, a machine body 1-2, crossbeams I1-3 and II1-4, longitudinal beams I1-5, II1-6, and III1-7, a guide rail I1-8, a vertical beam 1-9, a support base 1-10, and a top beam 1-11. Four self-locking casters are installed at the bottom of the machine body 1-2 to enable overall movement and fixation of the device. Crossbeams I1-3, II1-4, I1-5, II1-6, and III1-7 are sequentially installed aligned on the end faces of the machine body 1-2 to form a rigid support structure. The guide rail I1-8 is installed on the upper surface between longitudinal beams I1-5 and II1-6, providing support for the material conveying module 2. For sliding guidance, vertical beam 1-9 is installed on the lower end face of longitudinal beam II1-6 and longitudinal beam III1-7. Support seat 1-10 is fixed on the upper end face of the bottom of the body shape 1-2 to fix assembly execution module 3 and ensure structural stability during assembly. Vision sensor 4-1 is installed on top beam 1-11 to provide visual monitoring angle for the assembly station. Universal wheel 1-1 is connected to body shape 1-2 by screws. The contact surfaces between crossbeam I1-3, crossbeam II1-4, longitudinal beam I1-5, longitudinal beam II1-6, longitudinal beam III1-7, guide rail I1-8, body shape 1-2, vertical beam 1-9, support seat 1-10, and top beam 1-11 are connected by welding.

[0008] Specifically, the material conveying module 2 includes a guide rail II 2-1, a hopper 2-2, a pushing component 2-3, an elastic drive component 2-4, a part to be assembled 2-5, a connector 2-6, and a reset component 2-7. The guide rail II 2-1 is installed at the bottom of the hopper 2-2 to provide low-friction sliding guidance for the pushing component 2-3. The hopper 2-2 stores the part to be assembled 2-5 and the connector 2-6. The pushing component 2-3 pushes the parts to the assembly station under the action of the elastic drive component 2-4. The reset component 2-7 drives the pushing component 2-3 to reset when the hopper 2-2 is replenished, thus achieving rapid material feeding.

[0009] Specifically, the assembly execution module 3 includes a drive slide 3-1, an electric wrench 3-2, and a clamping spring 3-3. The drive slide 3-1 is mounted on the vertical beam 1-9 of the machine body 1 by bolts, and its bottom is mounted on the support base 1-10, forming a double fixed structure. This drives the electric wrench 3-2 to achieve lifting and lowering movements. The electric wrench 3-2 is fixed on the movable side of the drive slide 3-1 and serves as the core execution component to output torque. Its output parameters are controlled in real time by the controller 4-5. The clamping spring 3-3 is installed at the bottom of the electric wrench 3-2 to buffer assembly impacts and prevent damage to parts.

[0010] Specifically, the sensing and control module 4 includes a vision sensor 4-1, a vibration sensor 4-2, a torque sensor 4-3, a displacement sensor 4-4, and a controller 4-5. The vision sensor 4-1 is installed on the top beam 1-11 of the machine body 1, with its lens facing the assembly station, to collect images of the parts to be assembled 2-5 and the connectors 2-6 for positioning and attitude recognition. The vibration sensor 4-2 is installed on the electric wrench 3-2 and the machine body 1 to monitor the vibration spectrum during the assembly process and identify abnormalities such as jamming and abnormal noises. The torque sensor 4-3 is installed between the output shaft and the actuator of the electric wrench 3-2 to collect assembly torque data in real time. The displacement sensor 4-4 is installed at the connection between the movable and fixed sides of the drive slide 3-1 to monitor the lifting and lowering displacement of the assembly tool. The controller 4-5 is placed in a control cabinet near the operating area and deploys a deep reinforcement learning model with a DQN+LSTM fusion architecture. It receives data from multiple sensors through a data acquisition card and outputs drive slide displacement commands and electric wrench torque commands to achieve closed-loop adaptive control.

[0011] Specifically, the assembly workflow is as follows: The pushing component 2-3 of the material conveying module 2, under the action of the elastic drive component 2-4, pushes the parts to be assembled 2-5 and the connecting parts 2-6 to the assembly station; subsequently, the vision sensor 4-1 in the sensing and control module 4 acquires images of the assembly station, performs positioning and attitude recognition, and sends signals to the controller 4-5; the controller 4-5 then drives the assembly execution module 3 to start working, driving the slide 3-1 to lower the electric wrench 3-2 until its bottom clamping spring 3-3 contacts and clamps the parts to be assembled 2-5; the electric wrench 3-2 then... The device rotates to perform the assembly action. During this process, the torque sensor 4-3 collects the tightening torque of the connector 2-6 in real time, the vibration sensor 4-2 monitors the vibration spectrum of the assembly process, and the displacement sensor 4-4 provides feedback on the feed depth of the electric wrench 3-2. The deep reinforcement learning model in the controller 4-5 performs real-time fusion and analysis of the above multimodal data, dynamically optimizes the displacement command output to the drive slide 3-1 and the torque command of the electric wrench 3-2, forming a closed-loop adaptive control loop to ensure that the connector 2-6 is screwed in along the optimal path, and autonomously adjusts the strategy when an abnormality is detected.

[0012] The beneficial effects of this invention are:

[0013] This invention utilizes some standard components, effectively reducing the overall manufacturing cost. Furthermore, the procurement and replacement of standardized parts are more convenient, further reducing economic investment in production and subsequent maintenance. Adopting a modular design, the entire machine comprises multiple independent modules, each of which can be installed individually before being assembled. This design not only simplifies the manufacturing process and facilitates mass production but also greatly facilitates subsequent maintenance and component replacement, reducing maintenance difficulty and time costs. Through the collaborative work of the sensing and control modules, following a closed-loop control logic of "perception-decision-execution-feedback," real-time data acquisition and dynamic strategy optimization during the assembly process are achieved, enabling operators to monitor the assembly status in real time and promptly issue warnings and handle anomalies. The assembly execution module is ingeniously designed; driven by the slide table, the electric wrench precisely performs assembly operations. Combined with a clamping spring to buffer assembly impact, it significantly improves assembly efficiency while ensuring structural stability, enabling rapid assembly of various specifications of parts. The material conveying module can adapt to the pushing requirements of different parts and connectors to be assembled. Combined with the flexible action of the assembly execution module and the adaptive adjustment of the sensing and control module, the device can adapt to the assembly tasks of various parts with different shapes and specifications, which enhances the applicability of the equipment and meets diverse assembly needs. Attached Figure Description

[0014] Figure 1 This is a flowchart of the invention;

[0015] Figure 2 This is a flowchart of the deep reinforcement learning process of this invention;

[0016] Figure 3 This is an isometric drawing of the overall structure of the present invention;

[0017] Figure 4 This is a left view of the overall structure of the present invention;

[0018] Figure 5 This is a schematic diagram of the body of the present invention;

[0019] Figure 6 This is a schematic diagram of the material conveying module of the present invention;

[0020] Figure 7 This is a schematic diagram of the assembly execution module of the present invention;

[0021] Figure 8 This is a schematic diagram of the vision sensor of the present invention;

[0022] Figure 9 This is a schematic diagram of the vibration sensor of the present invention;

[0023] Figure 10 This is a schematic diagram of the torque sensor of the present invention;

[0024] Figure 11This is a schematic diagram of the displacement sensor of the present invention;

[0025] The following components are labeled in the diagram: Body 1, Casters 1-1, Body Outer Shape 1-2, Crossbeam I 1-3, Crossbeam II 1-4, Longitudinal Beam I 1-5, Longitudinal Beam II 1-6, Longitudinal Beam III 1-7, Guide Rail I 1-8, Vertical Beam 1-9, Support Base 1-10, Top Beam 1-11, Material Conveying Module 2, Guide Rail II 2-1, Hopper 2-2, Pushing Component 2-3, Elastic Drive Component 2-4, Parts to be Assembled 2-5, Connector 2-6, Reset Component 2-7, Assembly Execution Module 3, Drive Slide 3-1, Electric Wrench 3-2, Compression Spring 3-3, Sensing and Control Module 4, Vision Sensor 4-1, Vibration Sensor 4-2, Torque Sensor 4-3, Displacement Sensor 4-4, Controller 4-5. Detailed Implementation

[0026] The present invention will be further described below with reference to the accompanying drawings and embodiments, but the scope of the present invention is not limited to the description.

[0027] like Figures 1-11 As shown, an adaptive assembly device based on multimodal perception and deep reinforcement learning includes a body 1, a material conveying module 2, an assembly execution module 3, and a sensing and control module 4. The material conveying module 2 is installed on the guide rails I1-8 of the body 1 to achieve rapid material loading. The multimodal sensor group in the sensing and control module 4 collects visual images, vibration signals, torque data, and displacement information of the assembly process in real time. The controller 4-5 in the sensing and control module 4 preprocesses and fuses the collected multimodal data and inputs it into the deep reinforcement learning model for state recognition and decision analysis. Based on the output of the deep reinforcement learning model, an assembly strategy is dynamically generated to control the electric wrench 3-2 and the drive slide 3-11 in the assembly execution module 3 to perform assembly actions. Based on the real-time feedback data of the multimodal sensor group, a closed-loop adaptive control is formed to complete the assembly task.

[0028] Furthermore, the machine body 1 includes casters 1-1, a machine body outline 1-2, crossbeams I1-3 and II1-4, longitudinal beams I1-5, II1-6, and III1-7, guide rails I1-8, vertical beams 1-9, support bases 1-10, and top beams 1-11. Four self-locking casters are installed at the bottom of the machine body outline 1-2 to enable overall movement and fixation of the device. Crossbeams I1-3, II1-4, I1-5, II1-6, and III1-7 are sequentially installed aligned on the end faces of the machine body outline 1-2 to form a rigid support structure. Guide rails I1-8 are installed on the upper surface between longitudinal beams I1-5 and II1-6, serving as the material conveying module 2. Providing a sliding guide, vertical beam 1-9 is installed on the lower end face of longitudinal beams II1-6 and III1-7. Support base 1-10 is fixed on the upper end face of the bottom of the body shape 1-2 to fix the assembly execution module 3 and ensure structural stability during assembly. Top beam 1-11 is equipped with vision sensor 4-1 to provide a visual monitoring angle for the assembly station. Universal wheel 1-1 is connected to body shape 1-2 by screws. The contact surfaces between crossbeam I1-3, crossbeam II1-4, longitudinal beam I1-5, longitudinal beam II1-6, longitudinal beam III1-7, guide rail I1-8, body shape 1-2, vertical beam 1-9, support base 1-10, and top beam 1-11 are connected by welding.

[0029] Furthermore, the material conveying module 2 includes a guide rail II 2-1, a hopper 2-2, a pushing component 2-3, an elastic drive component 2-4, a part to be assembled 2-5, a connector 2-6, and a reset component 2-7. The guide rail II 2-1 is installed at the bottom of the hopper 2-2 to provide low-friction sliding guidance for the pushing component 2-3. The hopper 2-2 stores the part to be assembled 2-5 and the connector 2-6. The pushing component 2-3 pushes the parts to the assembly station under the action of the elastic drive component 2-4. The reset component 2-7 drives the pushing component 2-3 to reset when the hopper 2-2 is replenished, thereby achieving rapid material feeding.

[0030] Furthermore, the assembly execution module 3 includes a drive slide 3-1, an electric wrench 3-2, and a clamping spring 3-3. The drive slide 3-1 is mounted on the vertical beam 1-9 of the machine body 1 by bolts, and its bottom is mounted on the support base 1-10, forming a double fixed structure. This drives the electric wrench 3-2 to achieve lifting and lowering movements. The electric wrench 3-2 is fixed on the movable side of the drive slide 3-1 and serves as the core execution component to output torque. Its output parameters are controlled in real time by the controller 4-5. The clamping spring 3-3 is installed at the bottom of the electric wrench 3-2 to buffer assembly impacts and prevent damage to parts.

[0031] Furthermore, the sensing and control module 4 includes a vision sensor 4-1, a vibration sensor 4-2, a torque sensor 4-3, a displacement sensor 4-4, and a controller 4-5. The vision sensor 4-1 is installed on the top beam 1-11 of the machine body 1, with its lens facing the assembly station, to collect images of the parts for positioning and posture recognition. The vibration sensor 4-2 is installed on the electric wrench 3-2 and the machine body 1 to monitor the vibration spectrum during the assembly process and identify abnormalities such as jamming and unusual noises. The torque sensor 4-3 is installed between the output shaft and the actuator of the electric wrench 3-2 to collect assembly torque data in real time. The displacement sensor 4-4 is installed at the connection between the movable and fixed sides of the drive slide 3-1 to monitor the lifting and lowering displacement of the assembly tool. The sensor is connected to the data acquisition card via a CAN bus. The controller 4-5 is placed in a control cabinet near the operating area. Internally, a deep reinforcement learning model with a DQN+LSTM fusion architecture is deployed. Its data interaction and control process is optimized as follows: The raw data collected by multiple sensors is initially processed by the data acquisition card, including filtering, noise reduction, and format conversion, and then transmitted to the controller 4-5; the controller further performs preprocessing such as normalization and time alignment on the data to form a standardized input vector, which is then fed into the deep learning model; the model infers the optimal control strategy based on historical and real-time data and feeds it back to the controller; the controller generates specific control commands such as driving the slide displacement and electric wrench torque according to the model output, and sends them to each actuator through the industrial bus; during the operation of the actuator, the sensors collect assembly status data in real time and transmit it back to the controller, forming a closed-loop control link of "perception-decision-execution-feedback" to ensure that the assembly strategy dynamically adapts to changes in working conditions.

[0032] Furthermore, the deep reinforcement learning workflow is as follows:

[0033] Multimodal data acquisition: Image information I (RGB format, 640×480 resolution) of the installation position and assembly quality of the parts is acquired by a vision sensor; vibration acceleration x(t) of the parts during loading, unloading and assembly is acquired by a vibration sensor; torque data τ(t) of the applied torque during assembly is acquired by a torque sensor; and displacement data d(t) of the parts is acquired by a displacement sensor.

[0034] 1. Data preprocessing:

[0035] 1.1 Image Data Processing:

[0036] (1) Convert RGB to grayscale, retain brightness information and remove color redundancy.

[0037] ;

[0038] Where R(x,y), G(x,y), and B(x,y) are the red, green, and blue channel pixel values ​​of the original image at coordinates (x,y); (x,y) represents the pixel value after grayscale conversion; the coefficients (0.299,0.587,0.114) are the international standard brightness conversion weights, adapted to the characteristic that the human eye is most sensitive to green.

[0039] (2) Convert the grayscale image to a black and white binary image to highlight the boundary between the part outline and the background:

[0040] ;

[0041] Where: T is the threshold; 255 represents white, and 0 represents black; in the adaptive threshold scenario... It can be calculated using local means: The offset is constant.

[0042] (3) Size adjustment.

[0043] The image is scaled from 640×480 to 224×224 using bilinear interpolation while preserving characteristic geometric relationships. Let the target image pixel count be... The corresponding floating-point coordinates of the original image are ,but:

[0044] ;

[0045] in: Let (x, y) be the integer coordinates of the four neighborhoods around (x, y) in the original image. Weight The feature decays linearly with increasing distance, ensuring a smooth transition of features after scaling.

[0046] (4) Histogram equalization: Divide the image into (8×8) sub-blocks, each sub-block having a size of (28×28); for the gray-level histogram H(k) of each sub-block, if a certain gray-level frequency (H(k)>T) (T is the cropping threshold, take T=2.0), then the excess part is evenly distributed to other gray-levels.

[0047] ;

[0048] in: .

[0049] Cumulative Distribution Function (CDF) Mapping:

[0050] ;

[0051] in: ( (Total number of pixels in the sub-block).

[0052] 1.2 Vibration signal processing:

[0053] (1) Use a low-pass filter H(f) to filter out high-frequency noise:

[0054] ;

[0055] in: The original vibration data (time domain signal); This is the transfer function (frequency domain characteristics) of the Butterworth low-pass filter. This represents the convolution operation; This is the filtered signal.

[0056] (2) For signals containing missing values By effectively constructing an interpolation function, the imputation value is calculated for the missing position t:

[0057] ;

[0058] in: This represents the piecewise cubic Hermitian interpolation algorithm; A time index for valid data points; This is valid vibration data corresponding to the time index.

[0059] (3) Normalize the data to the range [-1, 1]:

[0060] ;

[0061] in: Normalized vibration data : The mean value of vibration data; Standard deviation of vibration data; : minute value.

[0062] 1.3 Torque and Displacement Data Processing:

[0063] (1) Physical quantity conversion: Converting the raw electrical signals (such as voltage, current, or digital pulses) output by the sensor into engineering units with physical meaning:

[0064] ;

[0065] in: Torque calibration factor; The sensor's original voltage value; Zero-point offset voltage; Displacement conversion factor; Encoder pulse counting.

[0066] (2) Moving average filtering suppresses random high-frequency noise in the signal and preserves the true physical change trend.

[0067] τ filtered [n]= 1 M ∑ k=0 M−1 τ [n−k] ;

[0068] in: Window length (dynamically adjustable, default M=5) Current sampling point.

[0069] (3) Calculate displacement velocity and determine the motion state (uniform speed / acceleration / stuck) by the rate of change of displacement.

[0070] v d [n]= d[n]−d[n−1] Δt (Δt=0.01s) ;

[0071] in: Displacement velocity at the nth sampling time (unit: mm / s, reflecting the speed of movement of components or actuators); Displacement sensor reading at the nth sampling time (unit: mm, such as the real-time position of the drive slide or the assembly feed position of the component). The displacement sensor reading at the (n-1)th sampling time (position at the previous time, compared with...) (forming a "current-historical" positional difference) Sampling time interval.

[0072] (4) Normalization process to eliminate dimensional differences and compress the data to the [0,1] interval to meet the input requirements of machine learning models.

[0073] ;

[0074] in: : Minimum / maximum values ​​of historical torque data; The normalized torque value at time t; The raw torque data collected by the torque sensor at time t.

[0075] 2. Feature Engineering:

[0076] 2.1 Vibration signal feature extraction:

[0077] (1) Temporal characteristics:

[0078] Peak acceleration:

[0079] ;

[0080] in: Normalized vibration acceleration signal (eliminating dimensional effects); The maximum absolute value of the signal is taken to reflect the maximum impact intensity during vibration.

[0081] Root mean square value:

[0082] ;

[0083] in: Number of signal sampling points; The sum of squares of the normalized signal.

[0084] Kurtosis coefficient reflects the sharpness of the impulse component in a signal; a higher kurtosis indicates a more concentrated impulse.

[0085] ;

[0086] in: Original vibration signal; Signal mean; Signal standard deviation; Fourth-order central moment (reflects the degree to which the signal deviates from the mean).

[0087] The margin factor reflects the contrast between the impact signal and the background vibration; the larger the ratio, the more prominent the impact is relative to the background.

[0088] ;

[0089] in: Peak acceleration; Root mean square value.

[0090] (2) Frequency domain characteristics:

[0091] FFT spectrum: X(f) = |F{ x normalized (t)}| Main frequency components: f main =arg m f∈[10,500] | (f)| Bandwidth energy ratio: E ratio = ∑ f=200 500 | X(f) | 2 ∑ f=10 500 | X(f) | 2 ;

[0092] in: Fourier transform (converting a time-domain signal to the frequency domain); The amplitude spectrum of a frequency domain signal, where f is the frequency and |X(f)| represents the signal energy amplitude at frequency f; arg m f∈[10,500] : Within the frequency range of 10-500Hz, take the frequency corresponding to the maximum value of |X(f)|; The frequency at which energy is most concentrated in a vibration signal reflects the characteristics of the main vibration source. The total energy within the 200-500Hz frequency band; The total energy within the 10-500Hz frequency band; the ratio reflects the proportion of collision energy in the total vibration energy, and is used to judge the severity of the collision between parts.

[0093] 3. Model Building:

[0094] 3.1 LSTM Subnetwork:

[0095] (1) Input layer: Receives multimodal time series data at each time step t ( ), concatenate them into an input vector:

[0096] Sseq(t)=[x(t),τ(t),d(t)]∈ R 3 ;

[0097] Where: t: time step (forming the sequence) S seq =[ s seq (1),…, s seq (Txt) : Acceleration (in m / s²) collected by the vibration sensor at time t; : Assembly torque (in N·m) collected by the torque sensor at time t; Displacement of the slide table (in mm) acquired by the displacement sensor at time t.

[0098] (2) LSTM layer: Forgotten Gate: f t =σ( W f ⋅[ h t−1 , x t ]+ b f ) Input Gate: i t =σ( W i ⋅[ h t−1 , x t ]+ b i ) C ~ t =tanh( W C ⋅[ h t−1 , x t ]+ b C ) Status Update: C t = f t ⊙ C t−1 + i t ⊙ C ̃ t Output gate: o t =σ( W o ⋅[ h t−1 , x t ]+ b o ) h t = o t ∘tanh( C t ) ;

[0099] in: The output of the forget gate at time t; The weight matrix of the forget gate; The hidden state at any given moment; The bias term of the forget gate; : sigmoid activation function output layer; Input gate output (0~1, controlling the proportion of new information included). : Weight matrix of input gate and candidate cell state; : Corresponding bias term; : Candidate cell state at time t (new information); Hyperbolic tangent activation function; : Cell state at time t (stored for long-term memory); : Cellular state at any given moment; Element-wise product (multiplying elements one by one to filter information); Output gate output (0~1, controlling the output ratio of cell state); The weight matrix and bias of the output gate; : The hidden state at time t (output to the next layer of temporal features).

[0100] (3) Output layer:

[0101] ;

[0102] in: The final output time-series feature vector (including dynamic features of vibration, torque, and displacement). The hidden state of the last time step T; The weight matrix and biases of the fully connected layer; Linear rectified function (activates nonlinear features and suppresses gradient vanishing).

[0103] 3.2 CNN Subnetworks:

[0104] (1) Input layer:

[0105] ;

[0106] (2) Convolutional layer:

[0107] ;

[0108] in: : Feature map output after convolution; K: 3×3 convolution kernel (sliding window, extracting local texture / edge features); Convolution operation (dot product of kernel and local region of image); : Convolutional layer bias term.

[0109] (3) Pooling layer:

[0110] ;

[0111] in: Feature map after pooling (reduced size, retaining key features). Max pooling.

[0112] (4) Output layer:

[0113] ;

[0114] in: : Output visual feature vector (8-dimensional, including information such as part position deviation and attitude angle). : Weight matrix and bias of fully connected layer.

[0115] 3.3DQN Main Network:

[0116] (1) Feature fusion:

[0117] ;

[0118] in: : Fuse feature vectors; Feature splicing.

[0119] (2) Fully connected layer:

[0120] ;

[0121] ;

[0122] in: : Intermediate feature vector; : Weight matrix of fully connected layer; : Corresponding bias term.

[0123] (3) Q-value output layer:

[0124] ;

[0125] in: : Perform an action in state s Q value; Current assembly status; Assembly action; : Weight matrix and bias of the output layer.

[0126] 4. Model Training:

[0127] 4.1 Loss Function:

[0128] DQN updates parameters by minimizing the prediction error of the Q value:

[0129] L= E [(r+γ max a' Q '(s',a')−Q(s,a) ) 2 ] ;

[0130] Where: L: loss value; E [⋅] :expect; : Execute action The immediate reward afterwards; Discount factor (0.9, the weight of future reward decay, with higher weight for closer futures). The next state output by the target network Next action Q value; : The Q value currently output by the network.

[0131] 4.2 Reward Function:

[0132] ;

[0133] in: Weights (0.4, 0.3, 0.3 respectively, set according to task priority).

[0134] Accuracy Bonus (the smaller the displacement deviation, the higher the bonus):

[0135] ;

[0136] in: Actual displacement; Target displacement; (Reward coefficient); (Control the decay rate).

[0137] Torque bonus (rewarded if within the acceptable range):

[0138] ;

[0139] Abnormal penalty (points deducted for vibration exceeding threshold):

[0140] ;

[0141] in: Peak vibration value; Vibration threshold; (Penalty coefficient); : Counting function.

[0142] Minimize the loss using the Adam optimizer:

[0143] ;

[0144] in: Model parameters at time t; Learning rate; Loss function pair The gradient.

[0145] 4.3 Adaptive Decision Making and Optimization:

[0146] Comprehensive scoring model:

[0147] ;

[0148] in: The assembly pass rate output by the CNN; : Probability of anomaly occurrence in LSTM output; S: Final quality score (range [0,1]).

[0149] 5. Dynamic control strategy:

[0150] Visual positioning compensation:

[0151] ;

[0152] in: Position compensation amount; Horizontal deviation compensation coefficient; Vertical deviation compensation coefficient; Horizontal deviation pixel value; Vertical deviation pixel value.

[0153] Dynamic torque adjustment:

[0154] ;

[0155] in: The adjusted new torque value; Reference torque value; 5-point torque gradient; Gradient response coefficient; The symbol for gradient.

[0156] 6. Self-learning model update:

[0157] 6.1 Learning Trigger Mechanism:

[0158] ;

[0159] Where: 1 indicates that a model update is triggered, and 0 indicates that it is not triggered; Current assembly serial number; Indicator function (1 when Si < 0.85, 0 otherwise); summation range: the last 50 assemblies (i from k-49 to k).

[0160] 6.2 Dynamic Update of Data Pool: D new = ;

[0161] in: The training dataset before and after the update; The oldest sample in the data pool; Current sample.

[0162] 6.3 Incremental Learning Formula:

[0163] ;

[0164] in: Model parameters before and after the update; Learning rate; For parameters The gradient; The number of samples updated each time; Loss function (here, cross-entropy loss); The true label of sample i; The predicted score for sample i; L2 regularization coefficient; The L2 norm of the parameter.

[0165] 6.4LSTM Anomaly Detector:

[0166] p err =softmax(W⋅ h T +b)[2] ;

[0167] in: The normalization function outputs a multi-class probability distribution (here, [normal probability, warning probability, abnormal probability]). Weight matrix; The hidden state at the final time step of the LSTM; Bias term; Take the third value from the softmax output (index starts from 0), which is the anomaly probability.

[0168] 6.5 Adaptive threshold adjustment:

[0169] ;

[0170] in: The threshold for quality scoring; High-frequency vibration energy (when E high frequency > 100, the threshold is lowered to improve sensitivity).

[0171] 7. Model fine-tuning and optimization:

[0172] 7.1 CNN Visual Model Update:

[0173] ;

[0174] in: Parameters of the fully connected layer before and after the update; Cross-entropy loss (ytrue is the true qualified label, pok is the predicted qualified probability).

[0175] 7.2 LSTM Timing Model Update:

[0176] ;

[0177] in: LSTM parameters before and after the update; Mean squared error loss (perrpred is the predicted anomaly probability, perrtrue is the true anomaly label).

[0178] 8. Security Deployment Mechanism:

[0179] CRC check:

[0180] ;

[0181] Wherein: CRC32: 32-bit cyclic redundancy check value; XOR operation; The i-th byte of the model parameters; CRC polynomial.

[0182] Furthermore, the material conveying module 2 automatically pushes the parts to be assembled 2-5 and the connectors 2-6 to the assembly station; the vision sensor 4-1 in the sensing and control module 4 then acquires images for positioning and attitude recognition, and sends signals to the controller 4-5; the controller 4-5 controls the drive slide 3-1 to drive the electric wrench 3-2 to descend and clamp the parts, and then the electric wrench 3-2 performs the assembly operation; the torque sensor 4-3, vibration sensor 4-2 and displacement sensor 4-4 synchronously acquire the physical data of the assembly process and feed it back to the controller 4-5; the deep reinforcement learning model in the controller 4-5 dynamically optimizes the control commands in real time based on these multimodal data to achieve adaptive closed-loop assembly control.

[0183] Matters not covered in this invention are common knowledge.

[0184] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the invention. All equivalent changes or modifications made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.

Claims

1. An adaptive assembly device based on multimodal perception and deep reinforcement learning, characterized in that: The assembly includes a body (1), a material conveying module (2), an assembly execution module (3), and a sensing and control module (4). The material conveying module (2) is installed on the guide rail I (1-8) of the body (1) to convey the parts to be assembled (2-5) and the connectors (2-6) to the assembly station. The multimodal sensor group in the sensing and control module (4) collects visual images, vibration signals, torque data and displacement information of the assembly process in real time. The controller (4-5) in the sensing and control module (4) preprocesses and fuses the collected multimodal data and inputs it into a deep reinforcement learning model for state recognition and decision analysis. Based on the output of the deep reinforcement learning model, an assembly strategy is dynamically generated. According to the strategy, the drive slide (3-1) and electric wrench (3-2) in the assembly execution module (3) are controlled to perform assembly actions, and a closed-loop adaptive control is formed based on the real-time feedback data of the multimodal sensor group to complete the assembly task. The assembly execution module (3) includes a drive slide (3-1), an electric wrench (3-2), and a clamping spring (3-3). The fixed side of the drive slide (3-1) is rigidly connected to the side of the vertical beam (1-9) in the body (1) by high-strength bolts. At the same time, the bottom of the drive slide (3-1) is rigidly connected to the top surface of the support base (1-10). The electric wrench (3-2) is fixed on the movable side of the drive slide (3-1), and the clamping spring (3-3) is installed at the bottom of the electric wrench (3-2). The sensing and control module (4) includes a multimodal sensor group and a controller (4-5). The multimodal sensor group includes a vision sensor (4-1), a vibration sensor (4-2), a torque sensor (4-3), and a displacement sensor (4-4). The vision sensor (4-1) is installed on the top beam (1-11) of the machine body (1), with its lens facing the assembly station. It is used to collect images of the parts to be assembled (2-5) and the connectors (2-6) and to perform positioning and posture recognition. The vibration sensor (4-2) is installed on the electric wrench (3-2) of the assembly execution module (3) and on the machine body (1). It is used to monitor the assembly process. The vibration spectrum during the process is used to identify jamming, abnormal noise, and other abnormal states; the torque sensor (4-3) is installed between the output shaft and the execution end of the electric wrench (3-2) to collect assembly torque data in real time; the displacement sensor (4-4) is installed at the connection between the moving side and the fixed side of the drive slide (3-1) of the assembly execution module (3) to monitor the lifting displacement of the assembly tool; the controller (4-5) deploys a deep reinforcement learning model with a DQN+LSTM fusion architecture, receives multi-sensor data through the data acquisition card, and outputs displacement commands for the drive slide (3-1) and torque commands for the electric wrench (3-2) to achieve closed-loop adaptive control.

2. The adaptive assembly device based on multimodal perception and deep reinforcement learning according to claim 1, characterized in that: The machine body (1) includes casters (1-1), a body outline (1-2), crossbeam I (1-3), crossbeam II (1-4), longitudinal beam I (1-5), longitudinal beam II (1-6), longitudinal beam III (1-7), guide rail I (1-8), vertical beam (1-9), support base (1-10), and top beam (1-11); four self-locking casters (1-1) are installed at the bottom of the body outline (1-2), and crossbeam I (1-3), crossbeam II (1-4), longitudinal beam I (1-5), and longitudinal beam II (1-6) are also present. 1-6) The longitudinal beam III (1-7) is installed in sequence on the end face of the fuselage (1-2), the guide rail I (1-8) is installed on the upper end face between the longitudinal beam I (1-5) and the longitudinal beam II (1-6), the vertical beam (1-9) is installed on the lower end face of the longitudinal beam II (1-6) and the longitudinal beam III (1-7), the support base (1-10) is fixed on the upper end face of the bottom of the fuselage (1-2), and the top beam (1-11) is located on the top of the body (1) for installing the vision sensor (4-1) in the sensing and control module (4).

3. The adaptive assembly device based on multimodal perception and deep reinforcement learning according to claim 1, characterized in that: The material conveying module (2) includes a guide rail II (2-1), a hopper (2-2), a pushing component (2-3), an elastic drive component (2-4), and a reset component (2-7). The guide rail II (2-1) is embedded in the inner bottom of the hopper (2-2) along its length. The pushing component (2-3) slides with the guide rail II (2-1) through a slider. The parts to be assembled (2-5) and the connecting parts (2-6) are stacked in the hopper (2-2) in a preset posture and located on the pushing path of the pushing component (2-3). One end of the elastic drive component (2-4) is connected to the inner wall of the hopper (2-2), and the other end abuts against the pushing component (2-3). The reset component (2-7) is a tension spring structure with a stroke limit and its two ends are respectively connected to the pushing component (2-3) and the end of the hopper (2-2).

Citation Information

Patent Citations

  • Robotic system and method for mounting component assembly

    CN118829511A

  • Screw machine control management method and system, electronic equipment and storage medium

    CN120406123A