Modular device automatic assembly positioning and attitude adjustment method and system

CN122807953APending Publication Date: 2026-09-25HANGZHOU KAIYUAN ENVIRONMENTAL PROTECTION ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611302613.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-26
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,现有的模块化设备自动装配技术仍然存在一些明显的缺陷和不足

Benefits of technology

本发明提供的模块化设备自动装配定位与姿态调整方法,通过多角度深度扫描获取实时位姿参数,结合蒙特卡洛树搜索算法构建装配轨迹预测模型,能够准确预测装配过程中出现的碰撞风险,提高了装配的安全性和可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807953A_ABST
    Figure CN122807953A_ABST
Patent Text Reader

Abstract

The application provides a modular equipment automatic assembly positioning and posture adjusting method and system, relates to the technical field of equipment adjustment, and comprises the following steps: acquiring real-time posture parameters of a modular equipment through multi-angle scanning of a depth camera, constructing an assembly track prediction model by using a Monte Carlo tree search algorithm to generate an optimal assembly track, dynamically compensating position and posture deviation values by combining an adaptive prediction encoder, and controlling a mechanical arm to accurately adjust the assembly posture. The application can effectively improve assembly precision and efficiency, reduce collision risks in the assembly process, and realize real-time suppression of environmental disturbances.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to equipment adjustment technology, and more particularly to a method and system for automatic assembly, positioning and attitude adjustment of modular equipment. Background Technology

[0002] With the rapid development of industrial automation, automated assembly technology for modular equipment has been widely applied in manufacturing. This technology utilizes robotic automation systems to achieve precise positioning and assembly of parts, significantly improving production efficiency and product quality. Traditional modular equipment assembly processes rely primarily on manual operation or simple mechanical positioning devices. However, as precision manufacturing demands increasingly higher assembly accuracy, automated assembly systems require greater positioning precision and adaptability. Currently, the field of industrial automated assembly is beginning to employ technologies such as computer vision, intelligent algorithms, and adaptive control to improve the accuracy and stability of the assembly process.

[0003] However, existing modular automated assembly technologies still have some significant drawbacks and shortcomings. First, current technologies lack sufficient positioning accuracy in complex environments, especially under conditions of changing light and surface reflections. Traditional vision systems struggle to acquire accurate pose information, leading to positioning deviations during assembly. Second, existing assembly path planning methods lack dynamic adaptability, mostly relying on preset assembly paths and failing to automatically adjust to the optimal assembly trajectory based on actual working conditions. When encountering obstacles or workpiece position deviations, collisions or assembly failures are prone to occur. Finally, traditional assembly systems have weak resistance to environmental disturbances. Under the influence of external factors such as vibration and temperature changes, they struggle to maintain stable assembly accuracy and cannot achieve real-time compensation for position and orientation deviations, impacting assembly quality and efficiency. Summary of the Invention

[0004] The present invention provides a method and system for automatic assembly, positioning and attitude adjustment of modular equipment, which can solve the problems in the prior art.

[0005] A first aspect of the present invention provides a method for automatic assembly, positioning, and attitude adjustment of modular equipment, comprising: The assembly reference position information of the modular equipment to be assembled is obtained, and the modular equipment to be assembled is scanned from multiple angles using a depth camera to obtain the real-time pose parameters of the modular equipment. An assembly trajectory prediction model is constructed based on the Monte Carlo tree search algorithm. The real-time pose parameters and the assembly reference position information are input into the assembly trajectory prediction model to generate multiple candidate assembly trajectories. The collision risk of each candidate assembly trajectory is evaluated, and the assembly trajectory with the lowest collision risk is selected as the optimal assembly trajectory. Calculate the positional deviation and attitude deviation between the target position and the current position of the modular equipment based on the optimal assembly trajectory; An adaptive predictive encoder is used to dynamically compensate for the position deviation and the attitude deviation. The adaptive predictive encoder optimizes the compensation parameters online based on historical assembly data to achieve real-time suppression of environmental disturbances. The control robot arm adjusts the position and attitude of the modular equipment to be assembled based on the compensated position and attitude deviation values ​​until the modular equipment reaches the assembly reference position.

[0006] The modular device to be assembled is scanned from multiple angles using a depth camera to obtain its real-time pose parameters, including: Multiple depth cameras are arranged to form a ring scanning array. The scanning angle of each depth camera is adaptively adjusted based on the structural feature information of the modular equipment to be assembled, so as to ensure that the scanning field of view of the depth camera completely covers all surfaces of the modular equipment to be assembled. Depth images acquired by each depth camera are collected, and the depth images are preprocessed to obtain feature point cloud data of the modular device to be assembled. The feature point cloud data is input into a pre-trained feature matching network. The feature matching network uses epipolar geometric constraints to match feature points from multiple perspectives, generating six-degree-of-freedom pose parameters for the modular device to be assembled, and outputting the real-time pose parameters of the modular device to be assembled.

[0007] The feature point cloud data is input into a pre-trained feature matching network, which uses epipolar geometry constraints to match feature points from multiple perspectives, generating six-degree-of-freedom pose parameters for the modular device to be assembled, including: The feature point cloud data is aligned with the coordinate system based on the preset pose reference data, and the spatial position information and descriptor information of the feature points are extracted. The matching cost between the spatial location information and the descriptor information is calculated based on epipolar geometric constraints. A reprojection error function for feature point pairs is constructed, and the reprojection error is optimized to obtain the initial matching result of the feature points. The feature point pairs in the initial matching results are filtered, and matching point pairs that do not meet the requirements are removed according to the epipolar geometric constraints; The selected feature point pairs are iteratively optimized. In each iteration, the reprojection error function of the feature point pairs is reconstructed. The reprojection error is minimized to obtain the six-degree-of-freedom pose parameters of the modular device to be assembled.

[0008] An assembly trajectory prediction model is constructed based on the Monte Carlo tree search algorithm. The real-time pose parameters and the assembly reference position information are input into the assembly trajectory prediction model to generate multiple candidate assembly trajectories, including: An assembly trajectory state space is constructed based on the Monte Carlo tree search algorithm. The assembly feasibility is evaluated for each motion state according to the assembly trajectory state space, and a search tree structure for the assembly trajectory is generated. The number of visits and reward values ​​of assembly trajectory nodes are determined based on the search tree structure. An iterative search is performed starting from the root node of the search tree structure. In each iteration, the node with the highest reward value is selected for expansion until the search depth reaches a preset search threshold. Each search path is then output as a candidate assembly trajectory.

[0009] An assembly trajectory state space is constructed based on the Monte Carlo tree search algorithm. Assembly feasibility is evaluated for each motion state according to this state space, generating a search tree structure for the assembly trajectory, including: Using the assembly start position as the root node, an assembly trajectory state space is constructed. Based on the Monte Carlo tree search algorithm, search nodes are generated in the assembly trajectory state space. According to the assembly state of the parent node, the Monte Carlo sampling method is used to generate the candidate assembly state of the child node, and an initial number of visits and a reward value are assigned to each search node. Obtain the geometric constraint information of the equipment to be assembled and its surrounding environment. Based on the geometric constraint information, perform an assembly feasibility assessment on each search node, calculate the collision risk value and the assembly difficulty value, and use the weighted sum of the collision risk value and the assembly difficulty value as the node score. Based on the node score, prune and optimize the search tree to generate the final assembly trajectory search tree structure.

[0010] An adaptive predictive encoder is used to dynamically compensate for the position deviation and attitude deviation values. The adaptive predictive encoder optimizes the compensation parameters online based on historical assembly data to achieve real-time suppression of environmental disturbances, including: An adaptive predictive encoder is constructed. The sequence of compensation parameters from the historical assembly process and the actual assembly error sequence are input into the adaptive predictive encoder to generate initial compensation parameters for compensating the position deviation value and the attitude deviation value. The initial compensation parameters are applied to the current assembly process to obtain the compensated real-time assembly error; The initial compensation parameters are optimized and adjusted online based on the real-time assembly error until the real-time assembly error is less than a preset error threshold, thereby obtaining the final compensation parameters and achieving real-time suppression of environmental disturbances.

[0011] A second aspect of the present invention provides an automatic assembly positioning and attitude adjustment system for modular equipment, comprising: The first unit is used to obtain the assembly reference position information of the modular equipment to be assembled, and to use a depth camera to perform multi-angle scanning of the modular equipment to be assembled to obtain the real-time pose parameters of the modular equipment. The second unit is used to construct an assembly trajectory prediction model based on the Monte Carlo tree search algorithm. The real-time pose parameters and the assembly reference position information are input into the assembly trajectory prediction model to generate multiple candidate assembly trajectories. The collision risk of each candidate assembly trajectory is evaluated, and the assembly trajectory with the lowest collision risk is selected as the optimal assembly trajectory. The third unit is used to calculate the position deviation and attitude deviation values ​​between the target position and the current position of the modular equipment based on the optimal assembly trajectory. The fourth unit is used to dynamically compensate the position deviation value and the attitude deviation value using an adaptive predictive encoder. The adaptive predictive encoder optimizes the compensation parameters online based on historical assembly data to achieve real-time suppression of environmental disturbances. The fifth unit is used to control the robotic arm to adjust the position and attitude of the modular equipment to be assembled according to the compensated position deviation value and attitude deviation value, until the modular equipment reaches the assembly reference position.

[0012] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0013] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0014] The beneficial effects of this application are as follows: The modular equipment automatic assembly positioning and attitude adjustment method provided by the present invention obtains real-time pose parameters through multi-angle depth scanning and constructs an assembly trajectory prediction model by combining Monte Carlo tree search algorithm, which can accurately predict the collision risk that occurs during the assembly process and improve the safety and reliability of assembly.

[0015] Based on the adaptive predictive encoder, the method dynamically compensates for position and attitude deviations. It can optimize compensation parameters in real time based on historical assembly data, effectively suppress the influence of environmental disturbances, significantly improve assembly accuracy, and adapt to complex and changing working environments.

[0016] This invention realizes the intelligent and automated assembly process of modular equipment, which can complete high-precision assembly tasks without human intervention, reduce labor costs, improve production efficiency, and the system has good adaptability, which can be widely used in the intelligent manufacturing field of various modular equipment. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the automatic assembly, positioning, and attitude adjustment method for modular equipment according to an embodiment of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0020] Figure 1 This is a flowchart illustrating the automatic assembly, positioning, and attitude adjustment method for modular equipment according to an embodiment of the present invention. Figure 1 As shown, the method includes: The assembly reference position information of the modular equipment to be assembled is obtained, and the modular equipment to be assembled is scanned from multiple angles using a depth camera to obtain the real-time pose parameters of the modular equipment. An assembly trajectory prediction model is constructed based on the Monte Carlo tree search algorithm. The real-time pose parameters and the assembly reference position information are input into the assembly trajectory prediction model to generate multiple candidate assembly trajectories. The collision risk of each candidate assembly trajectory is evaluated, and the assembly trajectory with the lowest collision risk is selected as the optimal assembly trajectory. Calculate the positional deviation and attitude deviation between the target position and the current position of the modular equipment based on the optimal assembly trajectory; An adaptive predictive encoder is used to dynamically compensate for the position deviation and the attitude deviation. The adaptive predictive encoder optimizes the compensation parameters online based on historical assembly data to achieve real-time suppression of environmental disturbances. The control robot arm adjusts the position and attitude of the modular equipment to be assembled based on the compensated position and attitude deviation values ​​until the modular equipment reaches the assembly reference position.

[0021] In one optional implementation, a depth camera is used to perform multi-angle scanning of the modular device to be assembled to obtain the real-time pose parameters of the modular device, including: Multiple depth cameras are arranged to form a ring scanning array. The scanning angle of each depth camera is adaptively adjusted based on the structural feature information of the modular equipment to be assembled, so as to ensure that the scanning field of view of the depth camera completely covers all surfaces of the modular equipment to be assembled. Depth images acquired by each depth camera are collected, and the depth images are preprocessed to obtain feature point cloud data of the modular device to be assembled. The feature point cloud data is input into a pre-trained feature matching network. The feature matching network uses epipolar geometric constraints to match feature points from multiple perspectives, generating six-degree-of-freedom pose parameters for the modular device to be assembled, and outputting the real-time pose parameters of the modular device to be assembled.

[0022] When deploying multiple depth cameras to form a circular scanning array, a depth camera with a resolution of 1280×720 pixels, a frame rate of 30fps, and a depth measurement range of 0.5-4.5 meters can be selected. In practical applications, depending on the size of the modular equipment to be assembled, up to eight depth cameras can be deployed, with an angle of approximately 45 degrees between adjacent cameras, forming a complete circular scanning array. These depth cameras are mounted on adjustable supports, the height of which can be adjusted within the range of 1.0-2.0 meters to accommodate modular equipment of different heights.

[0023] Based on the structural feature information of the modular equipment to be assembled, the scanning angles of each depth camera are adaptively adjusted. The system extracts the 3D structural features of the target equipment from a pre-established CAD model database of modular equipment, including its geometric dimensions, surface irregularities, and edge features. Then, based on the extracted feature information, the system calculates the optimal scanning angle for each depth camera. For example, for areas with complex surface structures, such as grooves and holes, the camera's pitch angle can be adjusted to 15-30 degrees to ensure these areas are fully captured; for flat surface areas, the camera's pitch angle can be maintained within the range of 0-10 degrees.

[0024] Before performing a full scan, a virtual scan simulation is conducted to check if the current camera configuration can completely cover the device surface. If coverage blind spots exist, the system will automatically calculate and suggest adjustments to the position and angle of specific cameras. For example, if 20% of the bottom of the device is detected as uncovered, the system will instruct the tilt angle of the two adjacent cameras to be increased from 5 degrees to 25 degrees to ensure complete coverage. In practical applications, the camera position adjustment accuracy can reach ±5mm, and the angle adjustment accuracy can reach ±1 degree, ensuring scan quality.

[0025] When acquiring depth images from each depth camera, all depth cameras are triggered synchronously, with the acquisition frequency set to 10Hz to ensure real-time performance. To reduce ambient light interference, diffuse light sources can be placed around the scanning area, with the illumination intensity controlled between 800-1000 lux. The acquired raw depth images have a resolution of 1280×720 pixels and a depth accuracy of ±2mm.

[0026] The depth image is preprocessed to obtain feature point cloud data for the modular device to be assembled. Gaussian filtering is applied to the depth image, and a 5×5 convolution kernel is used to eliminate noise. Then, depth thresholding is performed to remove points that are too close to the camera (less than 0.5 meters) or too far (greater than 4.5 meters), which helps to remove background and outlier points. Next, voxel downsampling is performed to mesh the original point cloud with a grid size of 5mm×5mm×5mm, retaining only one point within each voxel to reduce data volume and maintain feature integrity. Subsequently, normal vector estimation is performed, and principal component analysis is conducted on the local neighborhood (radius of 10mm) of each point to calculate the surface normal vector. Finally, feature extraction is performed using the FPFH (FastPointFeatureHistograms) algorithm to calculate the feature descriptor for each point, with a feature dimension of 33, describing the local geometric characteristics of the point cloud. Through these preprocessing steps, the original point cloud data (typically containing millions of points) is reduced to tens of thousands of high-quality feature points.

[0027] When the feature point cloud data is input into a pre-trained feature matching network, the network employs a deep learning architecture, including an encoder-decoder structure. The encoder consists of four convolutional layers, each using 64, 128, 256, and 512 3×3 filters with a stride of 2; the decoder uses corresponding deconvolutional layers to recover the spatial dimensions of the feature map. The network is pre-trained on a large-scale industrial parts dataset containing 10,000 modular device samples with different poses. During training, a batch size of 16, a learning rate of 0.0001, and 100 training epochs are used.

[0028] The feature matching network employs epipolar geometric constraints to match feature points from multiple perspectives, calculating the fundamental matrix between adjacent perspectives and establishing epipolar constraint relationships. For a feature point in any perspective, when searching for candidate matching points in other perspectives, the search range is limited to the vicinity of the corresponding epipolar line (distance threshold set to 2 pixels), significantly improving matching efficiency and accuracy. The network outputs a matching confidence score for each pair of feature points, and the system retains matching pairs with a confidence score greater than 0.85.

[0029] When generating the six-DOF pose parameters of the modular device to be assembled, the RANSAC (Random Sample Consensus) algorithm is used to estimate the relative transformation matrix between cameras based on high-quality feature matching point pairs. In practical applications, the number of iterations is set to 1000, and the interior point threshold is set to 3mm, which can effectively filter out abnormal matches. Subsequently, all relative transformation matrices are fused through a global optimization algorithm to construct a consistent global coordinate system, and the six-DOF pose parameters (three translation components and three rotation angles) of the device relative to the reference coordinate system are calculated.

[0030] The system outputs real-time pose parameters of the modular device to be assembled, including translational components in the X, Y, and Z directions (accuracy better than ±2mm) and three rotational angles around the X, Y, and Z axes (accuracy better than ±0.5 degrees). These pose parameters are transmitted to the assembly control system in real-time in JSON format, with an update frequency of 10Hz, ensuring precise control of the assembly process. The latency of the entire scanning and pose calculation process is controlled within 100 milliseconds, meeting real-time control requirements.

[0031] In one optional implementation, the feature point cloud data is input into a pre-trained feature matching network, which uses epipolar geometric constraints to match feature points from multiple viewpoints, generating six-degree-of-freedom pose parameters for the modular device to be assembled, including: The feature point cloud data is aligned with the coordinate system based on the preset pose reference data, and the spatial position information and descriptor information of the feature points are extracted. The matching cost between the spatial location information and the descriptor information is calculated based on epipolar geometric constraints. A reprojection error function for feature point pairs is constructed, and the reprojection error is optimized to obtain the initial matching result of the feature points. The feature point pairs in the initial matching results are filtered, and matching point pairs that do not meet the requirements are removed according to the epipolar geometric constraints; The selected feature point pairs are iteratively optimized. In each iteration, the reprojection error function of the feature point pairs is reconstructed. The reprojection error is minimized to obtain the six-degree-of-freedom pose parameters of the modular device to be assembled.

[0032] The implementation method of multi-view feature point matching with epipolar geometric constraints in feature matching network will explain in detail how to input feature point cloud data into a pre-trained feature matching network and match multi-view feature points through epipolar geometric constraints to generate six-degree-of-freedom pose parameters of the modular device to be assembled.

[0033] Acquire feature point cloud data, which contains point cloud information collected from multiple perspectives of the modular device to be assembled. This point cloud data can be acquired through sensing devices such as depth cameras or LiDAR. For example, in a practical application, a robotic arm module can be scanned from four different angles, acquiring approximately 5,000 feature points from each perspective, forming a feature point cloud dataset of approximately 20,000 points in total.

[0034] Before inputting the acquired feature point cloud data into the pre-trained feature matching network, coordinate system alignment preprocessing is required. Specifically, the system aligns the feature point cloud data to a coordinate system based on preset pose reference data. The preset pose reference data is the standard reference coordinate system information determined during the system calibration phase. In practical applications, the workbench plane can be selected as the XY plane, and the origin position can be defined. For example, the lower left corner of the assembly area can be set as the origin, with the X-axis along the long side of the workbench, the Y-axis along the wide side of the workbench, and the Z-axis perpendicular to the workbench and upwards. All feature point cloud data are transformed to this unified coordinate system through rigid body transformation; the transformation matrix can be determined using three or more calibration points. After coordinate system alignment, the spatial position information of each feature point is extracted, including three-dimensional coordinates (x, y, z) and descriptor information. Descriptor information is a high-dimensional vector characterizing the local geometric properties of the feature point, typically with dimensions of 128 or 256, used for subsequent feature matching.

[0035] After coordinate system alignment and feature information extraction, the matching cost between spatial location information and descriptor information is calculated based on epipolar geometric constraints. Epipolar geometric constraints are geometric relationships in multi-view imaging, stipulating that a point in one view must lie on the corresponding epipolar line in another view. In practical applications, this constraint relationship is expressed by calculating the fundamental matrix or the essential matrix. For the correspondence of feature points between any two views i and j, the distance from the point to the epipolar line can be calculated as a geometric consistency measure. Simultaneously, the cosine similarity or Euclidean distance between feature descriptors is calculated as an appearance similarity measure. The geometric consistency measure and the appearance similarity measure are considered together to construct a matching cost function. For example, the geometric consistency weight can be set to 0.7, and the appearance similarity weight to 0.3, and the matching cost can be calculated comprehensively.

[0036] Based on the calculated matching cost, a reprojection error function for feature point pairs is constructed. This function represents the distance between the actual observed point and the feature point after projecting it from one viewpoint to another using the currently estimated transformation matrix. Specifically, squared Euclidean distance can be used as the error metric. For matching m feature point pairs across n viewpoints, an overall reprojection error function is constructed. The reprojection error is then optimized using algorithms such as gradient descent to obtain the initial matching results for the feature points. In practical applications, a maximum number of iterations of 100 and a convergence threshold of 0.01 can be set to balance computational efficiency and accuracy.

[0037] The initial matching results contained erroneous matches, requiring further filtering. Unsuitable matching point pairs were removed based on epipolar geometric constraints. Specific filtering criteria included: a point-to-epidural distance threshold test (set to a threshold of 2 pixels); a descriptor similarity threshold test (set to a similarity threshold of 0.85); and a local consistency test to check for consistent matching patterns within the neighborhood of the matching point pairs. These filtering criteria effectively removed abnormal matches, improving the accuracy of subsequent pose estimation. In a real-world example, the initial matching yielded 3000 feature point correspondences; after filtering, approximately 2500 high-quality matches were retained.

[0038] The selected feature point pairs are iteratively optimized to accurately solve for the six-DOF pose parameters of the modular device to be assembled. In each iteration, the reprojection error function of the feature point pair is reconstructed based on the currently estimated pose parameters. Robust optimization methods, such as the Levenberg-Marquardt algorithm, are used to minimize the reprojection error. During iterative optimization, a dynamic weight adjustment strategy can be employed to assign smaller weights to matching point pairs with larger reprojection errors to reduce the impact of outliers. The iteration termination condition is set as the parameter update amount being less than a preset threshold (e.g., 0.001) or reaching the maximum number of iterations (e.g., 50). The final transformation matrix obtained is the six-DOF pose parameters of the modular device to be assembled relative to the reference coordinate system, including three translation parameters (tx, ty, tz) and three rotation parameters (rx, ry, rz). In practical applications, the pose estimation accuracy can achieve a translation error of less than 1 mm and a rotation error of less than 0.5 degrees, meeting the requirements for precision assembly.

[0039] In one optional implementation, an assembly trajectory prediction model is constructed based on the Monte Carlo tree search algorithm. The real-time pose parameters and the assembly reference position information are input into the assembly trajectory prediction model to generate multiple candidate assembly trajectories, including: An assembly trajectory state space is constructed based on the Monte Carlo tree search algorithm. The assembly feasibility is evaluated for each motion state according to the assembly trajectory state space, and a search tree structure for the assembly trajectory is generated. The number of visits and reward values ​​of assembly trajectory nodes are determined based on the search tree structure. An iterative search is performed starting from the root node of the search tree structure. In each iteration, the node with the highest reward value is selected for expansion until the search depth reaches a preset search threshold. Each search path is then output as a candidate assembly trajectory.

[0040] The real-time pose parameters of the equipment to be assembled include three-dimensional spatial position coordinates and Euler angle attitude parameters. The assembly reference position information includes the target assembly position coordinates and target assembly attitude parameters. During the construction of the assembly trajectory state space, the current pose of the equipment to be assembled is used as the root node of the search tree, and child node states are generated in a six-degree-of-freedom space through Monte Carlo sampling. For each sampled state, the motion increment from the parent node to the child node is calculated according to the kinematic constraints of the equipment, ensuring that the motion increment meets the maximum speed and acceleration limits of the equipment.

[0041] When assessing the assembly feasibility for each motion state, two aspects need to be considered: collision detection and assembly constraints. Collision detection is achieved by constructing a three-dimensional envelope of the equipment to be assembled and its surrounding environment. A collision risk is determined when the minimum distance between two envelopes is less than a safety threshold. Assembly constraint assessment is based on the relative positional relationship of assembly features, calculating the distance and angular deviations between mating surfaces, and normalizing the deviation values ​​to obtain the constraint satisfaction.

[0042] The search tree structure employs an expansion strategy driven by both node visit counts and node reward values. Node visit counts are initialized to 1, and node reward values ​​are initialized to 0. When expanding a node, unvisited child nodes are prioritized. Once all child nodes have been visited, the expansion direction is selected based on the node's overall score. The node's overall score consists of the node reward value and the exploration item (visit count), with the exploration item having a weighting coefficient of 0.5.

[0043] During the iterative expansion of the search tree, starting from the root node, the node with the highest overall score is selected for expansion each time. For the selected node, multiple candidate child nodes are generated through Monte Carlo sampling, and the assembly feasibility of each child node is evaluated. The evaluation result of the child node is used as an immediate reward, multiplied by a decay factor related to the node depth, and then passed to the parent node to update the parent node's cumulative reward value.

[0044] When the search depth reaches a preset threshold or all feasible expansion directions have been explored, candidate assembly trajectories are extracted from the search tree. The extraction process starts from the root node, selecting the child node with the highest reward value each time, until a leaf node is reached, and the sequence of nodes traversed is taken as a candidate trajectory. Through multiple extraction processes, multiple different candidate assembly trajectories are generated, each trajectory containing a complete sequence of pose parameters.

[0045] To ensure the smoothness of the assembly trajectory, the extracted candidate trajectories undergo post-processing. Cubic spline interpolation is used to smooth the position trajectory, and spherical linear interpolation is used to smooth the attitude trajectory. The smoothed trajectories still need to meet assembly feasibility constraints; trajectories that do not meet these constraints are discarded. The final output contains no more than five candidate assembly trajectories, and each trajectory is guaranteed to meet both kinematic constraints and assembly feasibility requirements.

[0046] In practical applications, during the movement of a part to be assembled from the workbench to the assembly station, three feasible assembly trajectories were generated using the method described above. These trajectories all avoided fixtures and other components in the workspace, ensuring collision safety during the movement. Near the assembly station, the trajectories gradually converged to the preset assembly path, ensuring the part could smoothly enter the assembly position. The maximum positional deviation during the assembly process was kept within 0.5 mm, and the maximum attitude deviation was kept within 0.5 degrees, meeting the assembly accuracy requirements.

[0047] In one optional implementation, an assembly trajectory state space is constructed based on a Monte Carlo tree search algorithm. Assembly feasibility is evaluated for each motion state according to the assembly trajectory state space, and the resulting search tree structure for the assembly trajectory includes: Using the assembly start position as the root node, an assembly trajectory state space is constructed. Based on the Monte Carlo tree search algorithm, search nodes are generated in the assembly trajectory state space. According to the assembly state of the parent node, the Monte Carlo sampling method is used to generate the candidate assembly state of the child node, and an initial number of visits and a reward value are assigned to each search node. Obtain the geometric constraint information of the equipment to be assembled and its surrounding environment. Based on the geometric constraint information, perform an assembly feasibility assessment on each search node, calculate the collision risk value and the assembly difficulty value, and use the weighted sum of the collision risk value and the assembly difficulty value as the node score. Based on the node score, prune and optimize the search tree to generate the final assembly trajectory search tree structure.

[0048] In this embodiment, the assembly trajectory state space is a multi-dimensional space that includes all states during the assembly process. For a six-degree-of-freedom assembly task, the state space includes six dimensions: position coordinates (x, y, z) and attitude angles (α, β, γ). The assembly trajectory is defined as a path from the starting position to the target position in this state space.

[0049] When constructing the assembly trajectory search tree, the assembly starting position is set as the root node. The state of the root node can be represented as S0=(x0, y0, z0, α0, β0, γ0), where (x0, y0, z0) represent the starting position coordinates, and (α0, β0, γ0) represent the starting attitude angles. Initially, the root node is assigned 1 visit N(S0) and 0 reward Q(S0).

[0050] Based on the Monte Carlo tree search algorithm, the following four phases are executed iteratively: selection, expansion, simulation, and backpropagation. In the selection phase, starting from the root node, the UCB1 (Upper Confidence Bound 1) formula is used to balance exploration and utilization, selecting the optimal child node until a leaf node is reached. The calculation of the UCB1 value considers the node's current evaluation value and visit frequency, allowing the algorithm to focus on nodes with high evaluation values ​​while also exploring nodes with fewer visits. Specifically, for node i, its UCB1 value equals the node's average reward plus an exploration term, which is inversely proportional to the node's visit frequency and directly proportional to the total visit frequency of its parent nodes, multiplied by an exploration constant C (in this embodiment, C is set to 1.414).

[0051] During the expansion phase, child nodes are generated from the current leaf node. The generation of child nodes uses a Monte Carlo sampling method, which starts from the current node state Sᵢ and randomly samples multiple candidate states within a preset range. For example, for position coordinates, sampling can be performed uniformly within a spherical space around the current position; for attitude angles, sampling can be performed within an angle space near the current attitude. In practical applications, the sampling range is set as follows: position coordinates within 5 cm of the current position, and attitude angles within 10 degrees of the current attitude. Each expansion generates 10 candidate child nodes, and each child node is assigned an initial number of visits N(Sᵢ). j )=1, initial reward value Q(S) j )=0.

[0052] For each generated candidate child node, an assembly feasibility assessment is performed, considering two key factors: collision risk value and assembly difficulty value. The collision risk value R(S) j The collision risk value is determined by calculating the minimum distance between the equipment to be assembled and its surrounding environment. When the minimum distance is less than the safety threshold (set to 1 cm in this embodiment), the collision risk value increases as the distance decreases; when the minimum distance is greater than the safety threshold, the collision risk value is 0. Specifically, when the minimum distance d is less than the safety threshold, the collision risk value R(S) is... j ) = (1 - d / safety threshold) 2 When d is greater than or equal to the safety threshold, R(S) j )=0.

[0053] Assembly difficulty value D(S) jThe assembly difficulty mainly considers the complexity and operational difficulty of the assembly path, which consists of the following factors: path length, number of direction changes, and difficulty of navigating narrow passages. In this embodiment, the assembly difficulty value is calculated as follows: First, the Euclidean distance from the current node to the target position is calculated, denoted as dgoal; second, the cumulative turning angle θtotal of the current path is calculated; finally, the passage narrowness factor η is calculated based on environmental constraints (η ranges from 0 to 1, with a larger η indicating a narrower passage). The assembly difficulty value D(S) j ) = w1×dgoal + w2×θtotal + w3×η, where w1=0.5, w2=0.3, and w3=0.2 are weighting coefficients.

[0054] Node's overall score (Score) j The score is obtained by weighting the collision risk value and the assembly difficulty value: Score(S j )=λ1×R(S j )+λ2×D(S j ), where λ1=0.6 and λ2=0.4 are weighting coefficients. The lower the score, the better the node.

[0055] During the simulation phase, a randomized policy simulation is performed starting from the newly expanded child nodes until a termination condition is met (such as reaching the target position or reaching the maximum simulation depth). In this embodiment, the maximum simulation depth is set to 100 steps. During the simulation, an action (a small change in position and attitude) is randomly selected at each step, and the immediate reward for that step is calculated. The immediate reward is inversely proportional to the node score, i.e., reward = -Score(S). The total simulation reward is the sum of the immediate rewards for all steps.

[0056] During the backpropagation phase, the total reward obtained from the simulation is backpropagated from the leaf nodes to the root node, updating the visit count and reward value of all nodes on the path. For each node S on the path, its visit count N(S) = N(S) + 1 is updated, and its reward value Q(S) = (Q(S) × (N(S) - 1) + reward) / N(S) is updated.

[0057] By repeatedly executing the above four stages, the search tree continuously grows and optimizes. The algorithm stops when the preset number of iterations (10,000 in this embodiment) or the computation time limit (60 seconds in this embodiment) is reached. To further optimize the search tree, the system performs pruning based on node scores. The pruning rule is: if the score of a node exceeds a preset threshold (0.8 in this embodiment), or if there are no nodes with scores lower than the threshold in the subtree of that node, then that node and its subtree are deleted.

[0058] The path with the lowest score from the root node to the leaf node is selected as the final assembly trajectory, which is represented as a series of state points {S0, S1, S2, ..., S}. n}, where S0 is the initial state, S n This is the target state.

[0059] Experimental data show that in a standard assembly task, this method can reduce the risk of collision by an average of 42% and the assembly time by 25%, which is a significant advantage over traditional methods.

[0060] In one optional implementation, an adaptive predictive encoder is used to dynamically compensate for the position deviation value and the attitude deviation value. The adaptive predictive encoder optimizes the compensation parameters online based on historical assembly data to achieve real-time suppression of environmental disturbances, including: An adaptive predictive encoder is constructed. The sequence of compensation parameters from the historical assembly process and the actual assembly error sequence are input into the adaptive predictive encoder to generate initial compensation parameters for compensating the position deviation value and the attitude deviation value. The initial compensation parameters are applied to the current assembly process to obtain the compensated real-time assembly error; The initial compensation parameters are optimized and adjusted online based on the real-time assembly error until the real-time assembly error is less than a preset error threshold, thereby obtaining the final compensation parameters and achieving real-time suppression of environmental disturbances.

[0061] When constructing an adaptive predictive encoder, a neural network structure with an input layer, hidden layers, and an output layer needs to be designed. The input layer receives the sequence of compensation parameters and the actual assembly error sequence from the historical assembly process. The hidden layer contains multiple neurons for feature extraction, and the output layer generates initial compensation parameters to compensate for position and pose deviations. Specifically, the input layer has 12 nodes, corresponding to the compensation parameters and assembly errors for six degrees of freedom; the hidden layer uses a two-layer structure with 24 and 18 nodes respectively; and the output layer has 6 nodes, corresponding to the compensation parameters for six degrees of freedom. The ReLU activation function is chosen for the neural network, which helps to solve the gradient vanishing problem during training.

[0062] When training the adaptive predictive encoder, at least 200 sets of historical assembly data are collected. Each set includes the compensation parameters used in the assembly process and the corresponding actual assembly error. These data are divided into training and validation sets in an 8:2 ratio. The training process uses a mini-batch gradient descent method with a batch size of 16. The initial learning rate is set to 0.001, and a learning rate decay strategy is used, reducing the learning rate to 0.9 times the original value every 50 training epochs. The training process continues until the error on the validation set no longer decreases significantly, which typically requires 300-500 training epochs.

[0063] For real-world assembly scenarios, training samples should cover a variety of environmental disturbances, including temperature variations (from 15°C to 35°C), vibration interference (frequency 0-100Hz, amplitude 0-2mm), and electromagnetic interference. This training allows the adaptive predictive encoder to adapt to assembly tasks under different environments, improving the system's robustness.

[0064] After training, the adaptive predictive encoder can generate initial compensation parameters to compensate for position and attitude deviations based on the input sequence of compensation parameters from the historical assembly process and the actual assembly error sequence. For example, for an X-axis position deviation of 0.35mm, a Y-axis position deviation of -0.28mm, a Z-axis position deviation of 0.42mm, an X-axis attitude deviation of 0.15°, a Y-axis attitude deviation of -0.21°, and a Z-axis attitude deviation of 0.08°, the adaptive predictive encoder can generate initial compensation parameters of -0.37mm for the X-axis, 0.30mm for the Y-axis, -0.45mm for the Z-axis, -0.16° for the X-axis, 0.23° for the Y-axis, and -0.09° for the Z-axis.

[0065] Applying initial compensation parameters to the current assembly process is achieved by setting position and attitude offsets in the robot control system. The generated initial compensation parameters are then superimposed onto the robot's target pose in a reverse compensation manner, correcting the robot's trajectory. For example, if the target position is (100mm, 200mm, 300mm) and the initial X-axis compensation value is -0.37mm, the corrected X-axis target position will be 99.63mm. After the compensation parameters are applied, the real-time assembly error is obtained using precision sensors (such as laser trackers, vision systems, or force sensors). Laser trackers offer a measurement accuracy of ±0.015mm, vision systems approximately ±0.05mm, and force sensors approximately ±0.1N and ±0.01Nm, respectively.

[0066] After obtaining the real-time assembly errors, the initial compensation parameters are optimized and adjusted online based on these errors. The optimization process adopts an incremental adjustment strategy, with the step size of each adjustment proportional to the real-time assembly error, but a maximum step size limit is set to ensure system stability. For example, the adjustment step size coefficient for position error is set to 0.9, and the maximum adjustment step size is 0.2mm; the adjustment step size coefficient for attitude error is set to 0.85, and the maximum adjustment step size is 0.1°. In the aforementioned example, if the actual assembly error of the X-axis after compensation is 0.08mm, then the adjustment amount of the X-axis compensation value is 0.08mm × 0.9 = 0.072mm, and the adjusted X-axis compensation value becomes -0.37mm - 0.072mm = -0.442mm.

[0067] This optimization and adjustment continues until the real-time assembly error is less than a preset error threshold. The position error threshold is typically set to ±0.05mm, and the attitude error threshold is typically set to ±0.03°. When the real-time assembly error of all six degrees of freedom is less than the corresponding preset error threshold, the current compensation parameters become the final compensation parameters, and the system completes real-time suppression of environmental disturbances.

[0068] In practical applications, this method can effectively cope with various environmental disturbances. For example, when the ambient temperature rises from 20°C to 30°C, the assembly system experiences thermal expansion, causing a positional deviation of approximately 0.15mm and an attitude deviation of 0.08°. Through the dynamic compensation of the aforementioned adaptive predictive encoder, the assembly error can be controlled within ±0.03mm and ±0.02°. Similarly, when there is a 50Hz, 1mm amplitude vibration disturbance in the environment, the system can reduce the assembly error from the original ±0.25mm and ±0.15° to within ±0.04mm and ±0.025°.

[0069] In this way, the adaptive predictive encoder can dynamically optimize compensation parameters based on historical assembly data and real-time assembly errors, thereby achieving real-time suppression of environmental disturbances and significantly improving assembly accuracy and stability.

[0070] A second aspect of the present invention provides an automatic assembly positioning and attitude adjustment system for modular equipment, comprising: The first unit is used to obtain the assembly reference position information of the modular equipment to be assembled, and to use a depth camera to perform multi-angle scanning of the modular equipment to be assembled to obtain the real-time pose parameters of the modular equipment. The second unit is used to construct an assembly trajectory prediction model based on the Monte Carlo tree search algorithm. The real-time pose parameters and the assembly reference position information are input into the assembly trajectory prediction model to generate multiple candidate assembly trajectories. The collision risk of each candidate assembly trajectory is evaluated, and the assembly trajectory with the lowest collision risk is selected as the optimal assembly trajectory. The third unit is used to calculate the position deviation and attitude deviation values ​​between the target position and the current position of the modular equipment based on the optimal assembly trajectory. The fourth unit is used to dynamically compensate the position deviation value and the attitude deviation value using an adaptive predictive encoder. The adaptive predictive encoder optimizes the compensation parameters online based on historical assembly data to achieve real-time suppression of environmental disturbances. The fifth unit is used to control the robotic arm to adjust the position and attitude of the modular equipment to be assembled according to the compensated position deviation value and attitude deviation value, until the modular equipment reaches the assembly reference position.

[0071] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0072] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0073] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for automatic assembly, positioning, and attitude adjustment of modular equipment, characterized in that, include: The assembly reference position information of the modular equipment to be assembled is obtained, and the modular equipment to be assembled is scanned from multiple angles using a depth camera to obtain the real-time pose parameters of the modular equipment. An assembly trajectory prediction model is constructed based on the Monte Carlo tree search algorithm. The real-time pose parameters and the assembly reference position information are input into the assembly trajectory prediction model to generate multiple candidate assembly trajectories. The collision risk of each candidate assembly trajectory is evaluated, and the assembly trajectory with the lowest collision risk is selected as the optimal assembly trajectory. Calculate the positional deviation and attitude deviation between the target position and the current position of the modular equipment based on the optimal assembly trajectory; An adaptive predictive encoder is used to dynamically compensate for the position deviation and the attitude deviation. The adaptive predictive encoder optimizes the compensation parameters online based on historical assembly data to achieve real-time suppression of environmental disturbances. The control robot arm adjusts the position and attitude of the modular equipment to be assembled based on the compensated position and attitude deviation values ​​until the modular equipment reaches the assembly reference position.

2. The method according to claim 1, characterized in that, The modular device to be assembled is scanned from multiple angles using a depth camera to obtain its real-time pose parameters, including: Multiple depth cameras are arranged to form a ring scanning array. The scanning angle of each depth camera is adaptively adjusted based on the structural feature information of the modular equipment to be assembled, so as to ensure that the scanning field of view of the depth camera completely covers all surfaces of the modular equipment to be assembled. Depth images acquired by each depth camera are collected, and the depth images are preprocessed to obtain feature point cloud data of the modular device to be assembled. The feature point cloud data is input into a pre-trained feature matching network. The feature matching network uses epipolar geometric constraints to match feature points from multiple perspectives, generating six-degree-of-freedom pose parameters for the modular device to be assembled, and outputting the real-time pose parameters of the modular device to be assembled.

3. The method according to claim 2, characterized in that, The feature point cloud data is input into a pre-trained feature matching network, which uses epipolar geometry constraints to match feature points from multiple perspectives, generating six-degree-of-freedom pose parameters for the modular device to be assembled, including: The feature point cloud data is aligned with the coordinate system based on the preset pose reference data, and the spatial position information and descriptor information of the feature points are extracted. The matching cost between the spatial location information and the descriptor information is calculated based on epipolar geometric constraints. A reprojection error function for feature point pairs is constructed, and the reprojection error is optimized to obtain the initial matching result of the feature points. The feature point pairs in the initial matching results are filtered, and matching point pairs that do not meet the requirements are removed according to the epipolar geometric constraints; The selected feature point pairs are iteratively optimized. In each iteration, the reprojection error function of the feature point pairs is reconstructed. The reprojection error is minimized to obtain the six-degree-of-freedom pose parameters of the modular device to be assembled.

4. The method according to claim 1, characterized in that, An assembly trajectory prediction model is constructed based on the Monte Carlo tree search algorithm. The real-time pose parameters and the assembly reference position information are input into the assembly trajectory prediction model to generate multiple candidate assembly trajectories, including: An assembly trajectory state space is constructed based on the Monte Carlo tree search algorithm. The assembly feasibility is evaluated for each motion state according to the assembly trajectory state space, and a search tree structure for the assembly trajectory is generated. The number of visits and reward values ​​of assembly trajectory nodes are determined based on the search tree structure. An iterative search is performed starting from the root node of the search tree structure. In each iteration, the node with the highest reward value is selected for expansion until the search depth reaches a preset search threshold. Each search path is then output as a candidate assembly trajectory.

5. The method according to claim 4, characterized in that, An assembly trajectory state space is constructed based on the Monte Carlo tree search algorithm. Assembly feasibility is evaluated for each motion state according to this state space, generating a search tree structure for the assembly trajectory, including: Using the assembly start position as the root node, an assembly trajectory state space is constructed. Based on the Monte Carlo tree search algorithm, search nodes are generated in the assembly trajectory state space. According to the assembly state of the parent node, the Monte Carlo sampling method is used to generate the candidate assembly state of the child node, and an initial number of visits and a reward value are assigned to each search node. Obtain the geometric constraint information of the equipment to be assembled and its surrounding environment. Based on the geometric constraint information, perform an assembly feasibility assessment on each search node, calculate the collision risk value and the assembly difficulty value, and use the weighted sum of the collision risk value and the assembly difficulty value as the node score. Based on the node score, prune and optimize the search tree to generate the final assembly trajectory search tree structure.

6. The method according to claim 1, characterized in that, An adaptive predictive encoder is used to dynamically compensate for the position deviation and attitude deviation values. The adaptive predictive encoder optimizes the compensation parameters online based on historical assembly data to achieve real-time suppression of environmental disturbances, including: An adaptive predictive encoder is constructed. The sequence of compensation parameters from the historical assembly process and the actual assembly error sequence are input into the adaptive predictive encoder to generate initial compensation parameters for compensating the position deviation value and the attitude deviation value. The initial compensation parameters are applied to the current assembly process to obtain the compensated real-time assembly error; The initial compensation parameters are optimized and adjusted online based on the real-time assembly error until the real-time assembly error is less than a preset error threshold, thereby obtaining the final compensation parameters and achieving real-time suppression of environmental disturbances.

7. A modular equipment automatic assembly positioning and attitude adjustment system, used to implement the method of any one of claims 1-6, characterized in that, include: The first unit is used to obtain the assembly reference position information of the modular equipment to be assembled, and to use a depth camera to perform multi-angle scanning of the modular equipment to be assembled to obtain the real-time pose parameters of the modular equipment. The second unit is used to construct an assembly trajectory prediction model based on the Monte Carlo tree search algorithm. The real-time pose parameters and the assembly reference position information are input into the assembly trajectory prediction model to generate multiple candidate assembly trajectories. The collision risk of each candidate assembly trajectory is evaluated, and the assembly trajectory with the lowest collision risk is selected as the optimal assembly trajectory. The third unit is used to calculate the position deviation and attitude deviation values ​​between the target position and the current position of the modular equipment based on the optimal assembly trajectory. The fourth unit is used to dynamically compensate the position deviation value and the attitude deviation value using an adaptive predictive encoder. The adaptive predictive encoder optimizes the compensation parameters online based on historical assembly data to achieve real-time suppression of environmental disturbances. The fifth unit is used to control the robotic arm to adjust the position and attitude of the modular equipment to be assembled according to the compensated position deviation value and attitude deviation value, until the modular equipment reaches the assembly reference position.

8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.