Vision and laser collaborative assembly positioning method and system
By employing a vision-laser co-application assembly positioning method, combined with operator status, high-precision and high-stability assembly process control was achieved. This solved the problems of inaccurate positioning and operational errors in complex industrial scenarios, improving assembly efficiency and human-machine collaboration.
Patent Information
- Application Number
- CN202511510653.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2026-01-16
AI Technical Summary
Existing machine vision-based assembly guidance technologies are inaccurate in positioning and lack flexibility in complex industrial scenarios. They cannot make intelligent adjustments based on the operator's skill level, resulting in insufficient accuracy and a high error rate.
By employing a vision-laser synergy approach, dynamic compensation and collaborative positioning are achieved through the simultaneous acquisition of stereoscopic images and laser spot data. This, combined with real-time adjustments to the virtual-real guidance scheme based on the operator's status, generates a high-precision visual guidance scheme.
It achieves high-precision and high-stability assembly process control, reduces the cognitive load and error rate of operators, and improves assembly efficiency and human-machine collaboration effectiveness.
Smart Images

Figure CN121353404A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of industrial automation control, in particular to a visual and laser cooperative assembly positioning method and system. BACKGROUND
[0002] In modern high-end manufacturing industries, such as aerospace, automotive and precision instruments, the assembly process of products is increasingly complex, and the requirements for precision and quality are increasingly stringent. In order to improve the assembly efficiency and reduce human errors, various auxiliary assembly technologies have emerged. Assembly guidance system as an advanced production auxiliary tool, through providing real-time and accurate operation instructions to the operator, plays a key role in guiding operation, preventing mistakes and process tracing. Such systems usually belong to the category of industrial process control and monitoring, aiming to optimize and standardize manual operation links through informationization and automation means.
[0003] The existing assembly guidance technology mostly adopts augmented reality scheme based on machine vision. These schemes usually use a single or multiple cameras to shoot the assembly site, track the position of workpieces or tools through image recognition algorithms, and then superimpose virtual guidance information such as arrows, contour lines or three-dimensional models on real-time video streams on display devices or special AR glasses to guide operators. Some systems also use projection to project guidance graphics directly onto the workpiece surface, realizing a form of augmented reality guidance.
[0004] However, the above-mentioned existing technology still has some inherent limitations in practical application. On the one hand, the positioning method relying solely on vision has high requirements for the working environment. When the surface of the parts lacks sufficient texture, there is strong light reflection, or the on-site lighting conditions are unstable, the robustness and accuracy of the vision algorithm will decrease significantly, which may lead to positioning failure or shaking and deviation between virtual information and real objects, affecting the guidance effect. On the other hand, guidance through virtual information on the screen still requires the operator to make a visual mapping from two-dimensional or virtual three-dimensional space to real three-dimensional physical space, which may have perception bias and is difficult to achieve unambiguous indication of millimeter-level or even higher precision points. In addition, the instructions provided by most guidance systems are fixed and cannot be adjusted according to the proficiency or real-time state of different operators, lacking flexibility and intelligence. SUMMARY
[0005] To solve the above problems, the present application provides a visual and laser cooperative assembly positioning method and system, which synchronously acquires stereo images and laser spot data for dynamic compensation and cooperative positioning, and adjusts the virtual-real combined guidance scheme in real time in combination with the state of the operator, so as to realize high-precision, high-stability and adaptive assembly process control.
[0006] The above object can be achieved by the following solution: A visual and laser collaborative assembly positioning method and system, comprising: synchronously acquiring stereo image data of an assembly area and positioning light spot data projected by a laser, to generate multi-source input data; performing image analysis and feature matching based on the multi-source input data, identifying a part type and calculating an initial pose estimation of the part type in a camera coordinate system, to generate a visual positioning result; loading a preset three-dimensional geometric specification corresponding to an assembly task, extracting theoretical position data defined in the three-dimensional geometric specification, and comparing the visual positioning result with the theoretical position data, and correcting deviations by a dynamic compensation algorithm, to generate collaborative positioning data; combining the collaborative positioning data with a preset standard operation procedure file, to generate a visual guidance scheme containing display instructions and laser control instructions; according to the visual guidance scheme, synchronously driving a display device to execute the display instructions and driving a laser to execute the laser control instructions, to realize virtual-real combined assembly guidance; based on the assembly guidance, performing assembly while monitoring the assembly process in real time to generate operation progress data, and acquiring historical operation data associated with an operator's identity, and dynamically adjusting the detail level of the visual guidance scheme in combination with the operation progress data and the historical operation data.
[0007] Optionally, the generating multi-source input data comprises: synchronously collecting left and right views by a binocular camera to form stereo image data; controlling a laser to project a light beam with a specific pattern onto a part surface to form positioning light spot data; and aligning the stereo image data and the positioning light spot data in time and merging them into multi-source input data.
[0008] Optionally, the generating visual positioning result comprises: generating a three-dimensional point cloud based on the stereo image data and extracting surface contour features therefrom; matching the surface contour features with a preset part model library to identify a part type and obtain an initial pose estimation; taking accurate three-dimensional coordinates of the positioning light spot data in the three-dimensional point cloud and optimizing the initial pose estimation by using the accurate three-dimensional coordinates, to generate a visual positioning result.
[0009] Optionally, the generating collaborative positioning data comprises: calculating deviation values between the visual positioning result and the theoretical position data and recording the deviation values to form a deviation value sequence; analyzing based on the deviation value sequence to distinguish random noise and systematic drift; compensating and filtering based on the random noise and the systematic drift to generate a dynamic compensation transformation, and applying the dynamic compensation transformation to the visual positioning result to output collaborative positioning data.
[0010] Optionally, the generating the visual guidance scheme containing display instructions and laser control instructions comprises: determining a current assembly step to be executed according to the cooperative positioning data; extracting animation instructions and laser projection target points corresponding to the assembly step from the standard operation procedure file; compiling the animation instructions into display instructions and the laser projection target points into laser control instructions, and packaging them together into a visual guidance scheme.
[0011] Optionally, the method further comprises: performing graphic rendering based on the display instructions to present guidance information in the form of highlighting and three-dimensional animation on a display device; and performing coordinate system transformation on the target three-dimensional coordinates in the laser control instructions to drive a laser to project an indicating light spot on a physical workpiece.
[0012] Optionally, the dynamically adjusting the detailed level of the visual guidance scheme based on the operation progress data and the historical operation data comprises: generating a proficiency evaluation result of an operator based on the average assembly time and error frequency in the historical operation data; generating an instant state evaluation result based on operation continuity analysis in the operation progress data; and selecting and applying guidance information of a corresponding detailed level from a preset multi-level guidance content library according to the proficiency evaluation result and the instant state evaluation result.
[0013] Optionally, the method further comprises: loading a preset tolerance range file defining qualified installation parameters, comparing the cooperative positioning data with the tolerance range file to generate a quality determination result; and if the quality determination result is abnormal, triggering an alarm device and pausing the guidance process, and storing the operation progress data and the quality determination result to a database after binding them with a product number.
[0014] Optionally, the selecting and applying guidance information of a corresponding detailed level from a preset multi-level guidance content library comprises: when the proficiency evaluation result is low or the instant state evaluation result is abnormal, applying high detailed level guidance information containing step-by-step guidance animation and detailed prompts; when the proficiency evaluation result is high and the instant state evaluation result is normal, applying simplified guidance information containing only core position indication; and automatically switching between the high detailed level guidance information and the simplified guidance information according to changes in the instant state evaluation result during the operation process.
[0015] Based on the same inventive concept, the application also provides a visual and laser cooperative assembly positioning system, comprising: a multi-source data acquisition module for synchronously acquiring stereoscopic image data of an assembly area and positioning light spot data projected by a laser, and generating multi-source input data; a visual positioning recognition module for image analysis and feature matching based on the multi-source input data, recognizing a part type and calculating an initial pose estimation of the spatial position of the part type in a camera coordinate system, and generating a visual positioning result; a dynamic deviation compensation module for loading a preset three-dimensional geometric specification corresponding to an assembly task, extracting theoretical position data defined in the three-dimensional geometric specification, and comparing the visual positioning result with the theoretical position data, correcting deviations through a dynamic compensation algorithm, and generating cooperative positioning data; a visual scheme generation module for generating a visual guidance scheme containing display instructions and laser control instructions in combination with the cooperative positioning data and a preset standard operation procedure file; a virtual-real guidance execution module for driving a display device to execute the display instructions and driving a laser to execute the laser control instructions according to the visual guidance scheme, realizing virtual-real combined assembly guidance; and a dynamic guidance adjustment module for generating operation progress data while monitoring the assembly process in real time based on the assembly guidance, obtaining historical operation data associated with the identity of an operator, and dynamically adjusting the detail level of the visual guidance scheme in combination with the operation progress data and the historical operation data.
[0016] Compared with the prior art, the application has the following advantages: 1. The application effectively solves the problems of inaccurate positioning and drift of traditional visual positioning in complex industrial scenes such as weak texture and reflection through the cooperative work of visual and laser sensors and the combination of a dynamic compensation algorithm. The system not only uses stereoscopic vision to obtain a global contour for preliminary identification, but also uses laser spots as high-precision anchor points for pose optimization, and continuously corrects deviations through time sequence analysis, thereby ensuring that the positioning result has high accuracy and long-term stability, providing a reliable data basis for subsequent accurate guidance and quality control.
[0017] 2. The virtual-real combined assembly guidance method provided by the application improves the intuitiveness of guidance information and the accuracy of operation. By rendering an animation on a display device to explain the operation process and action essentials, and driving a laser to directly project an indication light spot of a key point on a physical workpiece, the macro process guidance and micro position guidance are perfectly combined. This dual-mode guidance information complements each other, reduces the cognitive load of the operator and the misoperation caused by visual errors, and improves the execution efficiency of complex assembly tasks.
[0018] 3、The application establishes a set of adaptive guidance mechanism based on the state of the operator, realizes the intelligentization of man-machine cooperation. The system analyzes the historical data and real-time operation progress of the operator, evaluates its proficiency and instant state, and dynamically adjusts the detail level of the guidance information accordingly. This personalized and adaptive guidance strategy can provide detailed guidance for beginners and concise prompts for experts, avoiding the efficiency bottleneck or error risk caused by traditional fixed guidance for different levels of personnel, making the whole assembly control system more flexible, and improving the overall efficiency of man-machine cooperation.
[0019] Other features and advantages of the present application will be set forth in the following description, and in part will be apparent from the description, or can be learned by practice of the present application. The objects and other advantages of the present application will be realized and attained by the structure particularly pointed out in the written description and claims thereof as well as the appended drawings. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0021] Figure 1 is a flowchart of a visual and laser cooperative assembly positioning method according to an embodiment of the present application.
[0022] Figure 2 is a dynamic deviation compensation effect schematic diagram according to an embodiment of the present application.
[0023] Figure 3 is a virtual-real combined guidance mode schematic diagram according to an embodiment of the present application.
[0024] Figure 4 is a quality determination schematic diagram according to an embodiment of the present application.
[0025] Figure 5 is a structural schematic diagram of a visual and laser cooperative assembly positioning system according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0027] Referring to Figure 1 One embodiment of the present application proposes a visual and laser collaborative assembly positioning method, which synchronously acquires stereo image and laser spot data for dynamic compensation and collaborative positioning, and combines the state of the operator to adjust the virtual-real combined guidance scheme in real time, thereby achieving high-precision and high-stability adaptive assembly process control.
[0028] The method of the embodiment specifically comprises: Synchronously acquiring stereo image data and laser positioning spot data of an assembly area to generate multi-source input data; Performing image analysis and feature matching based on the multi-source input data, identifying the type of a component and calculating the initial pose estimation of the spatial position of the type of the component in the camera coordinate system to generate a visual positioning result; Loading a preset three-dimensional geometric specification corresponding to an assembly task, extracting theoretical position data defined in the three-dimensional geometric specification, and comparing the visual positioning result with the theoretical position data to correct the deviation through a dynamic compensation algorithm to generate collaborative positioning data; Combining the collaborative positioning data and a preset standard operation procedure file to generate a visual guidance scheme containing display instructions and laser control instructions; According to the visual guidance scheme, synchronously driving a display device to execute the display instructions and driving a laser to execute the laser control instructions to achieve virtual-real combined assembly guidance; Based on the assembly guidance, performing assembly while monitoring the assembly process in real time to generate operation progress data, and acquiring historical operation data associated with the identity of an operator, and dynamically adjusting the detail level of the visual guidance scheme in combination with the operation progress data and the historical operation data.
[0029] Specifically, first, the environment is perceived in a coordinated manner of vision and laser, through synchronous acquisition of stereoscopic image data of the assembly area and active positioning spot data of laser projection, forming multi-source input data with complementary information. Based on the multi-source input data, preliminary identification and positioning are carried out, and visual positioning results are generated. Subsequently, the real-time visual positioning results are continuously compared with the theoretical position data in the preset three-dimensional geometric specification, and a dynamic compensation algorithm is used to identify and correct systematic drift and random noise, thereby generating collaborative positioning data. On the basis of the collaborative positioning data, combined with the preset standard operation procedure file, a visual guidance scheme containing display instructions and laser control instructions is decided and generated. At the execution level, the display device and the laser are synchronously driven to realize assembly guidance combining virtual animation and physical spots. Finally, through a dynamic feedback mechanism, a closed loop is formed, which monitors the operation progress data in real time and combines the historical operation data of the operator to evaluate the personnel proficiency and the instant state, and then dynamically adjusts the detail level of the guidance scheme, realizing personalized and adaptive intelligent guidance. The coordinated perception mode of vision and laser reduces the limitations of a single sensor in a complex industrial environment, and the application of the dynamic compensation algorithm effectively reduces the positioning drift in long-term work, together ensuring that the pose data on which the guidance information depends has high accuracy and stability. Secondly, in the guidance mode, by synchronously combining virtual three-dimensional animation and physical laser spots, multi-modal complementation of information presentation is realized, virtual guidance provides a macroscopic understanding of the operation process and action, and physical spots provide unambiguous accurate guidance of key points, reducing the cognitive load and operation error rate of the operator, and improving the assembly efficiency and one-time success rate. Finally, by introducing a dynamic guidance adjustment mechanism based on operator proficiency and real-time state, detailed guidance can be provided for beginners to speed up learning and ensure quality, and concise prompts can be provided for skilled workers to avoid interference, thereby optimizing the human-machine collaborative experience.
[0030] Optionally, the generating multi-source input data comprises: synchronously collecting left and right views through a binocular camera to form stereoscopic image data; controlling the laser to project a light beam with a specific pattern onto the surface of the part to form positioning spot data; aligning the stereoscopic image data and the positioning spot data in time and merging them into multi-source input data.
[0031] Specifically, first, a set of acquisition system including hardware and control logic is deployed, the core of which includes a strictly calibrated binocular camera and a controllable laser. The binocular camera is composed of a left camera and a right camera separated by a fixed baseline distance in space. The internal and external parameters of the camera, i.e. the focal length, principal point, distortion coefficient of each camera, and the relative rotation and translation relationship between the two cameras, have been accurately determined in advance through a standard calibration process. During data acquisition, through a synchronous trigger signal or a synchronous timestamp mechanism, the left camera and the right camera are ensured to be exposed at the same time, synchronously collecting images of the assembly area to form left and right views respectively. These two two-dimensional images with parallax together constitute the stereo image data at that moment, providing a basis for subsequent three-dimensional reconstruction. At the same time, the laser projects a light beam with a specific pattern, such as a high-brightness crosshair, a circular spot, or a specific coded dot matrix, onto the surface of the part to be positioned. This specific pattern of light beam will form a clear and high-contrast reflection spot on the surface of the part, and the image information of this spot is the positioning spot data. At the same time, the collection of stereo image data and the formation of positioning spot data must be accurately aligned in time. This is achieved by coupling the camera's collection trigger with the laser's projection control logic, ensuring that at the moment when the binocular camera captures the left and right views, the light spot projected by the laser also stably appears on the surface of the part. Finally, the stereo image data collected at the same time is combined with the positioning spot data contained in the two views to form a data unit. This data unit is stamped with a uniform timestamp to form a complete multi-source input data, providing synchronized and complementary information for subsequent identification, positioning and analysis. Through the above method, the multi-source input data generated realizes the effective fusion of two different sensing information. Synchronizing and merging the two in time makes the global passive visual information and the local active laser measurement information complementary, improving the robustness of subsequent pose calculation and the accuracy of final positioning result.
[0032] Optionally, the generating the visual positioning result comprises: generating a three-dimensional point cloud based on the stereo image data, and extracting surface contour features from the three-dimensional point cloud; matching the surface contour features with a preset part model library to identify part types and obtain an initial pose estimate; extracting accurate three-dimensional coordinates of the positioning spot data in the three-dimensional point cloud, and optimizing the initial pose estimate using the accurate three-dimensional coordinates to generate a visual positioning result.
[0033] Specifically, first, the received stereo image data is rectified and densely matched using the internal and external parameters of the binocular camera, and a disparity map is calculated through algorithms such as semi-global block matching (SGBM). Then, according to the principle of triangulation, the disparity map is converted into a dense three-dimensional point cloud, which constitutes a three-dimensional digital description of the surfaces of all objects in the current assembly area. Next, from the generated three-dimensional point cloud, the geometric properties of the point cloud such as normal vector and curvature are analyzed by algorithms to extract surface contour features that can stably describe the shape of the parts, such as edge lines, corner points, plane patches, or more advanced feature descriptions such as fast point feature histogram (FPFH). In the next step, the real-time extracted surface contour features are matched with a pre-set part model library. The part model library stores digital models of various standard parts and their corresponding surface contour features. Through robust matching algorithms such as random sample consensus (RANSAC), a search and comparison is performed in the feature space to identify the types of parts present in the scene and calculate an initial pose estimate from the model coordinate system to the camera coordinate system, which is usually represented by a rotation matrix and a translation vector. However, this initial pose estimate may have errors caused by point cloud noise, occlusion, or insignificant features. To refine the positioning, the light spot data is processed. Through image processing algorithms, the center point coordinates of the laser spots are located in the left and right views with sub-pixel accuracy, and the two two-dimensional coordinate points are triangulated using the calibrated camera parameters to calculate an accurate three-dimensional coordinate of the spot in the camera coordinate system. Finally, the accurate three-dimensional coordinate is used to optimize the initial pose estimate. The goal of optimization is to adjust the initial pose estimate so that the theoretical position of the laser projection point on the part model, after pose transformation, minimizes the Euclidean distance with the accurate three-dimensional coordinate of the laser spot. This optimization problem can be expressed as minimizing an objective function, for the objective function , , where, and are the rotation matrix and translation vector to be optimized, is the theoretical three-dimensional coordinate of the laser projection point in the part model coordinate system, which is predefined, is the accurate three-dimensional coordinate of the laser spot measured by binocular vision; represents the square of the modulus length. The objective function The smaller the value is, i.e. the higher the positioning accuracy is. This problem can be solved by an iterative optimization algorithm such as a nonlinear least squares method to obtain an optimized and high-precision transformation matrix, which is the final output of the visual positioning result. This method combines the ability of passive visual global perception and the precision advantage of active local positioning, and improves the accuracy and reliability of the final positioning result, especially in challenging industrial scenes such as weak texture, shiny surface or complex background, its performance is better than the positioning method relying on visual features alone.
[0034] Optionally, the generating the cooperative positioning data comprises: calculating a deviation value between the visual positioning result and the theoretical position data, and recording a deviation value sequence; analyzing based on the deviation value sequence, distinguishing random noise and systematic drift; compensating and filtering based on the random noise and the systematic drift, generating a dynamic compensation transformation, and applying the dynamic compensation transformation to the visual positioning result to output cooperative positioning data.
[0035] Specifically, first, the visual positioning result representing the measured pose of the part is obtained, and the theoretical position data of the part in the assembly task is loaded from the pre-set three-dimensional geometric specification; the three-dimensional geometric specification obtains the accurate three-dimensional model of the part through a high-precision three-dimensional scanner or CAD design software, and defines the theoretical position, attitude such as rotation and translation matrix of each part in the assembly body in combination with the coordinate system calibration of the actual assembly environment. The visual positioning result and the theoretical position data are both represented in the form of a 4x4 homogeneous transformation matrix, which contains rotation and translation information. The deviation is quantified by calculating the difference between the two matrices, specifically, at each time step, the deviation transformation matrix is calculated, which represents the transformation from the theoretical pose to the measured pose. For the deviation transformation matrix at the time step , wherein, represents the visual positioning result, which is the measured pose, a 4x4 homogeneous transformation matrix; represents the theoretical pose, which is the pre-set specification pose, a 4x4 homogeneous transformation matrix; represents matrix multiplication; represents the inverse matrix of the matrix A sequence of bias values is thus formed. Next, a time domain analysis is performed on the sequence of bias values to distinguish between two different types of errors. One is high frequency, random noise with a mean close to zero, mainly caused by image sensor noise and algorithmic instability. The other is low frequency, systematic drift with a persistent bias, possibly caused by device thermal expansion, mechanical vibration, or slow changes in ambient lighting. To achieve this distinction, a filtering algorithm such as a Kalman filter can be employed. The Kalman filter takes the true state of the bias, mainly reflecting the systematic drift, as a state variable, and the computed bias transformation matrix of each frame as an observation. By predicting the drift at the next time step using a state transition equation and correcting it using the current observation, the Kalman filter can effectively estimate the true value of the systematic drift from the observation sequence, which is full of random noise. This estimated systematic drift is the dynamic compensation transformation. Finally, to correct the visual positioning result, the inverse of the dynamic compensation transformation is applied to the current visual positioning result, thus generating the final co-location data. For calculating the final co-location data , , wherein, represents the visual positioning result obtained at time step , which is a 4x4 homogeneous transformation matrix; is the dynamic compensation transformation representing the current systematic drift output from an algorithm such as a Kalman filter, which is also a 4x4 homogeneous transformation matrix; is the inverse matrix of the dynamic compensation transformation. is the final output obtained after the compensation operation, i.e., the high-precision co-location data, which represents the best pose estimate of the part after drift correction. As shown in Figure 2 , the gray dashed line represents the "sequence of bias values between the visual positioning result and the theoretical position" without processing, which contains high-frequency random noise and slow-changing systematic drift, with overall fluctuations being large; the black dashed line represents the "systematic drift" analyzed and estimated from the original bias by an algorithm such as a Kalman filter, which reflects low-frequency, persistent errors caused by factors such as device thermal expansion and contraction; the black solid line shows the "co-location data bias after compensation", and it can be seen that after the systematic drift is subtracted, the bias value is stabilized around zero. It can resist slow positioning bias caused by environmental or device changes in real time, ensuring the consistency and reliability of positioning under long-time operation; the final output co-location data thus has higher time domain smoothness and approximation to the theoretical position, providing a solid data foundation for subsequent high-precision assembly guidance and quality control.
[0036] Optionally, the generating a visual guidance scheme including display instructions and laser control instructions comprises: determine a current assembly step to be executed according to the collaborative positioning data; extract animation instructions and laser projection target points corresponding to the assembly step from the standard operation procedure file; compile the animation instructions into display instructions and the laser projection target points into laser control instructions, and jointly encapsulate them into a visual guidance scheme.
[0037] Specifically, first, the collaborative positioning data generated in the previous stage is received and a current assembly step to be executed is determined. The collaborative positioning data describes the real-time pose of the parts to be assembled in the camera coordinate system. At the same time, a preset standard operation procedure file matching the current assembly task is loaded. The standard operation procedure file is a structured data document that divides the entire assembly process into a series of ordered steps and defines clear trigger conditions, virtual guidance content, and physical indication information for each step. First, the execution state is judged, and the real-time collaborative positioning data is compared with the target pose of each step defined in the standard operation procedure file. When the actual pose of a part meets the determination condition for the completion of a step, the process state is automatically advanced to the next assembly step to be executed. After determining the current assembly step to be executed, two types of core information strongly related to the step are retrieved and extracted from the standard operation procedure file. The first type is animation instructions, which describe how to guide the operation through three-dimensional graphics in an abstract language independent of specific rendering engines, such as specifying the path of a virtual part model moving from the current position to the target position, highlighting a specific mounting surface, or displaying relevant text prompts. The second type is laser projection target points, which are a set of three-dimensional coordinates defined in the part model coordinate system, accurately indicating key physical locations such as screw hole positions, alignment marker points, or glue application areas. Subsequently, the compilation stage is entered, and the extracted animation instructions are converted into a set of specific display instructions. These display instructions are commands that can be directly parsed and executed by the graphics rendering engine, including model loading, transformation matrix setting, material property modification, and animation keyframe information. At the same time, the laser projection target points are transformed between coordinate systems, from the model coordinate system to the global coordinate system where the laser is located, and finally compiled into laser control instructions that can directly drive the deflection of the laser scanning galvanometer. Finally, the generated display instruction set and laser control instruction set are encapsulated into a synchronized data package, the visual guidance scheme, and passed to the execution module of the next stage. Through real-time perception of collaborative positioning data to drive the automatic flow of the assembly process, the timeliness and accuracy of the guidance information are improved, and the delay and errors caused by traditional manual switching steps are reduced.
[0038] Optionally, the method further comprises: performing graphics rendering based on the display instructions to present guidance information in the form of highlighting and three-dimensional animation on a display device; The target three-dimensional coordinates in the laser control instruction are subjected to coordinate system transformation, and the laser is driven to project an indicating light spot on the physical workpiece.
[0039] Specifically, when receiving the visualization guidance scheme containing display instructions and laser control instructions, the two types of instructions are processed in parallel. On the one hand, the graphic rendering engine receives the display instructions, first sets the camera pose of the virtual scene to be consistent with the real-time pose of the physical binocular camera according to the established coordinate relationship between the camera and the display world, then parses the display instructions, loads the three-dimensional models specified in the display instructions, and uses the above-mentioned collaborative positioning data to update the positions and poses of these three-dimensional models in the virtual scene in real time, so as to realize visual alignment of the virtual models and the physical workpiece on the display device. For highlight instructions, the graphic rendering engine will modify the shader parameters of the corresponding virtual model surface, such as increasing the self-luminous intensity or changing its color, so that it stands out prominently in the picture. For animation instructions, a series of continuous transformation matrices are generated by a keyframe interpolation algorithm according to the defined start and end states, to drive the virtual model to smoothly demonstrate the assembly path or operation action. Finally, the virtual scene superimposed with the guidance information is synthesized with the real-time video stream, and the result is output to the display device to present an information-enhanced reality picture to the operator. On the other hand, the laser control instruction, the core content of which is the target three-dimensional coordinates, is initially defined in the model coordinate system of the parts. In order to drive the laser to accurately project the light spot on the physical workpiece, a series of coordinate system transformations must be performed. The target three-dimensional coordinates of the target point in the laser coordinate system , , wherein, is the homogeneous representation of the target three-dimensional coordinates in the model coordinate system extracted from the standard operation procedure file; is the real-time collaborative positioning data, representing the pose transformation matrix of the physical workpiece from the model coordinate system to the camera coordinate system; is a fixed transformation matrix representing the transformation relationship from the camera coordinate system to the laser coordinate system, which is obtained by offline calibration after device installation. The laser then converts this three-dimensional coordinate into the deflection angle instruction required by its internal scanning galvanometer system, drives the galvanometer to deflect the laser beam quickly and accurately, and finally projects a clear and high-brightness indicating light spot on the specified position of the physical workpiece. For example, Figure 3As shown, "virtual world" represents the content in the display device screen; a dark gray model is rendered in the screen in real-time image alignment with the "physical workpiece", and a light gray "animation guide" is superimposed thereon; "physical world" represents the real operating scene, and a laser beam emitted by the laser is shown in the figure, which projects a clear "laser indication spot" on the exact position of the "physical workpiece". For calculating the deflection angle of horizontal deflection and the deflection angle of vertical deflection , there are: , wherein, represents the target point coordinates in the laser coordinate system; represents the inverse tangent function. This method constructs a dual-mode augmented reality guide environment by synchronously executing virtual graphic rendering and physical laser projection. This virtual-real combined and mutually complementary guide mode reduces the cognitive load and misoperation risk of the operator, and improves the execution efficiency and one-time success rate of complex assembly tasks.
[0040] Optionally, the step of dynamically adjusting the detailed level of the visual guide scheme by combining the operation progress data and the historical operation data comprises: generating a proficiency evaluation result of the operator based on the average assembly time and error frequency in the historical operation data; generating an instant state evaluation result based on operation continuity analysis in the operation progress data; selecting and applying guide information of a corresponding detailed level from a pre-set multi-level guide content library according to the proficiency evaluation result and the instant state evaluation result.
[0041] Specifically, first, the historical operation data associated with the operator is retrieved from the database through the operator's identity authentication information. The historical operation data is a long-term cumulative performance record, which contains the key performance indicators of the operator when performing the same or similar assembly tasks in the past, mainly including the average assembly time and error frequency. Based on these data, a comprehensive proficiency evaluation result is calculated through a preset evaluation model, which quantifies the skill level of the operator into a level such as novice, ordinary or expert; the evaluation model is trained based on historical operation data such as average assembly time, error frequency and real-time operation continuity indicators such as pause frequency and trajectory jitter through logistic regression, decision tree or lightweight neural network, and the key hyperparameters include learning rate, tree depth or hidden layer node number, and the training data comes from the labeled behavior data of multiple operators at different proficiency levels, and is continuously optimized through cross-validation and actual assembly feedback to dynamically evaluate the proficiency and real-time state of the operator. At the same time, during the entire assembly process, operation progress data is monitored and recorded in real time, which not only includes the start and end timestamps of each step, but more importantly, through the analysis of the collaborative positioning data stream, the micro-dynamics of the operation are captured, such as the smoothness of the component movement trajectory, whether there is unnecessary back-and-forth or long-term pause, so as to perform operation continuity analysis. Based on the analysis, an immediate state evaluation result is generated to determine whether the current state of the operator is "normal and smooth", "hesitant" or "abnormal". Finally, the dynamic guidance adjustment module combines the long-term proficiency evaluation result with the short-term immediate state evaluation result, and selects and applies the most appropriate guidance information from a multi-level guidance content library according to a preset decision matrix. The multi-level guidance content library has different levels of guidance versions for each assembly step, such as high detail, standard detail and simplified version. According to the decision result, the selected version of the guidance content is compiled into a new visual guidance scheme, which is pushed and executed. Through the introduction of a double evaluation feedback loop, the assembly guidance has changed from "one size fits all" to "personalization". It not only sets an initial guidance strategy according to the historical proficiency of the operator before the task starts, but also adjusts it in real time according to the real-time performance during the operation, forming an intelligent and adaptive teaching system.
[0042] Optionally, the method further comprises: loading a preset tolerance range file defining qualified installation parameters, comparing the collaborative positioning data with the tolerance range file to generate a quality determination result; if the quality determination result is abnormal, triggering an alarm device and pausing the guidance process, and storing the operation progress data and the quality determination result to the database after binding the product number.
[0043] Specifically, first, a preset tolerance range file for the current station is loaded from the local or server, which defines the geometric parameters required to meet the qualified installation in a digital form, including the ideal values of the six degrees of freedom of the target pose, i.e. three translation components and three rotation components, and their respective acceptable positive and negative deviation ranges. The actual pose represented by the collected collaborative positioning data is mathematically compared with the ideal pose defined in the tolerance range file. This comparison is achieved by calculating the transformation difference between the actual pose and the ideal pose, and decomposing it into translation errors along three coordinate axes and rotation errors around three coordinate axes. Then, these calculated error values are compared with the threshold values in the tolerance range file one by one. If any error value exceeds its corresponding allowed range, an "abnormal" quality judgment result is generated, otherwise a "qualified" quality judgment result is generated. If the quality judgment result is abnormal, the linked alarm device is immediately triggered, such as lighting the warning light, emitting the beep sound or popping up the prominent error prompt on the display screen, and the instruction to pause the guidance process is automatically executed to prevent the operation from continuing until the operator corrects the error. If the quality judgment result is qualified, the operation continues as usual. As shown in FIG. 8, the assembly deviation distribution of a batch of products on a certain key dimension is shown; the ±0.2mm in the figure is the qualified tolerance range, i.e. the shaded rectangular part in the figure; the gray bar column is judged as qualified, and the deviations of most products fall within the shaded rectangle, so they are automatically judged as qualified; the black bar column is judged as abnormal, i.e. if the absolute value of the deviation value exceeds 0.2mm, it is judged as abnormal; the black vertical dotted line represents no deviation. Finally, whether the quality judgment result is qualified or abnormal, the complete operation progress data collected this time is bound with the quality judgment result and associated with the unique number of the product being assembled, and stored in the production traceability database as a complete quality record entry. This method realizes the upgrade from "guided assembly" to "integrated assembly of quality inspection" by embedding real-time, non-contact precision measurement and quality judgment links in the assembly process. It can immediately detect and prevent unqualified products from flowing into the next process, and effectively reduce the rework rate and scrap rate by immediately alarming and pausing the process when an abnormality is detected and immediately correcting it. Figure 4
[0044] Optionally, the selecting and applying of the guidance information with corresponding detailed level from the preset multi-level guidance content library comprises: when the proficiency evaluation result is low or the real-time state evaluation result is abnormal, applying high detailed level guidance information containing step-by-step guidance animation and detailed prompts; when the proficiency evaluation result is high and the real-time state evaluation result is normal, applying simplified guidance information containing only core position indication; automatically switching between the high detailed level guidance information and the simplified guidance information according to the change of the real-time state evaluation result during the operation process.
[0045] Specifically, first, when it is determined that the guidance level needs to be switched, the current evaluation state is checked. When the operator's proficiency evaluation result is rated as low level, such as "novice", or regardless of its proficiency, its real-time state evaluation result is identified as abnormal, such as "long pause" or "error attempt", the high-level guidance information is automatically retrieved from the multi-level guidance content library and applied. The content of this level of guidance information is the most abundant, usually containing step-by-step guidance animations that break down a complex operation into multiple sub-steps, each step clearly demonstrating the correct movement trajectory and alignment posture of the parts. At the same time, detailed text or icon prompts are superimposed on the screen, clearly indicating the key operation points, precautions or tools needed. If the operator's proficiency evaluation result is "normal", and the real-time state evaluation result is normal, some basic step guidance information will be reduced; when the operator's proficiency evaluation result is rated as high level, such as "expert", and its real-time state evaluation result remains "normal flow", simplified guidance information is selected and applied. This guidance aims to minimize the disturbance to skilled operators, and its content is simplified to the core elements. Usually, it omits step-by-step animations and detailed text instructions, and only highlights the relevant part installation area with static highlights, and uses a laser to directly project a core position indicator spot on the physical workpiece, such as the exact center point of a screw hole. This mode provides the necessary position confirmation for the operator, but leaves the pace and specific method of operation entirely to his own experience. In addition, the method also has the ability to dynamically switch during operation. For example, a skilled operator initially starts work with simplified guidance information mode, but at a certain step, hesitation or continuous small positioning deviation is detected through real-time state evaluation, and the real-time state evaluation result changes immediately. At this time, the guidance mode is automatically switched from simplified to high-level detail seamlessly, actively providing more comprehensive assistance. Once the operator successfully completes this step and the subsequent operation resumes smoothly, the real-time state evaluation result returns to normal, and the simplified guidance information is automatically switched back, thus realizing a closed loop of intelligent assistance that dynamically changes with the operator's real-time state. This intelligent switching between different levels of guidance information not only optimizes the operator's work experience and efficiency, but also flexibly adapts to a mixed team of different skill levels, thus improving the overall flexibility of the production line and the comprehensive performance of human-machine collaboration.
[0046] Based on the same inventive concept, as Figure 5 The present application also provides a visual and laser collaborative assembly positioning system, which comprises: A multi-source data acquisition module for synchronously acquiring stereoscopic image data and laser positioning spot data of the assembly area, and generating multi-source input data; a visual positioning recognition module configured to perform image analysis and feature matching based on the multi-source input data, to identify a part type and calculate an initial pose estimation of a spatial position of the part type in a camera coordinate system, and to generate a visual positioning result; a dynamic deviation compensation module configured to load a preset three-dimensional geometric specification corresponding to the assembly task, to extract theoretical position data defined in the three-dimensional geometric specification, to compare the visual positioning result with the theoretical position data, to correct deviations by a dynamic compensation algorithm, and to generate collaborative positioning data; a visual scheme generation module configured to generate a visual guidance scheme including display instructions and laser control instructions in combination with the collaborative positioning data and a preset standard operation procedure file; a virtual-real guidance execution module configured to drive a display device to execute the display instructions and drive a laser to execute the laser control instructions according to the visual guidance scheme, to realize virtual-real combined assembly guidance; a dynamic guidance adjustment module configured to monitor an assembly process in real time to generate operation progress data based on the assembly guidance, to obtain historical operation data associated with an operator identity, and to dynamically adjust a detail level of the visual guidance scheme in combination with the operation progress data and the historical operation data.
[0047] To verify the feasibility of the present application in implementation, the present application is applied to a high-precision fuel nozzle assembly line of an enterprise. The assembly task requires accurate installation of multiple fuel nozzles on a combustion chamber shell, and the assembly precision and consistency directly affect the combustion efficiency and safety of the engine. The traditional manual operation mode has problems such as low positioning accuracy, low operation efficiency, difficult quality traceability, and high dependence on the skill proficiency of the operator. The enterprise hopes to use the method and system of the present application to realize accurate guidance of the assembly process, real-time quality monitoring, and personalized assistance.
[0048] In this embodiment, the enterprise deploys a visual and laser collaborative assembly positioning system described in the present application at the assembly station. The system is composed of a multi-source data acquisition module including a binocular camera and a laser deployed above the workpiece, a display device, and a background processing server. The server runs a visual positioning recognition module, a dynamic deviation compensation module, a visual scheme generation module, a virtual-real guidance execution module, and a dynamic guidance adjustment module. The system provides virtual-real combined assembly guidance for the operator by real-time acquisition and analysis, and realizes whole-process quality control and data traceability.
[0049] To verify the effectiveness of the present application, production data for a month is recorded and analyzed, covering the assembly processes of different proficiency operators, a novice "Wang Gong" and an expert "Li Gong", in multiple production shifts.
[0050] When an operator places the fuel nozzle to be assembled into the predetermined positioning area of the combustion chamber shell, the system begins to execute the positioning and guidance process. First, the multi-source data acquisition module synchronously acquires the stereoscopic image data of the assembly area and the positioning light spot data projected on the nozzle surface by the laser, generating multi-source input data. Subsequently, the visual positioning recognition module generates a three-dimensional point cloud based on the stereoscopic image data and extracts surface contour features from it, identifies that the part is an "A-type fuel nozzle" by matching with the preset part model library, and calculates its initial pose estimation in the camera coordinate system. Then, the system extracts the accurate three-dimensional coordinates of the positioning light spot in the three-dimensional point cloud, and optimizes the initial pose estimation using the coordinates to generate a high-precision visual positioning result.
[0051] The dynamic deviation compensation module loads the preset three-dimensional geometric specifications, compares the visual positioning result with the theoretical position data defined in the specifications, calculates the deviation value and forms a deviation value sequence. For example, in the operation at 10:15 on April 16, 2025, the system detected a systematic drift of about 0.05 mm due to slight thermal expansion and contraction of the tooling fixture. Through the dynamic compensation algorithm, the system corrected the drift and generated the final collaborative positioning data, ensuring the long-term stability and absolute accuracy of the positioning result.
[0052] Based on the collaborative positioning data, the visualization scheme generation module determines the current assembly step to be performed as "align and tighten 3 M6 positioning bolts", and extracts the corresponding animation instructions and laser projection target points from the standard operation process file. The virtual-real guidance execution module synchronously drives the display device and the laser: the display device clearly demonstrates the small rotation and advance of the nozzle required to reach the precise alignment position in the form of highlights and three-dimensional animations; at the same time, the laser accurately projects high-intensity circular light spots on the three bolt hole positions of the combustion chamber shell physical workpiece, providing the operator with unambiguous physical position indications.
[0053] The dynamic guidance adjustment function of the present application has significant advantages in the present embodiment. For the novice Wang worker, the average assembly time in his historical operation data is longer and the error frequency is higher, so the system determines that his proficiency evaluation result is "low". Therefore, the system automatically applies high-detailed guidance information, and in addition to the laser spot and alignment animation, the screen also displays the tightening order of each bolt and the recommended torque value step by step. In one operation, the system monitors Wang's pause for more than 5 seconds at the second bolt through operation progress data, and the real-time state evaluation result becomes "abnormal", so the system immediately pops up a prompt on the screen: "Please confirm the use of 12N·m torque wrench", which effectively avoids errors. For the expert Li worker, the proficiency evaluation result is "high", so the system applies simplified guidance information, only highlighting the nozzle model on the display device and projecting the bolt hole position with laser, greatly reducing information interference and ensuring smooth operation.
[0054] After completing the key work station of tightening the bolt, the system automatically collects operation data including images, collaborative positioning data and time stamps, and compares them with the preset tolerance range file, i.e. translation tolerance ±0.1mm and rotation tolerance ±0.2°. In an assembly on April 18, 2025, due to a bolt not being tightened, the nozzle had a 0.3° tilt, which exceeded the tolerance range. The quality judgment result generated by the system is "abnormal", which immediately triggers the red alarm light on the work station and suspends the guidance process, and the screen displays "No. 3 bolt position is abnormal, please check again". At the same time, the abnormal operation data is bound to the product number "BN-20230810-072" and stored in the database, realizing real-time quality interception and complete data traceability.
[0055] It should be noted that the electrical connection between the above-mentioned units does not necessarily mean direct connection, and indirect connection mode can also be used as long as the purpose of the present application is achieved. The above-mentioned is only an exemplary embodiment of the present application, and cannot limit the scope of the present application.
[0056] That is, any equivalent changes and modifications made according to the teachings of the present application are still within the scope of the present application. Other embodiments of the present application will be readily apparent to those skilled in the art upon considering the description and practice of the principles disclosed herein. The present application is intended to cover any variations, uses or adaptive changes to the present application that follow the general principles of the present application and include common knowledge or conventional technical means in the art not disclosed by the present application.
Claims
1. A method of assembly positioning using vision and laser coordination, comprising: The method comprises: Synchronously acquiring stereoscopic image data of the assembly area and positioning light spot data projected by the laser, to generate multi-source input data; Performing image analysis and feature matching based on the multi-source input data, identifying the type of the parts and components and calculating the initial pose estimation of the spatial position of the type of the parts and components in the camera coordinate system, to generate visual positioning results; Loading the preset three-dimensional geometric specification corresponding to the assembly task, extracting the theoretical position data defined in the three-dimensional geometric specification, and comparing the visual positioning results with the theoretical position data, correcting the deviation by a dynamic compensation algorithm, to generate collaborative positioning data; Combining the collaborative positioning data with the preset standard operation procedure file, to generate a visual guidance scheme containing display instructions and laser control instructions; According to the visual guidance scheme, synchronously driving the display device to execute the display instructions and driving the laser to execute the laser control instructions, to realize virtual-real combined assembly guidance; Based on the assembly guidance, monitoring the assembly process in real time to generate operation progress data, and acquiring historical operation data associated with the identity of the operator, combining the operation progress data with the historical operation data to dynamically adjust the detail level of the visual guidance scheme.
2. The method of claim 1, wherein, The generation of multi-source input data comprises: Synchronously collecting left and right views by a binocular camera to form stereoscopic image data; Controlling the laser to project a light beam with a specific pattern onto the surface of the parts and components, to form positioning light spot data; Aligning the stereoscopic image data and the positioning light spot data in time, and merging them into multi-source input data.
3. A method of visual and laser cooperative assembly positioning according to claim 2, wherein, The generation of visual positioning results comprises: Generating a three-dimensional point cloud based on the stereoscopic image data, and extracting surface contour features therefrom; Matching the surface contour features with a preset parts and components model library to identify the type of the parts and components, to obtain an initial pose estimation; Extracting accurate three-dimensional coordinates of the positioning light spot data in the three-dimensional point cloud, and optimizing the initial pose estimation by using the accurate three-dimensional coordinates, to generate visual positioning results.
4. The method of claim 3, wherein, The generation of collaborative positioning data comprises: Calculating the deviation value between the visual positioning results and the theoretical position data, and recording to form a deviation value sequence; Analyzing based on the deviation value sequence, to distinguish random noise and systematic drift; Compensating and filtering based on the random noise and systematic drift, generating a dynamic compensation transformation, and applying the dynamic compensation transformation to the visual positioning results, to output collaborative positioning data.
5. A method of visual and laser cooperative assembly positioning according to claim 4, wherein, The generation of a visual guidance scheme containing display instructions and laser control instructions comprises: Determining the current assembly step to be executed according to the collaborative positioning data; Extracting animation instructions and laser projection target points corresponding to the assembly step from the standard operation procedure file; Compiling the animation instructions into display instructions, and compiling the laser projection target points into laser control instructions, to jointly encapsulate into a visual guidance scheme.
6. A vision and laser co-operative assembly positioning method according to claim 5, wherein, The method further comprises: Performing graphic rendering based on the display instructions, to present guidance information in the form of highlighting and three-dimensional animation on the display device; The target three-dimensional coordinates in the laser control instruction are subjected to coordinate system transformation, and a laser is driven to project an indicating light spot on a physical workpiece.
7. A vision and laser co-operative assembly positioning method according to claim 6, wherein, The operation progress data and the historical operation data are combined to dynamically adjust the detail level of the visual guidance scheme, which comprises: Based on the average assembly time and error frequency in the historical operation data, a proficiency evaluation result of the operator is generated; Based on the operation continuity analysis in the operation progress data, an instant state evaluation result is generated; According to the proficiency evaluation result and the instant state evaluation result, guidance information of a corresponding detail level is selected and applied from a preset multi-level guidance content library.
8. A vision and laser co-operative assembly positioning method according to claim 7, wherein, The method further comprises: A preset tolerance range file defining qualified installation parameters is loaded, the collaborative positioning data is compared with the tolerance range file, and a quality judgment result is generated; If the quality judgment result is abnormal, an alarm device is triggered and the guidance process is paused, and the operation progress data and the quality judgment result are bound to a product number and stored in a database.
9. The method of claim 7, wherein, The guidance information of a corresponding detail level selected and applied from the preset multi-level guidance content library comprises: When the proficiency evaluation result is low or the instant state evaluation result is abnormal, high-detail-level guidance information containing step-by-step guidance animation and detailed prompts is applied; When the proficiency evaluation result is high and the instant state evaluation result is normal, simplified guidance information containing only core position indication is applied; According to the change of the instant state evaluation result in the operation process, automatic switching is performed between the high-detail-level guidance information and the simplified guidance information.
10. A vision and laser collaborative assembly positioning system for use in a vision and laser collaborative assembly positioning method according to any one of claims 1-9, characterized by, The system comprises: A multi-source data acquisition module is configured to synchronously acquire stereoscopic image data of an assembly area and positioning light spot data projected by a laser, and generate multi-source input data; A visual positioning identification module is configured to perform image analysis and feature matching based on the multi-source input data, identify a part type, calculate an initial pose estimation of a spatial position of the part type in a camera coordinate system, and generate a visual positioning result; A dynamic deviation compensation module is configured to load a preset three-dimensional geometric specification corresponding to an assembly task, extract theoretical position data defined in the three-dimensional geometric specification, compare the visual positioning result with the theoretical position data, correct deviations through a dynamic compensation algorithm, and generate collaborative positioning data; A visual scheme generation module is configured to combine the collaborative positioning data and a preset standard operation process file, and generate a visual guidance scheme containing display instructions and laser control instructions; A virtual-real guidance execution module is configured to drive a display device to execute the display instructions and drive a laser to execute the laser control instructions according to the visual guidance scheme, so as to realize virtual-real combined assembly guidance; A dynamic guidance adjustment module is configured to generate operation progress data by monitoring an assembly process in real time based on the assembly guidance, obtain historical operation data associated with an operator's identity, and dynamically adjust the detail level of the visual guidance scheme by combining the operation progress data and the historical operation data.
Citation Information
Cited By
Part positioning and inclination angle detection method and system based on optical signal processing
CN121720450A