Screw correction method and system based on computer visual perception

By using multimodal image acquisition and fusion technology, combined with physical model-driven dynamic compensation, efficient and precise deviation correction in the screw assembly process was achieved, solving environmental interference and adaptability issues, and improving the accuracy and efficiency of automated assembly.

CN121921312APending Publication Date: 2026-04-24JIANGSU YEDI ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU YEDI ELECTRONICS CO LTD
Filing Date
2026-03-25
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies for automated screw assembly suffer from limitations such as the effectiveness of correction being affected by environmental interference, single-modal image acquisition leading to information loss, poor adaptability, and difficulty in simultaneously meeting the requirements for assembly efficiency and accuracy.

Method used

Multimodal image acquisition, including structured light, depth and polarization images, is employed to generate a multimodal fused image. The screw pose is located through a feature layer, the real-time correction amount is calculated, and environmental interference is eliminated through dynamic compensation driven by a physical model, thus realizing closed-loop correction execution.

Benefits of technology

It improves the accuracy and efficiency of screw assembly, reduces manual intervention, adapts to the assembly needs of different batches of workpieces, and enhances the stability and adaptability of the correction system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921312A_ABST
    Figure CN121921312A_ABST
Patent Text Reader

Abstract

The invention discloses a screw deviation rectifying method and system based on computer visual perception, and belongs to the technical field of industrial robot intelligent assembly.The screw deviation rectifying method comprises the steps that a collecting device conducts multi-modal collection on image data in the screw assembling process, and images collected in a multi-modal mode comprise a structured light image, a depth image and a polarization image; preprocessing the collected image data, generating a multi-modal fusion image, and generating a first feature layer and a second feature layer; positioning the spatial pose of the screw relative to the contact surface of the illuminating lamp; calculating a real-time deviation correction amount; and closed-loop correction execution is carried out through the execution mechanism, the correction effect is monitored in real time, and a correction control signal is fed back. In the implementation process of the technical scheme, the collected image is preprocessed to generate the multi-modal fusion image, and the correction amount of the screw is calculated according to the multi-modal fusion image in the subsequent analysis process, so that correction is performed on the screw assembly process, and the assembly effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent assembly technology for industrial robots, specifically a screw correction method and system based on computer vision perception. Background Technology

[0002] With the improvement of industrial automation, the precision problems in the traditional screw assembly process have become increasingly prominent. Manual assembly is inefficient and prone to errors. Therefore, more and more companies are starting to use automated equipment to achieve automated screw assembly.

[0003] In automated assembly processes, computer vision systems are typically used to position and correct screws. The position of the screw holes on the workpiece is captured by a camera or other sensors, and then an execution signal is sent to the execution device through a preset program. During this process, the accuracy of the execution signal is judged. If there is a deviation, an algorithm is used to make real-time adjustments to ensure the precision and stability of screw assembly.

[0004] However, during the implementation of the above technical solutions, the correction effect is often affected by environmental interference and equipment errors. In order to balance assembly efficiency and accuracy, single-modal image acquisition is usually adopted to reduce the analysis time of the correction process. However, this method will result in the loss of some information and cannot truly reflect the offset information of the screw. At the same time, single-modal images are difficult to cope with complex environmental changes. When different batches of workpieces need to be assembled, the parameters need to be reset, resulting in poor adaptability and making it difficult to meet the production requirements at the same time for assembly efficiency and correction accuracy.

[0005] Therefore, it is necessary to provide a screw correction method and system based on computer vision perception to solve the above problems.

[0006] It should be noted that the information disclosed in this background section is only for understanding the background technology of this application concept, and therefore may include information that does not constitute prior art. Summary of the Invention

[0007] Based on the aforementioned problems in the existing technology, the problem to be solved by this application is to provide a screw correction method and system based on computer vision perception, which can accurately calculate the screw correction amount by performing multimodal analysis on images during the assembly process, thereby improving assembly efficiency.

[0008] The technical solution adopted by this application to solve its technical problem is: a screw correction method based on computer vision perception, comprising: The acquisition device performs multimodal acquisition of image data during the screw assembly process. The multimodal acquired images include structured light images, depth images, and polarization images. The acquired image data is preprocessed to generate a multimodal fused image and generate a first feature layer and a second feature layer. Based on the first feature layer and the second feature layer, and combining the three-dimensional coordinate information of the structured light image with the absolute depth value of the depth image, the spatial pose of the screw relative to the contact surface of the lighting lamp is located. Based on the screw pose and the target pose of the assembly process, the real-time correction amount is calculated, and the influence of environmental interference on the correction command is eliminated through dynamic compensation driven by the physical model. The actuator performs closed-loop correction, and adaptive adjustments are made during the execution process. The correction effect is monitored in real time, and feedback correction control signals are provided.

[0009] In the implementation of the technical solution of this application, the acquired images are preprocessed to generate multimodal fusion images, and the screw correction amount is calculated based on the multimodal fusion images in the subsequent analysis process, thereby correcting the screw assembly process and improving the assembly effect.

[0010] Furthermore, the preprocessing of the acquired image data includes the following steps: Edge gradient features are extracted from structured light images, and a first feature layer is constructed based on the edge gradient features. This first feature layer is used to describe the material and shape information of the contact surface between the screw and the lighting lamp. Texture gradient features are extracted from the depth image, and a second feature layer is constructed based on the texture gradient features. This second feature layer is used to describe the uneven texture of non-metallic materials. Establish a synergistic relationship between polarization and depth, and decouple the texture gradient based on polarization and depth to obtain an anisotropic texture gradient map; The second feature layer is adaptively updated through an online learning mechanism, and the texture gradient features are fused with the original curvature information of the depth image.

[0011] Furthermore, the extraction of edge gradient features from structured light images includes: based on a known lighting model, introducing random perturbation parameters to generate a virtual perturbation lighting model, then simulating structured light projection through a physical rendering engine and adding additional noise, labeling the edge tags of the simulated image based on the geometric parameters of the virtual model to generate edge samples, and then inputting the generated edge samples into an edge detection network to extract edge gradient features from the structured light image.

[0012] Furthermore, extracting texture gradient features from depth images includes: First, depth denoising and smoothing are performed. Bilateral filtering is used to filter the original depth map, preserving edge information while suppressing Gaussian noise. The kernel size of the bilateral filter is adjusted according to the curvature of the lighting surface. Then, multi-scale curvature constraints are used to enhance the texture gradient of the depth image and generate a multi-scale texture gradient map.

[0013] Furthermore, multi-scale curvature constraints are achieved by constructing a multi-level scale pyramid, which includes at least three sub-graphs, each corresponding to a texture gradient at a different scale. After constructing the multi-level scale pyramid, the local curvature of each sub-graph is calculated, and a dynamic threshold is set according to the prior curvature distribution of the non-metallic material. Then, cross-scale fusion is performed, the gradient of each scale sub-graph after screening is calculated, and weights are assigned according to the curvature magnitude. The final texture gradient map is obtained by fusion.

[0014] Furthermore, generating anisotropic texture gradient maps includes the following steps: First, the polarization angle gradient and polarization degree gradient are extracted from the polarization image. The polarization angle gradient reflects the change in polarization direction of the reflected light from the surface, and the polarization degree gradient reflects the anisotropy of the surface roughness. Then, the correlation between depth and polarization is analyzed to construct the cross-correlation matrix between the depth gradient and the polarization gradient. Next, the anisotropic gradient is enhanced based on the correlation analysis results. The calculation direction of the depth gradient is adjusted, and the gradient value is enhanced. Finally, an anisotropic texture gradient map is generated.

[0015] Furthermore, the cross-correlation matrix contains depth gradient information at different polarization angles and offsets in different directions. If the offset in a specific direction is higher than in other directions, it indicates that there is anisotropic bump texture in that direction.

[0016] Furthermore, positioning the screw's spatial orientation relative to the light fixture's contact surface includes: The first feature layer, the second feature layer, and the anisotropic texture gradient map are aligned with the 3D point cloud of the structured light image and the absolute depth value of the depth image to generate multimodal feature vectors. Then, a cross-modal graph is constructed using the pixel coordinates of the structured light image as nodes and the Euclidean distance of the multimodal features as edge weights. The cross-modal association graph is then input into the multilayer perceptron decoder to output the pose parameters of the screw.

[0017] Furthermore, the calculation of the real-time correction amount includes: first, calculating the pose deviation, which includes translational deviations in various directions and rotational deviations of the screw relative to the target pose; then, weighting and sorting the translational and rotational deviations based on assembly requirements, prioritizing the compensation of deviations that have a greater impact on assembly accuracy; then, compensating for environmental interference through a dynamic compensation model, introducing an environmental interference observer to analyze vibration acceleration and temperature change rate, then calculating the compensation amount, and finally outputting the correction command compensation value.

[0018] A screw alignment system based on computer vision perception includes: The data acquisition module is used to acquire multimodal image data of the screw assembly process by the acquisition device. The multimodal acquired images include structured light images, depth images and polarization images. The acquired image data is preprocessed to generate a multimodal fused image and generate a first feature layer and a second feature layer. The positioning module is used to locate the spatial pose of the screw relative to the contact surface of the lighting lamp based on the first feature layer, the second feature layer, the three-dimensional coordinate information of the structured light image and the absolute depth value of the depth image. The correction calculation module is used to calculate the real-time correction amount based on the screw pose and the target pose of the assembly process, and eliminate the influence of environmental interference on the correction command through dynamic compensation driven by the physical model. The closed-loop correction execution module is used to perform closed-loop correction execution through the actuator, and to make adaptive adjustments during the execution process, monitor the correction effect in real time, and provide feedback correction control signals.

[0019] The beneficial effects of this application are as follows: This application provides a screw correction method and system based on computer vision perception, which generates a multimodal fusion image by preprocessing the acquired image, and calculates the screw correction amount based on the multimodal fusion image in the subsequent analysis process, thereby correcting the screw assembly process and improving the assembly effect.

[0020] In addition to the purposes, features, and advantages described above, this application has other purposes, features, and advantages. A further detailed description of this application will be provided below with reference to the figures. Attached Figure Description

[0021] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the overall process of a screw correction method based on computer vision perception in this application; Figure 2This is a schematic diagram of the module structure of a screw correction system based on computer vision perception in this application. Detailed Implementation

[0022] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0024] Example 1: As Figure 1 As shown, this application provides a screw correction method based on computer vision perception. This method is applied in the industrial assembly of motorcycle lights, specifically in the automated screw assembly of industrial robots. During the assembly process, computer vision perception or industrial vision processing is used to monitor and correct the screw assembly in real time, ensuring accurate screw positioning, improving assembly efficiency and precision, and reducing manual intervention. The method includes the following steps: Step 101: The acquisition device performs multimodal acquisition of image data during the screw assembly process. The multimodal acquired images include structured light images, depth images, and polarization images. The acquired image data is preprocessed to generate a multimodal fused image and generate a first feature layer and a second feature layer. In the automated assembly process of screws, it is necessary to continuously collect image data from the assembly line. In the assembly process of motorcycle lights, due to the complexity of the light structure and the influence of different materials on the image processing, traditional single-dimensional acquisition methods are difficult to apply. Therefore, multimodal acquisition is carried out to collect image data during the screw assembly process. Multimodal acquisition includes structured light acquisition, with the corresponding acquisition device being a structured light camera, which is used to obtain the three-dimensional information of the target; depth acquisition, with the corresponding device being a ToF sensor, which is used to obtain the depth information of the target; and polarization acquisition, with the corresponding device being an infrared polarization camera, which can detect some things that cannot be detected by conventional imaging, such as birefringence information, stress information, surface roughness, etc., thereby providing a comprehensive analysis of the light structure image. Due to the complex structure of motorcycle headlights, their outlines are typically curved, and the same visual surface may contain multiple areas of different materials, increasing the difficulty of image processing. Therefore, after acquiring the image data, preprocessing is required. The preprocessing process mainly involves fusing data from different modalities, enabling accurate identification during automated assembly without pre-inputting specific headlight parameters. Specifically, the preprocessing of the acquired image data further includes the following steps: Step A: Extract edge gradient features from the structured light image, and construct a first feature layer based on the edge gradient features. This first feature layer is used to describe the material and shape information of the contact surface between the screw and the lighting lamp. Structured light images are images containing three-dimensional information acquired through a structured light camera. Because lighting fixtures often contain metallic materials, these materials create distinct shadow-highlight edge gradient features in the structured light image, different from other materials. Extracting these features allows for accurate differentiation of the different materials and shapes of the contact surfaces between screws and lighting fixtures, providing foundational data for subsequent edge correction. Furthermore, to facilitate the assembly process of other batches of lighting fixtures, the extracted edge gradient features are constructed as a first feature layer, which can be directly used in subsequent assembly processes, improving versatility. Specific methods for extracting edge gradient features from structured light images include edge detection using the Sobel operator and optimization of edge continuity using the Canny algorithm, thereby extracting the edge gradient features. Additionally, since the curved surface structure of the lighting fixture may cause edge blurring, in this embodiment, the method for extracting edge gradient features from structured light images may further include the following steps: Based on the known lighting model, random perturbation parameters are introduced to generate a virtual perturbation lighting model. Then, the structured light projection is simulated through a physical rendering engine, and additional noise is added. The edge labels of the simulated image are labeled based on the geometric parameters of the virtual model to generate edge samples. The generated edge samples are then input into the edge detection network to extract edge gradient features from the structured light image. Traditional edge extraction methods require multiple iterations of optimization using various operators. This approach is not only time-consuming but also poorly adaptable to complex curved surfaces. Therefore, in this embodiment, the aforementioned method can be used. Based on a known lighting CAD model, random perturbation parameters are introduced, including surface curvature offset, material reflectivity fluctuations, and surface roughness changes, to generate a perturbed lighting model. Then, a physical rendering engine is used to simulate structured light projection. For example, engines such as Blender and Unity are used to simulate the distribution of structured light on the lighting surface, generating a structured light image containing 3D information. Additional noise, such as motion blur and uneven lighting, is added. After generating the virtual model, it will contain several... Geometric parameters, such as abrupt changes in surface normal vectors and material interface intersections, can be used to label the edges of the simulated image using a binary mask. Compared with existing technologies, this significantly reduces labeling costs. After labeling, edge samples are generated and input into an edge detection network. In this embodiment, the edge detection network can be a HED network, which is a highly efficient edge detection algorithm. Through multi-scale feature fusion, it captures detailed changes in the image and, through a virtual-to-real data loop, enables the model to learn richer edge features, improving robustness and accuracy in practical applications. Even in complex environments, it can accurately identify edges, reduce false detections and missed detections, and improve assembly efficiency. Step B: Extract texture gradient features from the depth image and construct a second feature layer based on the texture gradient features. This second feature layer is used to describe the uneven texture of non-metallic materials. In the curved surface structure of lighting fixtures, texture gradient features are particularly crucial, especially for non-metallic materials such as plastics and fibers inside the fixture. The uneven texture of these materials directly affects light scattering, influencing image acquisition and subsequent analysis and processing. Therefore, in this embodiment, texture gradient features are extracted from the depth image to numerically reflect the surface undulations of the material. A second feature layer is constructed based on these features, which can then be applied to more materials in subsequent processes. In existing technologies, extracting texture gradient features from depth images typically involves calculating grayscale gradients; however, this method is easily affected by noise, and the uneven texture of non-metallic surfaces can cause slight fluctuations in local depth values. Therefore, in this embodiment, the following method is used to extract texture gradient features from the depth image: First, depth denoising and smoothing are performed. Bilateral filtering is used to filter the original depth map, which preserves edge information while suppressing Gaussian noise. The kernel size of the bilateral filter is adjusted according to the curvature of the lighting surface. Specifically, a small kernel is used for areas with large curvature, and a large kernel is used for areas with small curvature. Then, the depth image is enhanced with texture gradient through multi-scale curvature constraints, and a multi-scale texture gradient map is generated. The bump texture of non-metallic materials is usually numerically represented as local undulations with low amplitude and high frequency, such as the injection molding texture of plastic parts and the weaving texture of fiber layers. In the prior art, a single-scale gradient calculation method is usually used, which is easy to lose some details. Therefore, in this embodiment, multi-scale curvature constraints are used to combine gradient information of different scales to dynamically capture bump features of different sizes. Specifically, multi-scale curvature constraints are achieved by constructing a multi-level scale pyramid, which includes at least three sub-images, each corresponding to the texture gradient at a different scale. The three sub-images in the multi-level scale pyramid correspond to the texture gradient features at coarse, medium, and small scales, respectively. The curvature range corresponding to the coarse-scale sub-image is 5 to 10 mm, the curvature range corresponding to the medium-scale sub-image is 2 to 5 mm, and the curvature range corresponding to the small-scale sub-image is 0.5 to 2 mm. After constructing the multi-level scale pyramid, adaptive curvature threshold filtering is required. The local curvature of each sub-image is calculated, and a dynamic threshold is set according to the prior curvature distribution of non-metallic materials. For example, the typical curvature range of plastic parts is 0.1 to 1.5 mm, and that of fiber layers is 0.05 to 0.3 mm. Through dynamic threshold filtering, significant texture features are retained, noise interference is eliminated, and only areas with curvature greater than the threshold are retained, while pseudo-curvature caused by noise or smooth areas is filtered out. Cross-scale fusion is performed by calculating the gradients of the selected sub-images at each scale and assigning weights according to the curvature, thus obtaining the final texture gradient map. After the selection is completed, cross-scale fusion is also required because texture features at different scales are complementary. After fusion, the surface details of the material can be more comprehensively reflected, improving the application effect of the feature layer. During fusion, the gradients of the selected sub-images at each scale are calculated. This process can be implemented using the Sobel operator, and weights are assigned according to the curvature. The weight values ​​are dynamically adjusted according to the curvature changes. For example, regions with sharp curvature changes are given higher weights, and regions with gentle curvature changes are given lower weights, ensuring that the fused texture gradient map can highlight significant features while retaining details, thus obtaining the final fused multi-scale texture gradient map. Step C: Establish the synergistic relationship between polarization and depth, and decouple the texture gradient based on polarization and depth to obtain an anisotropic texture gradient map; Within the internal materials of lighting fixtures, anisotropic textures can exist on non-metallic surfaces due to processing techniques. For example, the flow direction during injection molding can cause striped unevenness. The depth gradient of these textures varies in different polarization directions, leading to biases when analyzed independently. Therefore, a synergistic relationship between polarization and depth is established, and joint decoupling is performed based on polarization and depth information to eliminate anisotropic interference, thereby generating an anisotropic texture gradient map. Specifically, this process includes the following steps: First, the polarization angle gradient and polarization degree gradient are extracted from the polarization image. This method is obtained by calculating the Stokes vector. The polarization angle gradient reflects the change in polarization direction of the reflected light from the surface, and the polarization degree gradient reflects the anisotropy of the surface roughness. The Stokes vector consists of four parameters that describe the polarization state of light. For specific application schemes, please refer to the existing technology. In this embodiment, it will not be described in detail. Then, a correlation analysis is performed on depth and polarization to construct a cross-correlation matrix between depth gradient and polarization gradient. When performing the correlation analysis between depth and polarization, it is necessary to consider the variation law of depth gradient under different polarization states. Therefore, this is achieved through a cross-correlation matrix. The cross-correlation matrix contains depth gradient information under different polarization angles and offsets in different directions. If the offset in a specific direction is higher than in other directions, it indicates that there is anisotropic texture in that direction. Through this analysis, the distribution and intensity of anisotropic texture can be accurately identified and quantified. Next, based on the correlation analysis results, the anisotropic gradient is enhanced by adjusting the calculation direction of the depth gradient and enhancing the gradient value, ultimately generating an anisotropic texture gradient map. After constructing the cross-correlation matrix between the depth gradient and the polarization gradient, the calculation direction of the depth gradient can be adjusted according to the maximum correlation direction obtained from the matrix, and the gradient value can be enhanced by using the correlation coefficient method to ensure that the gradient changes of the anisotropic texture under different polarization states are effectively enhanced, thereby obtaining an anisotropic texture gradient map and further optimizing the image fusion effect. Step D: Adaptively update the second feature layer through an online learning mechanism, and fuse each texture gradient feature with the original curvature information of the depth image; After acquiring each texture gradient feature, it is necessary to fuse it with the original curvature information of the depth image as input to the second feature layer. An online learning mechanism is then used to adaptively update the second feature layer, making the feature representation more accurate and dynamically adapting to texture changes in different scenes. This process can also adapt to texture changes of different batches of non-metallic materials. The process includes normalizing each texture gradient feature and concatenating it with the depth image to generate an initial feature layer. Then, a lightweight neural network is introduced, using the label of the current batch of non-metallic materials as supervision, adjusting the channel weights, and outputting the final second feature layer. This method does not require pre-storing material parameters and can quickly adapt to the texture features of new materials with a small number of samples, thereby achieving efficient feature extraction and updating. Step 102: Based on the first feature layer and the second feature layer, and combining the three-dimensional coordinate information of the structured light image with the absolute depth value of the depth image, the spatial pose of the screw relative to the contact surface of the lighting lamp is located. The first feature layer is the edge gradient feature, and the second feature layer is the texture gradient feature. By combining the three-dimensional coordinate information of the structured light image with the absolute depth value of the depth image, the spatial pose of the screw relative to the contact surface of the lighting lamp can be accurately located. This avoids the feature ambiguity and scale distortion problems in traditional single-modal positioning, improves positioning accuracy and stability, and ensures that the spatial pose information of the screw can be accurately captured under different lighting and material conditions, providing data support for subsequent automated assembly. Specifically, the method includes: aligning the first feature layer, the second feature layer, and the anisotropic texture gradient map with the 3D point cloud of the structured light image and the absolute depth value of the depth image to generate multimodal feature vectors; then constructing a cross-modal graph by using the pixel coordinates of the structured light image as nodes and the Euclidean distance of the multimodal features as edge weights. The cross-modal graph is a graph structure that effectively integrates multi-source information. By iteratively updating the node coordinates through graph optimization algorithms, it can accurately fit the spatial pose of the screw and eliminate errors between multi-source data. The graph is then input into a multilayer perceptron decoder to output the screw's pose parameters, which include 3D coordinates and rotation angles around each axis, thereby describing the screw's spatial pose. The multilayer perceptron decoder is a deep learning-based model that can effectively fuse multimodal features and accurately decode the screw's spatial pose information through layer-by-layer nonlinear transformation, ensuring high-precision positioning in different environments. Step 103: Calculate the real-time correction amount based on the screw pose and the target pose of the assembly process, and eliminate the influence of environmental interference on the correction command through dynamic compensation driven by the physical model. After obtaining the screw pose, it is necessary to compare it with the target pose and calculate the real-time correction amount. Then, a control signal is generated based on the real-time correction amount to drive the actuator to make adjustments. Since there will be environmental interference during the correction process, such as vibration and temperature drift, a dynamic compensation mechanism driven by the physical model is also needed to correct these interference factors in real time. Specifically, the above process includes: first, calculating the pose deviation, which includes translational deviations in various directions and rotational deviations of the screw relative to the target pose; then, weighting and sorting the translational and rotational deviations based on assembly requirements, prioritizing compensation for deviations that have a greater impact on assembly accuracy. For example, during screw assembly, if the translational deviation is large, the X and Y axis positions are adjusted first, and then the Z axis height and rotation around the axis are corrected, so that the screw gradually approaches the target pose and reduces assembly errors. Then, environmental disturbances are compensated by a dynamic compensation model. An environmental disturbance observer is introduced to analyze vibration acceleration and temperature change rate. The compensation amount is then calculated, and finally, the correction command compensation value is output. An environmental interference observer is an intelligent sensor that can monitor environmental changes in real time. It identifies and quantifies interference factors through data fusion technology. In this embodiment, the interference factors are vibration and temperature drift. Therefore, vibration acceleration and temperature change rate are used as the analysis sources. The compensation amount can be calculated using a state-space model. By comparing and analyzing historical data and real-time data, the compensation strategy is dynamically adjusted to obtain the optimal compensation value and ensure the accurate execution of the correction command. Step 104: Perform closed-loop correction through the actuator, make adaptive adjustments during the execution process, monitor the correction effect in real time, and provide feedback correction control signals.

[0025] During the correction process, an actuator is required. In this embodiment, the actuator can be an industrial robotic arm or a servo motor. After the execution is completed, adaptive adjustment is required, assembly image data is re-acquired, and a new round of pose recognition and correction is performed to form closed-loop control. The correction process can refer to existing technologies, such as converting correction commands into joint control quantities of the robotic arm and generating joint motion trajectories to complete screw position planning. Adaptive adjustment involves inputting the image data of this correction process into an edge detection network and a graph neural network for deep learning analysis to identify the trend of pose deviation changes. Based on historical correction data, a reinforcement learning function is constructed to achieve long-term performance optimization. The independent variables of this reinforcement learning function include adjustment accuracy, adjustment speed, and energy consumption. Each independent variable has a weight coefficient. By optimizing the combination of weight coefficients, the correction strategy is dynamically adjusted to improve assembly efficiency and accuracy, ultimately achieving adaptive closed-loop control.

[0026] Example 2: Figure 2 As shown, this application also proposes a screw correction system based on computer vision perception. This system operates the correction method in Embodiment 1, and includes: The data acquisition module is used to acquire multimodal image data of the screw assembly process by the acquisition device. The multimodal acquired images include structured light images, depth images and polarization images. The acquired image data is preprocessed to generate a multimodal fused image and generate a first feature layer and a second feature layer. The positioning module is used to locate the spatial pose of the screw relative to the contact surface of the lighting lamp based on the first feature layer, the second feature layer, the three-dimensional coordinate information of the structured light image and the absolute depth value of the depth image. The correction calculation module is used to calculate the real-time correction amount based on the screw pose and the target pose of the assembly process, and eliminate the influence of environmental interference on the correction command through dynamic compensation driven by the physical model. The closed-loop correction execution module is used to perform closed-loop correction execution through the actuator, and to make adaptive adjustments during the execution process, monitor the correction effect in real time, and provide feedback correction control signals.

[0027] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A screw alignment method based on computer vision perception, characterized in that: include: The acquisition device performs multimodal acquisition of image data during the screw assembly process. The multimodal acquired images include structured light images, depth images, and polarization images. The acquired image data is preprocessed to generate a multimodal fused image and generate a first feature layer and a second feature layer. Based on the first feature layer and the second feature layer, and combining the three-dimensional coordinate information of the structured light image with the absolute depth value of the depth image, the spatial pose of the screw relative to the contact surface of the lighting lamp is located. Based on the screw pose and the target pose of the assembly process, the real-time correction amount is calculated, and the influence of environmental interference on the correction command is eliminated through dynamic compensation driven by the physical model. The actuator performs closed-loop correction, and adaptive adjustments are made during the execution process. The correction effect is monitored in real time, and feedback correction control signals are provided.

2. The screw alignment method based on computer vision perception according to claim 1, characterized in that: Preprocessing the acquired image data further includes the following steps: Edge gradient features are extracted from the structured light image, and a first feature layer is constructed based on the edge gradient features. This first feature layer is used to describe the material and shape information of the contact surface between the screw and the lighting lamp. Texture gradient features are extracted from the depth image, and a second feature layer is constructed based on the texture gradient features. This second feature layer is used to describe the uneven texture of non-metallic materials. Establish a synergistic relationship between polarization and depth, and decouple the texture gradient based on polarization and depth to obtain an anisotropic texture gradient map; The second feature layer is adaptively updated through an online learning mechanism, and the texture gradient features are fused with the original curvature information of the depth image.

3. The screw correction method based on computer vision perception according to claim 2, characterized in that: Extracting edge gradient features from structured light images involves: based on a known lighting model, introducing random perturbation parameters to generate a virtual perturbation lighting model; then simulating structured light projection using a physical rendering engine and adding additional noise; labeling the edge of the simulated image based on the geometric parameters of the virtual model to generate edge samples; and then inputting the generated edge samples into an edge detection network to extract edge gradient features from the structured light image.

4. The screw correction method based on computer vision perception according to claim 2, characterized in that: Extracting texture gradient features from depth images includes: First, depth denoising and smoothing are performed. Bilateral filtering is used to filter the original depth map, preserving edge information while suppressing Gaussian noise. The kernel size of the bilateral filter is adjusted according to the curvature of the lighting surface. Then, multi-scale curvature constraints are used to enhance the texture gradient of the depth image and generate a multi-scale texture gradient map.

5. The screw alignment method based on computer vision perception according to claim 3, characterized in that: Multi-scale curvature constraints are achieved by constructing a multi-level scale pyramid, which includes at least three sub-graphs. Each sub-graph corresponds to the texture gradient at a different scale. After constructing the multi-level scale pyramid, the local curvature of each sub-graph is calculated, and a dynamic threshold is set according to the prior curvature distribution of the non-metallic material. Then, cross-scale fusion is performed. The gradient of each scale sub-graph after screening is calculated, and weights are assigned according to the curvature magnitude. The final texture gradient map is obtained by fusion.

6. The screw alignment method based on computer vision perception according to claim 2, characterized in that: Generating anisotropic texture gradient maps includes the following steps: First, the polarization angle gradient and polarization degree gradient are extracted from the polarization image. The polarization angle gradient reflects the change in polarization direction of the reflected light from the surface, and the polarization degree gradient reflects the anisotropy of the surface roughness. Then, the correlation between depth and polarization is analyzed to construct the cross-correlation matrix between the depth gradient and the polarization gradient. Next, the anisotropic gradient is enhanced based on the correlation analysis results. The calculation direction of the depth gradient is adjusted, and the gradient value is enhanced. Finally, an anisotropic texture gradient map is generated.

7. The screw correction method based on computer vision perception according to claim 6, characterized in that: The cross-correlation matrix contains depth gradient information at different polarization angles and offsets in different directions. If the offset in a specific direction is higher than in other directions, it indicates that there is anisotropic texture in that direction.

8. The screw alignment method based on computer vision perception according to claim 2, characterized in that: Positioning the screw relative to the contact surface of the light fixture includes: The first feature layer, the second feature layer, and the anisotropic texture gradient map are aligned with the 3D point cloud of the structured light image and the absolute depth value of the depth image to generate multimodal feature vectors. Then, a cross-modal graph is constructed using the pixel coordinates of the structured light image as nodes and the Euclidean distance of the multimodal features as edge weights. The cross-modal association graph is then input into the multilayer perceptron decoder to output the pose parameters of the screw.

9. The screw alignment method based on computer vision perception according to claim 1, characterized in that: The calculation of real-time correction amount includes: first, calculating the pose deviation, which includes translational deviation in various directions and rotational deviation of the screw relative to the target pose; then, weighting and sorting the translational and rotational deviations based on assembly requirements, prioritizing the compensation of deviations that have a greater impact on assembly accuracy; then, compensating for environmental interference through a dynamic compensation model, introducing an environmental interference observer to analyze vibration acceleration and temperature change rate, then calculating the compensation amount, and finally outputting the correction command compensation value.

10. A screw alignment system based on computer vision perception, used to implement the screw alignment method based on computer vision perception as described in any one of claims 1 to 9, characterized in that: include: The data acquisition module is used to acquire multimodal image data of the screw assembly process by the acquisition device. The multimodal acquired images include structured light images, depth images and polarization images. The acquired image data is preprocessed to generate a multimodal fused image and generate a first feature layer and a second feature layer. The positioning module is used to locate the spatial pose of the screw relative to the contact surface of the lighting lamp based on the first feature layer, the second feature layer, the three-dimensional coordinate information of the structured light image and the absolute depth value of the depth image. The correction calculation module is used to calculate the real-time correction amount based on the screw pose and the target pose of the assembly process, and eliminate the influence of environmental interference on the correction command through dynamic compensation driven by the physical model. The closed-loop correction execution module is used to perform closed-loop correction execution through the actuator, and to make adaptive adjustments during the execution process, monitor the correction effect in real time, and provide feedback correction control signals.