Tracking and positioning method for AR assembly guidance of precision electronic product

Through the combination of perspective classification and multi-branch optical flow network, the occlusion and truncation problems of the estimation of the central position estimation of complex electronic equipment assembly are solved, and the fast and accurate tracking and positioning of target objects is achieved, and the assembly efficiency and accuracy are improved.

CN120339386APending Publication Date: 2025-07-18CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510326456.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

During the assembly process of complex electronic equipment, the 6-degree of freedom posture is easily affected by occlusion and cutoff, resulting in low assembly efficiency, large human work dependence and long assembly cycle.

Method used

The pose initialization method based on view angle classification is adopted, and the data set is generated using the YOLOv8 object detection framework and Open3D, and the pose iterative optimization is combined with the multi-branch optical flow network to improve the robustness and speed of the algorithm.

Benefits of technology

It realizes fast and robust tracking and positioning of target objects in complex scenarios, improves assembly efficiency and accuracy, reduces manual intervention, and shortens assembly cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339386A_ABST
    Figure CN120339386A_ABST
Patent Text Reader

Abstract

The invention discloses a tracking and positioning method for AR assembly guidance of a precision electronic product. A 6D pose estimation task is decomposed into pose initialization and pose iterative optimization. The method comprises the following steps: firstly, converting pose initialization into a visual angle classification task, and using Open3D to integrate an electronic case three-dimensional model and VOC2012 data into a data set for generating visual angle classification; then classifying the visual angles based on a target detection framework YOLOv8, and obtaining a label visual angle close to the current visual angle; and converting a label view angle obtained by YOLOv8 classification into corresponding 3D rotation, and calculating 3D translation by combining the internal reference of the camera and the center coordinate of the detection frame. After a rough initial 6D pose of a target object is obtained, a currently observed image and a rendered image of a model under the current pose are spliced as input, and a pose optimizer based on a pre-training simple optical flow network is used for predicting relative pose transformation between the images and the rendered image to solve a more accurate 6D pose. And higher precision can be obtained in multiple iterations, and finally, the electronic equipment can be tracked and positioned by the network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and specifically relates to a tracking and positioning method for AR assembly guidance of precision electronic products. Background Art

[0002] For technologies such as complex electronic equipment, in dense large-scale electromechanical systems, due to their complex structures and high precision requirements, the assembly quality directly affects the performance of the products. Moreover, in the development of complex products, the assembly workload accounts for 20% - 70% of the entire research and development workload, with an average of 45%. Currently, product assembly is usually mainly manual, and the assembly time accounts for 40% - 60% of the entire machine development. The assembly cycle is long, and the assembly efficiency directly affects the product development cost. Assembly guidance enabled by AR is a highly potential new method that can display assembly guidance information for operators in real time, virtually demonstrate the assembly process, guide the assembly personnel in an intuitive manner, and provide real-time feedback and guidance. It is of great significance for reducing the mental burden of assembly personnel and improving the product assembly quality and efficiency. Among them, 6-degree-of-freedom (6DoF) pose estimation is one of the core technologies of the AR assembly assistance system, and it is the key to realizing the real-time and accurate superposition of virtual guidance information in the real assembly environment. Summary of the Invention

[0003] Aiming at the problem that tracking and positioning in the assembly process of complex electronic products are easily affected by occlusion and truncation, by means of a mature 2D object detection algorithm, a rough initial pose is predicted based on the RGB image input to quickly restore the pose after the target object tracking is lost. At the same time, by introducing depth prior information with a known preset viewing angle, and then using the three-dimensional model information of the target object to assist in iterative optimization to obtain the accurate 6D pose, combined with a reparameterized multi-branch module to improve the operation speed of the algorithm, the fast and robust tracking and positioning of the target object in a complex operation scenario are realized.

[0004] In view of this, the technical solution adopted by the present invention is: a tracking and positioning method for AR assembly guidance of precision electronic products, including the following steps:

[0005] Step 1: Divide 48 uniform viewpoints according to the 12 vertices of the regular icosahedron as the camera view classification standard;

[0006] Step 2: Use the Open3D library to load the three-dimensional model of the electronic equipment, and randomly rotate and move the camera slightly for rendering under 48 camera views;

[0007] Step 3: Use the VOC2012 dataset as the background, synthesize the background for the rendered image and save it, and generate annotation data for the category and border;

[0008] Step 4: Train using the object detection framework YOLOv8 on the synthetic dataset to obtain a model capable of predicting the rough perspective information of the target object;

[0009] Step 5: Convert the labeled perspective classified by YOLOv8 into the corresponding 3D rotation, and calculate the 3D translation by combining the camera intrinsics and the center coordinates of the detection box, which is used as the initial value for subsequent optimization;

[0010] Step 6: Replace the convolutional module of the simple optical flow network with a multi-branch module to enhance the network's expressive ability and improve the fitting ability of the model during training, and train using the flyingchairs optical flow dataset;

[0011] Step 7: Based on the trained simple optical flow network model, perform pose optimization training using the synthetic dataset of the 3D model of the electronic equipment to obtain a model capable of iteratively optimizing the rough pose;

[0012] Step 8: Use structural reparameterization to merge the branches of the multi-branch module of the simple optical flow network, significantly improving the inference speed;

[0013] Step 9: Combine the trained pose initial network (obtained in Step 4) and the pose optimization network (obtained in Step 8) to predict the 6D pose of the electronic equipment for real-time tracking and positioning.

[0014] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned tracking and positioning method for AR assembly guidance of precision electronic products are implemented.

[0015] The present invention proposes a pose initialization method based on perspective classification, synthesizes and produces a dataset for perspective classification, uses the YOLOv8 object detection framework as a pose initializer and restorer, and calculates the initial pose by combining the perspective classification results to achieve rough positioning of the target in complex working scenarios and recovery after pose loss.

[0016] Convert pose initialization into a perspective classification task, use Open3D to synthesize the 3D model of the electronic chassis and the VOC2012 dataset to generate a dataset for perspective classification, so as to improve the algorithm robustness and reduce the difficulty of obtaining training data. Convert the labeled perspective classified by YOLOv8 into the corresponding 3D rotation, and calculate the 3D translation by combining the camera intrinsics and the center coordinates of the detection box. After the target object tracking is lost, the rough initial pose of the target object can be quickly restored based on this.

[0017] The present invention proposes an improved 6D pose iterative optimization method based on optical flow guidance. Aiming at the problems of poor real-time performance of current mainstream optical flow prediction networks, which are difficult to be applied to actual scenarios, and low accuracy of lightweight optical flow networks, the method proposes to introduce a multi-branch module to replace the original downsampling block of the simple optical flow network to improve the model performance, and uses a pre-trained simple optical flow network as the backbone network for the pose optimization task. An online synthetic dataset for pose optimization is used for training. During the inference process, the network structure reparameterization technology is combined to merge multiple convolutional branches, which can improve the inference speed without loss and achieve real-time target 6D pose optimization, providing technical support for the rapid and accurate assembly of electronic equipment under the assistance of AR guidance. Description of the Drawings

[0018] Figure 1 It is the overall framework diagram of the method of the present invention;

[0019] Figure 2 It is the view classification training sample diagram;

[0020] Figure 3 It is the structural diagram of the pose optimization network based on optical flow;

[0021] Figure 4 It is the pose iterative optimization framework diagram;

[0022] Figure 5 It is the schematic diagram of the multi-branch module;

[0023] Figure 6 It is the schematic diagram of the multi-branch operator fusion. Detailed Embodiment

[0024] The tracking and positioning method for the AR assembly guidance of precision electronic products, as Figure 1 shown, the method includes the following steps:

[0025] Step 1: An icosahedron is a regular polyhedron composed of 20 equilateral triangles, with a total of 30 edges, 20 faces, and 12 vertices. Since the vectors from the center to each vertex of a regular polyhedron are uniformly distributed in 3D space, and the vertices of an icosahedron have a moderate density in space, the present invention uses the vertices of the icosahedron to assist in dividing the labels of classification perspectives. Assuming that the target object is located at the center of the world coordinate system and the distance between the camera optical center and the object center is d, the exact coordinates of the 12 vertices {P i |i ∈ [0, 11]} can be obtained from the calculation formula of the icosahedron geometric information. When the camera is located at a vertex P iWhen the camera is oriented towards the center of the target object, the camera rotates clockwise by 90°, 180°, and 270° respectively around the z-axis of its own coordinate system. Adding the state without rotation, a total of 4 reference viewpoints can be obtained under this vertex. A total of 48 reference viewpoints can be obtained for all vertices, denoted as {View_ref i×4+k | i ∈ [0, 11], n ∈ [0, 3]}, and record the rotation quaternion of the camera relative to the 3D model under each viewpoint, denoted as {q_ref i×4+k | i ∈ [0, 11], n ∈ [0, 3]}, as well as 48 images of the 3D model under the current viewpoints, denoted as {I_ref i×4+k | i ∈ [0, 11], n ∈ [0, 3]}. i represents the number of vertices. Since there are 4 reference viewpoints with different rotation states under each vertex i, n from 0 to 3 represents rotating 0°, 90°, 180°, and 270° clockwise around the z-axis of its own coordinate system respectively. Therefore, it is necessary to combine the rotation state parameter n to use {I_ref i×4+k | i ∈ [0, 11], n ∈ [0, 3]} to represent all reference viewpoints. Since the classification results of the input viewpoints and the reference viewpoints have similar appearances, the rotation quaternion under the reference viewpoint can be used as a rough estimate of the 3D rotation component of the input viewpoint, that is, R_init = q_ref cls , where cls is the predicted viewpoint category, R_init represents the initial estimate of the 3D rotation, and q_ref cls represents the rotation quaternion corresponding to the reference viewpoint category cls.

[0026] Step 2: Use the Open3D point cloud function library in Python to synthesize training data and generate annotation files in YOLO format. First, load the target 3D model, set a rendering window with a size of 640x640, and set the background color to pure black. For each reference viewpoint View_ref, apply the rotation quaternion corresponding to this viewpoint to the virtual camera, and then apply random rotations and translations within a certain range to the camera.

[0027] Step 3: Denote the rendered 3D model image as I. Since the image background color is single at this time, it is very easy to extract the contour C of the target 3D model, the bounding box coordinates (including the upper left corner coordinates (x0, y0), width w, and height h), and the mask M. The data with a simple black background is too simple and prone to overfitting. Use the public image dataset VOC2012 combined with the mask M to synthesize diverse backgrounds for I. The specific approach is as follows:

[0028] (1) Randomly select 1 image from the dataset and scale it to a size of 640x640 as the background image B,

[0029] (2) To make the transition of the boundary between the 3D model and the background smoother, the mask M is inverted and the edges of the mask are blurred using a 3x3 Gaussian kernel Gb (Gaussian blur), and then multiplied by the background image B.

[0030] (3) After performing the same Gaussian blur on the original mask M, it is multiplied by the 3D model rendering image I.

[0031] (4) The blurred background image and the blurred rendering image are added together to obtain the final result O. Expressed by the formula as O = Gb(M)×I + Gb(1 - M)×B. The result image and the bounding box coordinates are saved locally, and at the same time, the index of this view is used as the category of the image. Some images are as Figure 2 .

[0032] Step 4: The lightweight design and high efficiency of the Yolov8 framework enable it to quickly and accurately detect targets in real-time scenarios. By adopting YOLOv8 as the pose initializer, the real-time performance and accuracy of pose estimation are improved, providing a reliable basis for subsequent pose tracking and positioning.

[0033] Step 5: The 3D translation components of an object include the translation amounts in the x, y, and z-axis directions. Since the translation components in the xy-axis directions are affected by the translation amount in the z-axis direction (i.e., depth), the present invention decouples the 3D translation and adopts a method of first determining the z-axis component and then calculating the xy-axis components. Since the prior information of depth d is implicit in 48 reference views in the classification task, the specific method for solving the z-axis component is as follows:

[0034] First, determine and fix the depth d of the reference view, and then record the bounding box area of the target object under each reference view respectively. After the view classification is completed, the bounding box area {S_ref i×4+k |i∈[0,11],k∈[0,3]} of the object in the input image view is obtained through the output result of the classification network. According to the proportional relationship between the bounding box areas during classification, the input z-axis component z_init is obtained, as shown in Formula 3-1.

[0035]

[0036] S_input represents the pixel area of the bounding box obtained by detecting the target object in the input image, cls represents the view category of the target object detected in the input image, and S_ref cls represents the pixel area of the bounding box of the target object in the image under the standard reference view corresponding to cls.

[0037] After obtaining the z-axis component, the center point coordinates (x_center, y_center) of the bounding box are approximately regarded as the center point coordinates (u, v) of the target object. Combining with Formula 3-2, we can get:

[0038]

[0039] Among them, x_init and y_init are rough estimates of the x-axis and y-axis components respectively. K represents the internal parameter matrix of the camera, which can be obtained through calibration. Finally, the 3D translation component T_init of the target object is [x_init, y_init, z_init] T 。

[0040] Step 6: The simple optical flow network adopts a shared encoder-decoder architecture. Among them, the encoder part is used to extract features and contains 10 downsampling blocks. The number of input channels of the first downsampling block is 6. Each downsampling block is composed of 1 convolutional layer and 1 Leaky_ReLU activation function layer in series. The decoder part is used to predict the optical flow and is composed of 4 upsampling blocks, 1 convolutional layer and 1 transposed convolutional layer. Each upsampling block contains two parallel branches. One branch contains 1 transposed convolutional layer in series with 1 Leaky_ReLU activation function layer, and the other branch contains 1 convolutional layer in series with 1 transposed convolutional layer. Finally, they are cropped to the same size and then stitched together as the output. In the decoding stage, the output of the encoder feature extraction part will be fused, and multi-scale features will be fused together to predict multi-scale optical flow. Such connections run through the entire network, and there are a total of four fusion processes, as Figure 3 。

[0041] Step 7: Given the initial rough estimate of the 6D pose of the target object in the test image, the pose optimization network can give the relative pose transformation that conforms to the motion relationship between the rendered view of the target object's three-dimensional model in the given pose and the current observed image. The relative pose transformation can be used to improve the accuracy of the input pose and re-render the target with the optimized pose. After multiple iterations like this, the pose estimated by the network will become more and more accurate, as Figure 4 。

[0042] Step 8: First, merge the concatenated operators. The 1st and 4th branches are both separate convolutional operations and do not need to be merged on the branches. Figure 5 The operations of the multi-branch module from left to right can be expressed as:

[0043] y = Conv k×k (Conv 1×1 (x)) = W k×k (W 1×1 x + b 1×1 ) + b k×k (4 - 1)

[0044] Among them, Conv represents the convolution operation, W and b represent the weight and bias parameters of the convolution respectively, the subscript represents the size of the corresponding convolution kernel, x and y represent the input and output respectively, k represents the size of the convolution kernel, and formula (4-1) can be further expanded and combined into a single convolution:

[0045] y = W fused2 x + b fused2 (4-2)

[0046] Among them, W fused2 = W k×k W 1×1 , b fused2 = W k×k b 1×1 + b k×k , which respectively represent the weight and bias after convolution fusion. The operation of the third branch can be expressed as:

[0047] y = Avgpool k×k (Conv 1×1 (x)) = W_Avg k×k (W 1×1 x + b 1×1 ) + b_Avg k×k (4-3)

[0048] Among them, Avgpool represents the average pooling operation, W_Avg and b_Avg represent the convolution weight and bias equivalently converted from the average pooling operation, W_Avg is a matrix with a length and width of k×k and a value of 1 / k 2 , b_Avg is 0 and does not participate in the operation and can be ignored. Formula (4-3) can also be further expanded and combined into a single convolution:

[0049] y = W fused3 x + b fused3 (4-4)

[0050] Among them, W fused3 = W_Avg k×k W 1×1 , b fused3 = W_Avg k×k b 1×1 .

[0051] Finally, merge all the operators of the 4 parallel branches, and its merging operation can be expressed as formula (4-5):

[0052]

[0053] Among them, W fused_all = W 1×1 + W fused2 + W fused3 + Wk×k ,b fused_all =b 1×1 +b fused2 +b fused3 +b k×k , for the case where the shapes of the weight matrices are inconsistent, a 1x1 convolutional kernel can be equivalent to a k×k convolutional kernel after padding (k-1) / 2 circles of 0s around it. During the specific process of the multi-branch integration operation, the equivalent schematic diagrams of different operators are as Figure 6 shown.

[0054] Step 9: Load the pre-trained pose initialization and pose optimization models, obtain the image input through the camera, first obtain the rough pose through the pose initializer, and then perform iterative optimization through the pose optimizer to output the accurate 6D pose. According to the predicted 6D pose and the assembly relationship tree, render the model of the part to be assembled onto the display device to guide the assembly task through the animation.

Claims

1. A tracking and positioning method for precision electronic product AR assembly guidance, characterized in that, It includes the following steps: Step 1: Divide 48 uniform viewpoints according to the 12 vertices of the regular icosahedron as the camera view classification standard; Step 2: Use the Open3D library to load the 3D model of the electronic equipment, and randomly rotate and move the camera slightly for rendering under 48 camera views; Step 3: Use the VOC2012 dataset as the background, synthesize the background for the rendered images and save them, and generate annotation data for categories and bounding boxes; Step 4: Use the object detection framework YOLOv8 to train on the synthesized dataset to obtain a model that can predict the rough view information of the target object; Step 5: Convert the label view classified by YOLOv8 into the corresponding 3D rotation, and calculate the 3D translation by combining the camera internal parameters and the center coordinates of the detection box, which is used as the initial value for subsequent optimization; Step 6: Replace the convolutional module of the simple optical flow network with a multi-branch module and train it using the flyingchairs optical flow dataset; Step 7: Based on the trained simple optical flow network model, use the synthesized dataset of the 3D model of the electronic equipment for pose optimization training to obtain a model that can iteratively optimize the rough pose; Step 8: Merge the branches of the multi-branch module of the simple optical flow network using structural reparameterization; Step 9: Combine the trained initial pose network and pose optimization network to predict the 6D pose of the electronic equipment for real-time tracking and positioning.

2. The tracking and positioning method for AR assembly guidance of precision electronic products according to claim 1, characterized in that: In the said step 1, the center of the target three-dimensional model is the center of a regular icosahedron. The distance between the camera optical center and the object center is d, and the camera is located at a vertex P of the regular icosahedron. i , and the exact coordinates of the 12 vertices {P i | i ∈ [0, 11]} can be obtained according to the calculation formula of the geometric information of the regular icosahedron. Orient the camera towards the center of the three-dimensional model. At this time, the camera rotates 90°, 180°, and 270° clockwise respectively around the z-axis of its own coordinate system, plus the non-rotated state. A total of 4 reference viewpoints are obtained under this vertex. 48 reference viewpoints can be obtained for all vertices, denoted as {View_ref i×4+n | i ∈ [0, 11], n ∈ [0, 3]}, and record the rotation quaternion of the camera relative to the three-dimensional model under each viewpoint, denoted as {q_ref i×4+n | i ∈ [0, 11], n ∈ [0, 3]}, as well as 48 images of the three-dimensional model under the current viewpoint, denoted as {I_ref i×4+n | i ∈ [0, 11], n ∈ [0, 3]}. i represents the number of vertices, and n from 0 to 3 respectively represents rotating 0°, 90°, 180°, and 270° clockwise around the z-axis of its own coordinate system. Since the classification results of the input viewpoints and the reference viewpoints have similar appearances, the rotation quaternion under the reference viewpoint can be used as a rough estimate of the 3D rotation component of the input viewpoint, that is, R_init = q_ref cls , where cls is the predicted viewpoint category, R_init represents the initial estimate of the 3D rotation, and q_ref cls represents the rotation quaternion corresponding to the reference viewpoint category cls.

3. The tracking and positioning method for AR assembly guidance of precision electronic products according to claim 1, characterized in that: The steps for synthesizing the background for the rendered images in Step 3 include: (1) Denote the rendered 3D model image as I, and extract the contour C, bounding box coordinates, and mask M of the target 3D model; (2) Randomly select 1 image from the VOC2012 dataset and scale it to a size of 640x640 as the background image B; (3) Invert the mask M and use a 3x3 Gaussian kernel Gb to blur the mask edge, and then multiply it by the background image B; (4) Perform the same Gaussian blur on the original mask M as in step (3) and then multiply it by the 3D model rendered image I; (5) Add the blurred background image and the blurred rendered image to get the final result O, which is expressed by the formula O = Gb(M)×I + Gb(1 - M)×B.

4. The tracking and positioning method for AR assembly guidance of precision electronic products according to claim 1, characterized in that: When calculating the 3D translation in Step 5, first determine the z-axis component and then calculate the xy-axis components, specifically including: First, determine and fix the depth d of the reference perspective, and then record the bounding box area of the target object under each reference perspective. After the perspectives are classified, obtain the bounding box area {S_ref i×4+n |i∈[0,11],n∈[0,3]} of the object in the input image perspective through the output result of the classification network. Obtain the z-axis component z_init of the input according to the proportional relationship between the bounding box areas classified as shown in Equation 3-1; S_input represents the pixel area of the bounding box obtained by detecting the target object in the input image, cls represents the perspective category obtained by detecting the target object in the input image, and S_ref cls represents the pixel area of the bounding box of the target object in the image under the standard reference perspective corresponding to cls; After obtaining the z-axis component, approximate the center coordinates (x_center, y_center) of the bounding box as the center coordinates (u, v) of the target object. Combining with formula 3-2, we can get: Among them, x_init and y_init are the rough estimates of the x-axis and y-axis components respectively, and K represents the internal parameter matrix of the camera; finally, the 3D translation component T_init of the target object = [x_init, y_init, z_init] T .

5. The tracking and positioning method for precision electronic product AR assembly guidance according to claim 1, characterized in that: The optical flow network in Step 6 adopts a shared encoder-decoder architecture. Among them, the encoder part is used to extract features and contains 10 downsampling blocks. The decoder part is used to predict the optical flow and consists of 4 upsampling blocks, 1 convolutional layer, and 1 transposed convolutional layer. In the decoding stage, the output of the encoder feature extraction part is fused, and multi-scale features are fused together for multi-scale optical flow prediction, with a total of four fusion processes.

6. The tracking and positioning method for AR assembly guidance of precision electronic products according to claim 5, characterized in that: The downsampling block consists of a convolutional layer and a Leaky_ReLU activation function layer connected in series. The number of input channels of the first downsampling block is 6. The upsampling block contains two parallel branches. One branch consists of a transposed convolutional layer connected in series with a Leaky_ReLU activation function layer, and the other branch consists of a convolutional layer connected in series with a transposed convolutional layer. Finally, they are cropped to the same size and then concatenated as the output.

7. The tracking and positioning method for AR assembly guidance of precision electronic products according to claim 1, characterized in that: Step 8 specifically includes: First, merge the concatenation operators. The first and fourth branches are both individual convolutional operations and do not need to be merged on the branches. The operation of the second branch is expressed as: y = Conv k×k (Conv 1×1 (x)) = W k×k (W 1×1 x + b 1×1 ) + b k×k (4 - 1) Among them, Conv represents the convolutional operation, W and b represent the weights and bias parameters of the convolution respectively, and the subscript represents the size of the corresponding convolution kernel. Formula (4-1) is further expanded and merged into a single convolution: y = W fused2 x + b fused2 (4 - 2) Among which W fused2 = W k×k W 1×1 , b fused2 = W k×k b 1×1 + b k×k , which respectively represent the weights and biases after convolutional fusion; The operations for the third branch are expressed as: y = Avgpool k×k (Conv 1×1 (x)) = W_Avg k×k (W 1×1 x + b 1×1 ) + b_Avg k×k (4 - 3) Among them, Avgpool represents the average pooling operation, and W_Avg and b_Avg represent the convolution weights and biases equivalently converted from the average pooling operation. W_Avg is a matrix with a length and width of k×k and a value of 1 / k 2 , b_Avg is 0 and does not participate in the operation. The formula (4-3) can also be continuously expanded and combined into a single convolution: y = W fused3 x + b fused3 (4 - 4) where W fused3 = W_Avg k×k W 1×1 ,b fused3 = W_Avg k×k b 1×1 . Finally, merge all the operators of the four parallel branches, and its merging operation can be expressed as formula (4-5): where W fused_all = W 1×1 + W fused2 + W fused3 + W k×k , b fused_all = b 1×1 + b fused2 + b fused3 + b k×k For the case where the shapes of the weight matrices are inconsistent, a 1x1 convolutional kernel can be equivalent to a k×k convolutional kernel with (k-1) / 2 circles of 0s padded around it.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the tracking and positioning method for the precision electronic product AR assembly guidance described in any one of claims 1 to 7.