Visual positioning method for assembling mobile phone parts

By employing a de-reflection model that combines the Stokes vector method and a polarization fusion network, along with a joint deblurring network for reflection suppression and motion compensation, the problems of reflection and image blurring in traditional visual positioning are solved, achieving high-precision visual positioning in motion.

CN120912673BActive Publication Date: 2026-02-27信阳鑫达智能科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511042555.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2026-02-27
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

Traditional visual positioning solutions face challenges when assembling mobile phone parts, including image distortion caused by surface reflections and image blurring due to minute movements of parts during assembly, resulting in dynamic errors in the positioning results.

Method used

A de-reflection model employing the Stokes vector method and polarization fusion network is combined with a joint deblurring network for reflection suppression and motion compensation. A clear assembly map is generated through spatial, temporal, and frequency domain processing, and high-precision visual positioning is achieved using a pose mapping network.

Benefits of technology

It effectively reduces positioning errors caused by reflection and movement, ensuring high-precision visual positioning in motion and improving the accuracy and efficiency of assembly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912673B_ABST
    Figure CN120912673B_ABST
Patent Text Reader

Abstract

The application discloses a visual positioning method for mobile phone part assembly, relates to the field of image processing technology, and determines a part to be assembled, generates a reflection suppression assembly diagram by fusing three assembly diagrams under three polarization angles by using a Stokes vector method and a polarization fusion network, generates an enhanced assembly diagram by labeling the position of the part to be assembled in the reflection suppression assembly diagram by using a semantic segmentation network, generates compensation parameters of multiple frames by using a three-dimensional convolution network in combination with an optical flow field, generates a reconstructed assembly diagram by performing Fourier transform, filtering and inverse Fourier transform on the reflection suppression assembly diagram, compensates the enhanced assembly diagram based on the compensation parameters by using a cross-domain fusion network, fuses the reconstructed assembly diagram to generate a clear assembly diagram to remove motion blur, inputs the clear assembly diagrams of the multiple frames and a camera external parameter matrix into a pose mapping network, outputs target poses of the multiple frames, and predicts the target poses at the completion frame by using a time sequence network, and realizes intelligent assembly at the completion frame based on an intelligent control algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a visual positioning method for mobile phone part assembly. BACKGROUND

[0002] In the process of mobile phone part assembly, visual positioning technology is the core link to realize automatic and precise assembly. Through real-time detection and analysis of part position and attitude, visual positioning can effectively improve assembly efficiency, reduce labor cost, and ensure the consistency of product quality. Therefore, how to accurately determine the accuracy of part assembly position through visual image has become a key problem to improve the level of intelligent manufacturing.

[0003] The existing Chinese patent with authorization announcement No. CN117372528B discloses a visual image positioning method for mobile phone shell modular assembly. The pre-assembly positioning image and the reference image are collected by a CCD camera, the feature map is extracted by using a deep neural network model, and after positioning interactive comparison and analysis, it is determined whether the assembly is accurate. This scheme realizes automatic positioning detection through feature decomposition, sequence interactive comparison and classifier training, and reduces the dependence on manual work.

[0004] Currently, the traditional visual positioning scheme faces two technical bottlenecks when processing mobile phone part assembly: on the one hand, the reflection of the part surface is easy to cause local information distortion of the image, making it difficult to accurately extract the real features; on the other hand, the small movement of the part during assembly may cause image blur, resulting in dynamic error in the positioning result. SUMMARY

[0005] The present application proposes a visual positioning method and system for mobile phone part assembly to realize dynamic visual positioning and assembly considering reflection suppression and motion compensation, aiming at the deficiencies of the existing technology.

[0006] The technical solution to achieve the purpose of the present application is as follows:

[0007] The visual positioning method for mobile phone part assembly comprises the following steps:

[0008] Determine the parts to be assembled, and collect K frames of assembly images at polarization angles of 0°, 45° and 90° respectively to construct an assembly image stack Wherein, And are the assembly images of the kth frame at polarization angles of 0°, 45° and 90° respectively;

[0009] Calculate the polarization degree map of the kth frame based on the assembly images of the kth frame at three polarization angles using the Stokes vector method And polarization angle map Use polarization fusion network to fuse the polarization degree map of the kth frame And polarization angle map fusing the weight mask into the kth frame and generating a specular reflection suppression assembly graph of the kth frame by acting on the spliced result of the assembly graph of the kth frame under three polarization angles through convolution constructing a sequence of specular reflection suppression assembly graphs

[0010] the joint deblurring network adopts a semantic segmentation network in the spatial domain to label the specular reflection suppression assembly graph of the kth frame to generate an enhanced assembly graph of the kth frame extracting spatiotemporal features in the sequence of specular reflection suppression assembly graphs in the temporal domain using a three-dimensional convolution network and combining an optical flow field to generate a sequence of compensation parameters the specular reflection suppression assembly graph of the kth frame is sequentially subjected to Fourier transform, band-pass filtering and inverse Fourier transform to generate a reconstructed assembly graph of the kth frame based on the sequence of compensation parameters of the kth frame compensate for the enhanced assembly graph of the kth frame and the reconstructed assembly graph of the kth frame fusion to generate a clear assembly graph of the kth frame constructing a sequence of clear assembly graphs

[0011] the clear assembly graph of the kth frame is input into a pose mapping network together with a camera extrinsic matrix T to extract a target feature map of the kth frame belonging to a to-be-detected position in the clear assembly graph of the kth frame and flatten it to splice the encoding result of the camera extrinsic matrix T to input a multilayer perceptron to output a target pose Δ of the kth frame k , constructing a sequence of target poses

[0012] determining a completed frame, inputting the sequence of target poses into a time sequence network to generate a predicted target pose at the completed frame, and based on an intelligent control algorithm, controlling an execution device to adjust the to-be-assembled part to the predicted target pose at the completed frame and assemble it.

[0013] Further, the Stokes vector method comprises the following steps:

[0014] the assembly graph of the kth frame under polarization angles of 0°, 45° and 90° is respectively subjected to Gaussian filtering to generate denoised assembly graphs of the kth frame under the corresponding polarization angles;

[0015] summing the denoised assembly graphs of the kth frame under polarization angles of 0° and 90° to generate a total light intensity graph I of the kth frame k ; ​

[0016] the difference between the denoised assembly map of the kth frame under the polarization angle of 0° and 90° as the polarization map of the kth frame

[0017] subtracting the denoised assembly map of the kth frame under the polarization angle of 0° and 90° from twice the denoised assembly map of the kth frame under the polarization angle of 45° to generate the oblique polarization map of the kth frame

[0018] taking the square root of the sum of squares of the polarization map of the kth frame and the oblique polarization map of the kth frame and dividing by the total light intensity map I of the kth frame k to generate the degree of polarization map of the kth frame

[0019] taking the half of the arctangent value of the ratio of the oblique polarization map of the kth frame and the polarization map of the kth frame as the polarization angle map of the kth frame

[0020] Further, the polarization fusion network comprises a polarization encoding layer and a channel fusion layer;

[0021] The polarization encoding layer fuses the degree of polarization map of the kth frame and the polarization angle map of the kth frame using the linear layer of the multilayer perceptron and generates the polarization feature map F of the kth frame based on the ReLU function mapping k ;

[0022] The channel fusion layer normalizes the polarization feature map F of the kth frame k to the weight mask of the kth frame and performs channel convolution on the splicing result in the channel dimension of the assembly map of the kth frame under the three polarization angles to generate the anti-glare assembly map of the kth frame

[0023] Further, the polarization fusion network is jointly trained with the normal vector estimation network, the input of the normal vector estimation network is the assembly map of the kth frame under the three polarization angles, and the normal vector map of the kth frame is output in turn through convolution, downsampling and upsampling taking the normalized result of the normal vector map of the kth frame and the anti-glare assembly map of the kth frame to perform dot product, taking the square of the two-norm of the dot product result as the reflection prior loss of joint training, updating the parameters of the polarization fusion network and the normal vector estimation network to make the reflection prior loss continuously decrease, and when the reflection prior loss converges, the joint training is completed.

[0024] Specifically, the semantic segmentation network assembles the reflection-suppressed group image of the kth frame The assembled reflection-suppressed group image of the kth frame is divided into small blocks and ensured to have a pixel overlap area between adjacent small blocks, and a sliding convolution kernel is used to convolve each small block and arrange the convolution result according to the arrangement of the small blocks into a semantic embedding image of the kth frame The semantic embedding image of the kth frame is input into a Transformer improved block, and the output is a semantic feature map of the kth frame The semantic feature map of the kth frame is adjusted by bilinear interpolation The dimension of the semantic feature map of the kth frame is adjusted by bilinear interpolation The semantic class is segmented by a Softmax function to identify the reflection-suppressed group image of the kth frame The pixel points belonging to the to-be-assembled position in the reflection-suppressed group image of the kth frame are generated into an enhanced group image of the kth frame The Transformer improved block replaces the full connection layer and the ReLU function in the existing Transformer structure with a point-wise convolution and a GeLU function, respectively.

[0025] Further, a sequence of compensation parameters is generated The sequence of compensation parameters is generated by the following steps:

[0026] The three-dimensional convolution network uses a three-dimensional convolution kernel to perform spatiotemporal feature extraction on the sequence of compensation parameters The spatiotemporal feature map sequence is generated by batch normalization and ReLU function mapping

[0027] The pyramid optical flow mapping block constructs an L-layer pyramid according to the reflection-suppressed group image sequence The reflection-suppressed group image of the kth frame and the k+1th frame constructs an L-layer pyramid and calculates the forward preliminary optical flow map from the kth frame to the k+1th frame in time order and reverse order, respectively And the backward preliminary optical flow map The confidence mask M from the kth frame to the k+1th frame is generated by filtering and cycle consistency detection k,k+1 And the forward optical flow map v from the kth frame to the k+1th frame is obtained by correction k,k+1 And the backward optical flow map v k+1,k ;

[0028] The compensation parameter f1 of the 1st frame is set V And the compensation parameter of the Kth frame The forward optical flow map v from the 1st frame to the 2nd frame 1,2 And the backward optical flow map v from the K-1th frame to the Kth frame K,K-1 When 2≤k≤K-1, the compensation parameter of the kth frame is set The forward optical flow map v from the kth frame to the k+1th frame k,k+1 The sequence of compensation parameters is generated

[0029] Further, the processing steps of the pyramid optical flow mapping block include:

[0030] Obtaining the anti-reflective assembled image of the kth frame and the k+1th frame, respectively performing L-1 times downsampling, and constructing the sampling pyramid of the kth frame and the k+1th frame;

[0031] Setting the forward preliminary optical flow map of the kth frame to the k+1th frame at the Lth layer and the backward preliminary optical flow map which are all 0 images and are refined step by step from the Lth layer downwards;

[0032] For the lth layer, 1≤l≤L-1, a search box with the same dimension as the convolution image of the k+1th frame at the l+1th layer is generated at the center of the convolution image of the kth frame at the lth layer and the k+1th frame at the lth layer, the search box is moved in the convolution image of the k+1th frame at the lth layer, and the corresponding forward pixel offset is determined, and the up-sampling result of the forward preliminary optical flow map of the kth frame to the k+1th frame at the l+1th layer is superimposed to generate the forward preliminary optical flow map of the kth frame to the k+1th frame at the lth layer

[0033] The search box in the convolution image of the k+1th frame at the lth layer is moved back to the center position, the search box is moved in the convolution image of the kth frame at the lth layer, and the corresponding backward pixel offset is determined, and the up-sampling result of the backward preliminary optical flow map of the kth frame to the k+1th frame at the l+1th layer is superimposed to generate the backward preliminary optical flow map of the kth frame to the k+1th frame at the lth layer

[0034] After refining to the 1th layer, the process is stopped, and the forward preliminary optical flow map of the kth frame to the k+1th frame at the 1th layer and the backward preliminary optical flow map are taken as the forward preliminary optical flow map of the kth frame to the k+1th frame and the backward preliminary optical flow map of the kth frame to the k+1th frame

[0035] Further, the forward preliminary optical flow map of the kth frame to the k+1th frame and the backward preliminary optical flow map of the kth frame to the k+1th frame are filtered by bilateral filtering, and then the loop consistency detection is performed, the anti-reflective assembled image of the kth frame is forward predicted according to the forward preliminary optical flow map of the kth frame to the k+1th frame and is backward predicted again according to the backward preliminary optical flow map of the kth frame to the k+1th frame , and the reconstructed anti-reflective assembled image of the kth frame is restored and generated and is compared with the anti-reflective assembled image of the kth frame Compare the pixel differences at the same pixel coordinates (μ, ν). If the pixel difference is less than the difference threshold, set the confidence mask M from frame k to frame (k+1) to... k,k+1 The confidence level at pixel coordinates (μ,ν) is equal to 1, and 0 in all other cases. The confidence mask M from frame k to frame (k+1) is used. k,k+1 Compare with the forward preliminary optical flow maps from frame k to frame (k+1) respectively. and backward preliminary optical flow map Multiply to generate the forward optical flow graph v from frame k to frame (k+1). k,k+1 and backward optical flow map v k+1,k .

[0036] Specifically, when 1≤k≤K-1, the cross-domain fusion network utilizes the compensation parameters of the k-th frame. Determine the enhanced assembly graph of the k-th frame. The pixels in the image and the enhanced assembly map of the (k+1)th frame The corresponding pixel in the enhanced assembled image from frame k+1. Find the pixel values ​​of the four neighboring pixels around the corresponding pixel and update the enhanced assembly map of the k-th frame using inverse distance weighting. The pixel values ​​in the image are used to generate the motion-compensated assembly image of the k-th frame. And by using preset weights and reconstructing the graph of the k-th frame. Perform weighted summation to generate a clear assembly diagram of the k-th frame. Construct a clear assembly diagram sequence

[0037] Furthermore, the joint deblurring network is pre-trained adversarially with the blurrer and discriminator. Multiple pre-collected real-world sharp assembly images are input into the blurrer, which adds random Gaussian noise to the sharp assembly images and simulates the generation of multiple blurred assembly images through fully connected layers and deconvolutional layers. These multiple blurred assembly images are then deblurred by the joint deblurring network to generate multiple sharp assembly images. The discriminator extracts image features from the multiple sharp assembly images and their corresponding real-world sharp assembly images through convolutional layers and outputs the average similarity probability using the Sigmoid function. The parameters of the blurrer, joint deblurring network, and discriminator are continuously updated until the average similarity probability is maximized.

[0038] Specifically, pose mapping networks include residual networks, encoder-transformers, and regression networks;

[0039] A clear assembly diagram of the k-th frame from a residual network. Perform convolution to generate the global feature map of the k-th frame. The global feature map of the k-th frame The result of max pooling and pointwise convolution is combined with the global feature map of the k-th frame. Residual superposition, generate the feature map F of the kth frame k , based on the joint deblurring network, the feature map F of the kth frame is labeled and filtered k The features in the kth frame that do not belong to the position to be detected are generated as target feature maps

[0040] The encoding converter disassembles the camera extrinsic parameter matrix T to generate a rotation matrix R and a translation vector a, and converts the rotation matrix R to a quaternion vector q using a quaternion;

[0041] The regression network generates the target feature map of the kth frame Flattening and splicing the quaternion vector q and the translation vector a into a multilayer perceptron, and generating the target pose of the kth frame through two consecutive linear layers, batch normalization layers and ReLU functions k .

[0042] Compared with the prior art, the present application introduces a reflection correction and motion compensation mechanism to solve the technical problems of traditional visual positioning, and by constructing a de-reflection model based on the Stokes vector method and a polarization fusion network, the assembled area to be assembled in the assembled image can be accurately identified and suppressed, and a reflection suppression assembly image can be generated, thereby reducing the interference of reflection; a joint deblurring network is used to perform edge enhancement, motion modeling and filtering on the reflection suppression assembly image from the spatial, temporal and frequency domains respectively, and based on the processing results of the spatial, temporal and frequency domains, the image blur caused by the movement of the parts during assembly is dynamically compensated, providing a clear assembly image for subsequent movement trajectory capture of the position to be assembled and target pose prediction combined with the time dimension, effectively reducing the dynamic error of positioning, and ensuring that high-precision visual positioning can be achieved even when the mobile phone to be assembled is in a moving state. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 A flowchart of a visual positioning method for mobile phone part assembly is shown in the figure.

[0044] Figure 2 A flowchart of a pyramid optical flow mapping block is shown in the figure.

[0045] Figure 3 A flowchart of a process for generating forward and backward optical flow maps is shown in the figure.

[0046] Figure 4 A pose mapping network model is shown in the figure. DETAILED DESCRIPTION

[0047] The present application will be further described in detail below in conjunction with the drawings and examples.

[0048] Example 1

[0049] As Figure 1As shown, this invention discloses a visual positioning method for assembling mobile phone parts, comprising the following steps:

[0050] Once the parts to be assembled are identified, an industrial camera is used to acquire K frames of assembly images at polarization angles of 0°, 45°, and 90°, respectively, and an assembly image stack is constructed. in, and The assembly diagrams of the k-th frame at polarization angles of 0°, 45°, and 90°, respectively;

[0051] The Stokes vector method is used to preprocess the assembled image of the k-th frame at polarization angles of 0°, 45°, and 90°, and the polarization degree map of the k-th frame is calculated. and polarization angle diagram The polarization fusion network will convert the polarization degree map of the k-th frame. and polarization angle diagram The weight mask is transformed into the k-th frame through fusion. By convolving the assembly images of the k-th frame at polarization angles of 0°, 45°, and 90°, glare and artifacts caused by component reflections are compensated for, and a reflection-suppressed assembly image of the k-th frame is generated. Constructing a sequence of reflection suppression assembly diagrams

[0052] A joint deblurring network is used to process the image in the spatial, temporal, and frequency domains. In the spatial domain, a semantic segmentation network is used to annotate the reflection suppression assembly image of the k-th frame. The assembly positions in the image are analyzed, and edge details are enhanced to generate an enhanced assembly image for the k-th frame. A three-dimensional convolutional network is used in the temporal domain to extract the reflection suppression assembly map sequence. The spatiotemporal features of adjacent frames are combined with the optical flow field constructed by pyramid optical flow mapping to generate a sequence of compensation parameters. Assemble the reflection suppression diagram of the k-th frame. The reconstructed assembly image of the k-th frame is generated by transforming the image to the frequency domain using Fourier transform, followed by bandpass filtering and inverse Fourier transform. Using cross-domain fusion network based on compensation parameters of the k-th frame Enhanced assembly graph of frame k Perform motion compensation and assemble the image with the reconstructed graph of the k-th frame. Fusion generates a clear assembly image of the k-th frame. Construct a clear assembly diagram sequence

[0053] The clear assembly diagram of the k-th frame The pose mapping network is co-inputted with the pre-calibrated camera extrinsic parameter matrix T to extract the sharp assembled image of the k-th frame. Global feature map in And based on the annotation screening of the joint deblurring network, a target feature map of the kth frame is generated Encode and convert the camera extrinsic matrix T, and splice the flattened results of the target feature map of the kth frame Output the target pose of the kth frame based on a multi-layer perception mapping Construct a target pose sequence Wherein, p k and are the center world coordinates and three-dimensional deflection angle vectors of the position to be assembled at the kth frame, and the three-dimensional deflection angle vector Records the deflection angle of the position to be assembled about the three coordinate axes of the world coordinate system;

[0054] Based on the assembly duration of the execution device, determine the completion frame where the position to be assembled is located, and input the target pose sequence into a timing network to generate a predicted target pose at the completion frame, and control the execution device to accurately regulate the part to be assembled to the predicted target pose at the completion frame based on an intelligent control algorithm to achieve assembly, wherein the timing network includes a long short-term memory network, a recurrent neural network, a gated recurrent neural network, and a Transformer, and the intelligent control algorithm includes a PID control algorithm, a fuzzy control algorithm, various swarm intelligence algorithms, and reinforcement learning. Swarm intelligence algorithms include particle swarm optimization, greedy algorithm, mouse swarm optimization algorithm, grey wolf optimization algorithm and whale optimization algorithm.

[0055] Further, the Stokes vector method includes the following steps:

[0056] Use Gaussian filtering to remove noise in the kth frame assembly map at polarization angles of 0°, 45° and 90°, to generate a denoised assembly map at polarization angles of 0°, 45° and 90°.

[0057] The sum of the denoised assembly maps at polarization angles of 0° and 90° of the kth frame is taken as the total light intensity map of the kth frame The total light intensity map I k of the kth frame reflects the overall illumination intensity of the assembly scene;

[0058] The difference between the denoised assembly maps at polarization angles of 0° and 90° of the kth frame is taken as the polarization map of the kth frame The polarization map of the kth frame Reflects the intensity difference of polarized light in the horizontal and vertical directions, and is used to judge the directionality of the polarization direction;

[0059] Subtract the denoised assembly map at the polarization angle of 0° and 90° of the kth frame from twice the denoised assembly map at the polarization angle of 45° of the kth frame to generate the diagonal polarization map of the kth frame

[0060] the polarization map of the kth frame the slant polarization map of the kth frame square root of the sum of squares of the ratio of the slant polarization map of the kth frame k and the polarization map of the kth frame, to generate the polarization degree map of the kth frame wherein the polarization degree map of the kth frame reflects the reflection condition at each place in the assembled scene, when the polarization degree at the pixel coordinate (μ, v) of the kth frame , it indicates that the pixel coordinate (μ, v) is completely polarized light, when the polarization degree at the pixel coordinate (μ, v) of the kth frame , it indicates that the pixel coordinate (μ, v) is completely non-polarized light, when the polarization degree at the pixel coordinate (μ, v) of the kth frame , it indicates that the pixel coordinate (μ, v) is partially non-polarized light, since most of the mobile phone parts are metal materials, the polarization degree map of the kth frame actually reflects the interference of the reflection caused by metal materials on the collection of industrial cameras;

[0061] the slant polarization map of the kth frame and the polarization map of the kth frame half of the arctangent value of the ratio of the slant polarization map of the kth frame the polarization angle map of the kth frame reflects the angle between the polarization direction and the horizontal direction.

[0062] Further, the polarization fusion network comprises a polarization encoding layer and a channel fusion layer;

[0063] The polarization encoding layer fuses the polarization degree map of the kth frame and the polarization angle map of the kth frame and adjusts to the same dimension as the assembled image of the kth frame under the polarization angles of 0°, 45° and 90°, uses the ReLU function to increase the non-linear representation ability, and generates the polarization feature map F k of the kth frame;

[0064] The channel fusion layer normalizes the polarization feature map F k of the kth frame to the weight mask of the kth frame pastes the assembled images of the kth frame under the polarization angles of 0°, 45° and 90° in the channel dimension, and performs channel convolution based on the weight mask of the kth frame to selectively enhance or suppress the polarization characteristics of different pixel positions in different channels, to generate the reflection suppression assembled image of the kth frame wherein the weight mask of the kth frame This reflects the polarization characteristics of the assembled image of the k-th frame at the same pixel position under polarization angles of 0°, 45°, and 90°, which affects the reflection suppression of the fused assembled image of the k-th frame. The contribution of * is indicated by the convolution.

[0065] Furthermore, the polarization fusion network needs to be jointly trained with the normal vector estimation network. The normal vector estimation network adopts an encoder-decoder structure. The input is the assembled image of the k-th frame at polarization angles of 0°, 45°, and 90°. The encoder captures global features through convolution and downsampling, and the decoder restores the resolution through upsampling, outputting the normal vector map of the k-th frame. The normal vector map of the k-th frame The normalized result and the reflection suppression assembly diagram of the k-th frame Perform a dot product, and use the squared 2-norm of the dot product result as the reflection prior loss for joint training. The reflection prior loss utilizes the normal vector map of the k-th frame. The physical properties of light are used to constrain the specular reflection of light on the surface of metal parts from an energy perspective, thereby reducing the reflection suppression in the k-th frame. To address the dependence on specular reflection, the parameters of the polarization fusion network and the normal vector estimation network are jointly optimized based on the reflection prior loss to continuously reduce the reflection prior loss, thereby improving the reflection suppression assembly map of the k-th frame. It is closer to the ideal state of no reflection interference.

[0066] Specifically, the semantic segmentation network adopts the SegFormer structure, which assembles the reflection suppression map of the k-th frame. The image is divided into small blocks of the same shape, with a certain pixel overlap between adjacent blocks to ensure the continuity of semantic information extracted by subsequent convolutions. A 7×7 convolution kernel with 3×3 padding slides between different blocks with a stride of 4×4 and convolves with each block. The convolution results of each block are then arranged into a semantic embedding map for the k-th frame according to the block arrangement. The semantic embedding graph of the k-th frame The input is the Transformer Improvement Block, which replaces the fully connected layers of the multi-head attention mechanism in the existing Transformer structure with 1×1 pointwise convolutions to reduce the number of parameters and further refine semantic information. Simultaneously, it replaces the two fully connected layers and the ReLU function of the existing Transformer structure's feedforward network with two 1×1 pointwise convolutions and the GeLU function, respectively. The GeLU function is a Gaussian-based cumulative function, which, compared to the ReLU function, retains negative information while adaptively weighting the input based on its importance. The output of the Transformer Improvement Block is the semantic feature map of the k-th frame. The semantic feature map of the k-th frame is obtained by bilinear interpolation. The reflection-suppressed assembly map of the kth frame is adjusted The semantic class segmentation is performed by a Softmax function on the reflection-suppressed assembly map of the kth frame The semantic feature map of the kth frame is marked in the middle The pixel points belonging to the assembly position are refined to sharpen the edges, and the enhanced assembly map of the kth frame is generated The assembly position is the assembly area corresponding to the assembly part.

[0067] Further, the compensation parameter sequence is generated The method comprises the following steps:

[0068] The three-dimensional convolution network adopts a three-dimensional convolution kernel with a size of 3x3x3 and a padding of 1 to process the compensation parameter sequence The spatiotemporal feature extraction is performed with a step size of 1, and the spatiotemporal feature map sequence is generated by sequentially passing through batch normalization and a ReLU function The three-dimensional convolution kernel extracts the feature information in the 3x3 local area of adjacent frames in the convolution process, while taking into account the changes in time and space in the compensation parameter sequence

[0069] The pyramid optical flow mapping block constructs an L-layer pyramid according to the reflection-suppressed assembly map sequence The reflection-suppressed assembly map of the kth frame and the k+1th frame, and calculates the forward preliminary optical flow map from the kth frame to the k+1th frame and the backward preliminary optical flow map from the k+1th frame to the kth frame respectively The confidence mask M from the kth frame to the k+1th frame is generated by filtering and cycle consistency detection k,k+1 and the forward optical flow map v from the kth frame to the k+1th frame and the backward optical flow map v from the k+1th frame to the kth frame are obtained by correction k,k+1 k+1,k ;

[0070] The compensation parameter f1 of the 1st frame is set V and the compensation parameter of the Kth frame are respectively equal to the forward optical flow map v from the 1st frame to the 2nd frame and the backward optical flow map v from the K-1th frame to the Kth frame 1,2 K,K-1 When 2≤k≤K-1, the compensation parameter of the kth frame is set is equal to the forward optical flow map v from the kth frame to the k+1th frame, and the compensation parameter sequence is generated k,k+1

[0071] As shown in Figure 2 Further, the processing steps of the pyramid optical flow mapping block include:

[0072] ​​​​The reflection suppression assembly diagram of the kth frame and the k+1th frame is obtained, and L-1 times of downsampling is respectively performed to construct a sampling pyramid of the kth frame and the k+1th frame, wherein the lth downsampling is to generate a (l+1)th layer sampling diagram from the lth layer sampling diagram by using a convolution with a step of 2, and the dimension of the (l+1)th layer sampling diagram is half of the dimension of the lth layer sampling diagram, the sampling diagram includes the sampling diagrams of the kth frame and the k+1th frame, and l=1,…,L-1;

[0073] The forward preliminary optical flow diagram of the kth frame to the k+1th frame at the Lth layer is set and the backward preliminary optical flow diagram is a full 0 diagram and is refined step by step from the Lth layer downwards;

[0074] For the lth layer, 1≤l≤L-1, a search box with the same dimension as the (l+1)th layer convolution diagram of the kth frame is generated at the center of the lth layer convolution diagram of the kth frame and the k+1th frame, respectively, the search box in the lth layer convolution diagram of the k+1th frame is moved according to the forward pixel offset, a cost volume of the search box of the kth frame and the k+1th frame at the lth layer is constructed, the forward pixel offset corresponding to the minimum cost volume is calculated by three-dimensional convolution, and the up-sampling result of the forward preliminary optical flow diagram of the kth frame to the k+1th frame at the (l+1)th layer is superimposed to generate the forward preliminary optical flow diagram of the kth frame to the k+1th frame at the lth layer The cost volume is used to measure the pixel difference in the search box of the kth frame and the k+1th frame, when the cost volume is minimum, the pixels in the search box of the kth frame and the k+1th frame match, and the corresponding forward pixel offset is equal to the pixel movement between the kth frame and the k+1th frame, and the up-sampling is a transpose convolution with a step of 2;

[0075] The search box in the lth layer convolution diagram of the k+1th frame is moved back to the center position, the search box in the lth layer convolution diagram of the kth frame is moved according to the backward pixel offset, and the backward pixel offset corresponding to the minimum cost volume is calculated by three-dimensional convolution, and the up-sampling result of the backward preliminary optical flow diagram of the kth frame to the k+1th frame at the (l+1)th layer is superimposed to generate the backward preliminary optical flow diagram of the kth frame to the k+1th frame at the lth layer

[0076] Until the refinement is stopped at the first layer of the sampling pyramid, the forward preliminary optical flow diagram of the kth frame to the k+1th frame at the first layer and the backward preliminary optical flow diagram are respectively taken as the forward preliminary optical flow diagram of the kth frame to the k+1th frame and the backward preliminary optical flow diagram of the kth frame to the k+1th frame

[0077] As shown in Figure 3 , further, the forward preliminary optical flow diagram of the kth frame to the k+1th frame and backward preliminary optical flow map The anti-illumination assembly map of the kth frame is assembled by weakening the noise influence through bilateral filtering and performing cycle consistency detection According to the forward preliminary optical flow map from the kth frame to the k+1th frame The recursive anti-illumination assembly map of the k+1th frame is generated by forward prediction through bilinear interpolation and according to the backward preliminary optical flow map from the kth frame to the k+1th frame The reconstructed anti-illumination assembly map of the kth frame is generated by backward prediction through bilinear interpolation again Compare the anti-illumination assembly map of the kth frame and the reconstructed anti-illumination assembly map of the kth frame The pixel difference at the same pixel coordinate (μ,ν), if the pixel difference is less than the difference threshold, then the confidence mask M from the kth frame to the k+1th frame k,k+1 The confidence at the pixel coordinate (μ,ν) in the confidence mask M from the kth frame to the k+1th frame is equal to 1, if the pixel difference is greater than or equal to the difference threshold k,k+1 The confidence at the pixel coordinate (μ,ν) in the confidence mask M from the kth frame to the k+1th frame is equal to 0, the confidence reflects the reliability of the optical flow, that is, the reliable forward optical flow map and the backward optical flow map in the adjacent two frames must be mutually offset inverse calculation, the confidence mask M from the kth frame to the k+1th frame k,k+1 respectively multiplied by the forward preliminary optical flow map from the kth frame to the k+1th frame and the backward preliminary optical flow map to screen reliable optical flow, generate the forward optical flow map v from the kth frame to the k+1th frame k,k+1 and the backward optical flow map v k+1,k .

[0078] Specifically, when 1≤k≤K-1, the cross-domain fusion network uses the compensation parameter of the kth frame to determine the enhanced assembly map of the kth frame The corresponding pixel in the enhanced assembly map of the k+1th frame from the enhanced assembly map of the k+1th frame Find the pixel values of the four adjacent pixels around the corresponding pixel in the enhanced assembly map of the k+1th frame Update the pixel value in the enhanced assembly map of the kth frame through inverse distance weighting to weaken the abnormal influence of pixel movement at the position to be assembled, generate the motion compensation assembly map of the kth frame and the reconstructed assembly map of the kth frame Perform weighted summation through a preset weight, generate the clear assembly map of the kth frame Construct a clear assembly map sequence

[0079] Further, the joint deblurring network needs to be pre-trained with the blurrer and the discriminator. The actual clear assembly graph of multiple frames collected in advance is input into the blurrer. The blurrer superimposes random Gaussian noise in the clear assembly graph and sequentially passes through the full connection layer to reduce dimension and the deconvolution layer to increase dimension to simulate the information loss caused by mirror reflection and the position to be assembled. The blurrer generates the blurred assembly graph of multiple frames. The blurred assembly graph of multiple frames is deblurred by the joint deblurring network to generate the clear assembly graph of multiple frames. The discriminator extracts the image features of the clear assembly graph of multiple frames and the corresponding actual clear assembly graph through multiple convolution layers, and outputs the average similarity probability between the clear assembly graph of multiple frames and the corresponding actual clear assembly graph through the Sigmoid function. The parameters of the blurrer, the joint deblurring network and the discriminator are updated constantly through the adversarial training idea until the average similarity probability is maximized.

[0080] As shown in Figure 4 , specifically, the pose mapping network adopts a traditional pre-training method, which will not be described in detail. The pose mapping network includes a residual network, an encoding converter and a regression network.

[0081] The residual network convolves the clear assembly graph of the kth frame by using a convolution kernel with a size of 3x3 and a padding of 1 with a step of 1 to extract the global feature map of the kth frame The residual network reduces the dimension of the global feature map of the kth frame by maximum pooling, and restores the initial dimension of the global feature map of the kth frame by 1x1 pointwise convolution again, and performs residual superposition with the global feature map of the kth frame to generate the feature map F k of the kth frame. The annotation filtering based on the joint deblurring network filters out the features in the feature map F k of the kth frame that do not belong to the position to be detected, and generates the target feature map F

[0082] The encoding converter disassembles the camera extrinsic parameter matrix T to generate a rotation matrix R and a translation vector a. The rotation matrix R with a dimension of 3x3 is converted into a quaternion vector q with a dimension of 1x4 by using quaternion conversion. The quaternion conversion is as follows:

[0083]

[0084]

[0085] wherein q(i) is the i-th quaternion in the quaternion vector q, i∈{1,2,3,4}, length(q) is the module length of the quaternion vector q, R(1,1), R(2,2) and R(3,3) are the 1st, 2nd and 3rd values on the diagonal line of the rotation matrix R, respectively.

[0086] The regression network maps the target feature map of the kth frame is flattened into a target feature vector of the kth frame is concatenated with the quaternion vector q and the translation vector a in column and input into the multilayer perceptron, the multilayer perceptron is processed through two consecutive linear layers and batch normalization layers, and then is nonlinearly mapped based on the ReLU function, and outputs the target pose of the kth frame

[0087] The application discloses a visual positioning method for mobile phone part assembly, introduces a reflection correction and motion compensation mechanism to solve the technical problems of traditional visual positioning, constructs a reflection removal model based on Stokes vector method and polarization fusion network cooperative operation, can accurately identify and suppress the to-be-assembled area in the assembly image and reconstruct and generate a reflection suppression assembly image, thereby reducing the interference of reflection; a joint deblurring network is used to perform edge enhancement, motion modeling and filtering on the reflection suppression assembly image in the spatial domain, time domain and frequency domain respectively, and based on the processing results of the spatial domain, time domain and frequency domain, the image blur caused by part motion in the assembly process is dynamically compensated, a clear assembly image is provided for subsequent capture of the motion trajectory of the to-be-assembled position and combination of the time dimension target pose prediction, the dynamic error of positioning is effectively reduced, and high-precision visual positioning can be realized under the to-be-assembled mobile phone processing motion state.

[0088] The above only describes the preferred embodiments of the application, and the protection scope of the application is not limited to the above-described embodiments, and any technical scheme falling within the idea of the application belongs to the protection scope of the application. It should be noted that, for ordinary skilled persons in the art, some improvements and refinements without departing from the principles of the application are also considered to be within the protection scope of the application.

Claims

1. A visual positioning method for assembling parts of a mobile phone, characterized by, The method comprises the following steps: determining the parts to be assembled, acquiring a stack of assembly images at polarization angles of 0°, 45° and 90° respectively; calculating a polarization degree image and a polarization angle image of each frame based on the assembly images at the three polarization angles by using a Stokes vector method, fusing the polarization degree image and the polarization angle image of each frame into a weight mask by using a polarization fusion network, and performing convolution on the splicing result of the assembly images at the three polarization angles to generate a reflection-suppressed assembly image of each frame, thereby constructing a sequence of reflection-suppressed assembly images; using a semantic segmentation network to label the position to be assembled in the reflection-suppressed assembly image of each frame to generate an enhanced assembly image of each frame by using a joint deblurring network, extracting the spatiotemporal features in the sequence of reflection-suppressed assembly images by using a three-dimensional convolution network, and generating a sequence of compensation parameters in combination with an optical flow field, sequentially passing the reflection-suppressed assembly image of each frame through Fourier transform, band-pass filtering and inverse Fourier transform to generate a reconstructed assembly image of each frame, and fusing the enhanced assembly image of each frame with the reconstructed assembly image of each frame to generate a clear assembly image of each frame by using a cross-domain fusion network based on the compensation parameter of each frame.

2. The visual positioning method for assembling parts of a mobile phone as claimed in claim 1, wherein, The Stokes vector method comprises the following steps: filtering the assembly images at the polarization angles of 0°, 45° and 90° respectively to generate corresponding denoised assembly images; summing the denoised assembly images at the polarization angles of 0° and 90° to generate a total light intensity image of each frame; taking the difference between the denoised assembly images at the polarization angles of 0° and 90° as a polarization image of each frame; subtracting the denoised assembly images at the polarization angles of 0° and 90° from twice the denoised assembly image at the polarization angle of 45° to generate a diagonal polarization image of each frame; taking the square root of the sum of the polarization image and the diagonal polarization image of each frame and dividing by the total light intensity image of each frame to generate a polarization degree image of each frame; taking half of the arctangent value of the ratio between the diagonal polarization image and the polarization image of each frame as a polarization angle image of each frame.

3. The visual positioning method for assembling parts of a mobile phone as claimed in claim 1 wherein, The semantic segmentation network divides the reflection-suppressed assembly image of each frame into small blocks and ensures that there is a pixel overlap region between adjacent small blocks, performs convolution in each small block, arranges the convolution results into a semantic embedding image of each frame, further inputs a Transformer improved block to generate a semantic feature image of each frame, adjusts the dimension of the semantic feature image of each frame by using bilinear interpolation and performs semantic class segmentation by using a Softmax function to identify the pixel points belonging to the position to be assembled in the reflection-suppressed assembly image of each frame, thereby generating an enhanced assembly image of each frame, wherein the Transformer improved block replaces the full connection layer and the ReLU function in the existing Transformer structure with a point-wise convolution and a GeLU function respectively.

4. The visual positioning method for assembling parts of a mobile phone as claimed in claim 1 wherein, The method for generating the sequence of compensation parameters comprises the following steps: The three-dimensional convolution network extracts spatiotemporal features of the sequence of compensation parameters by using a three-dimensional convolution kernel, and generates a sequence of spatiotemporal feature maps by mapping through batch normalization and a ReLU function. The pyramid optical flow mapping block constructs a pyramid of L layers according to the anti-reflection assembly diagram of the kth frame and the k+1th frame of the anti-reflection assembly diagram sequence, and calculates a forward preliminary optical flow diagram and a backward preliminary optical flow diagram from the kth frame to the k+1th frame respectively, generates a confidence mask from the kth frame to the k+1th frame through filtering and cycle consistency detection, and corrects to obtain a forward optical flow diagram and a backward optical flow diagram from the kth frame to the k+1th frame; The compensation parameters of the 1st frame and the Kth frame are set to be equal to the forward optical flow diagram from the 1st frame to the 2nd frame and the backward optical flow diagram from the K-1th frame to the Kth frame respectively, when 2≤k≤K-1, the compensation parameter of the kth frame is set to be equal to the forward optical flow diagram from the kth frame to the k+1th frame, and a compensation parameter sequence is generated.

5. The visual positioning method for assembling parts of a handset as claimed in claim 4 wherein, The processing steps of the pyramid optical flow mapping block include: Obtaining the anti-reflection assembly diagrams of the kth frame and the k+1th frame, and performing L-1 times of downsampling respectively to construct the sampling pyramid of the kth frame and the k+1th frame; Setting the forward preliminary optical flow diagram and the backward preliminary optical flow diagram of the kth frame to the k+1th frame at the Lth layer to be all 0 diagrams and refining step by step from the Lth layer downward; For the lth layer, 1≤l≤L-1, a search box with the same dimension as the convolution diagram of the k+1th frame at the l+1th layer is generated at the center of the convolution diagram of the kth frame at the lth layer and the k+1th frame at the lth layer respectively, the search box is moved in the convolution diagram of the k+1th frame at the lth layer, and the forward offset corresponding to the minimum cost volume is determined, and the up-sampling result of the forward preliminary optical flow diagram of the kth frame to the k+1th frame at the l+1th layer is superimposed to generate the forward preliminary optical flow diagram of the kth frame to the k+1th frame at the lth layer; The search box in the convolution diagram of the k+1th frame at the lth layer is moved back to the center position, the search box is moved in the convolution diagram of the kth frame at the lth layer, and the backward pixel offset corresponding to the minimum cost volume is determined, and the up-sampling result of the backward preliminary optical flow diagram of the kth frame to the k+1th frame at the l+1th layer is superimposed to generate the backward preliminary optical flow diagram of the kth frame to the k+1th frame at the lth layer; After refining to the 1st layer, stop, and take the forward preliminary optical flow diagram and the backward preliminary optical flow diagram of the kth frame to the k+1th frame at the 1st layer as the forward preliminary optical flow diagram and the backward preliminary optical flow diagram of the kth frame to the k+1th frame respectively.

6. The visual positioning method for assembling parts of a mobile phone as claimed in claim 4, wherein, The forward preliminary optical flow diagram and the backward preliminary optical flow diagram from the kth frame to the k+1th frame are detected through filtering and cycle consistency detection, the anti-reflection assembly diagram of the kth frame is forward predicted according to the forward preliminary optical flow diagram from the kth frame to the k+1th frame and backward predicted again according to the backward preliminary optical flow diagram from the kth frame to the k+1th frame, a reconstructed anti-reflection assembly diagram of the kth frame is restored and generated, and the pixel difference at the same pixel coordinate is compared with the anti-reflection assembly diagram of the kth frame, if the pixel difference is less than a difference threshold, the confidence of the pixel coordinate in the confidence mask from the kth frame to the k+1th frame is equal to 1, otherwise it is equal to 0, the confidence mask from the kth frame to the k+1th frame is multiplied by the forward preliminary optical flow diagram and the backward preliminary optical flow diagram of the kth frame to the k+1th frame respectively, and the forward optical flow diagram and the backward optical flow diagram from the kth frame to the k+1th frame are generated.

7. The visual positioning method for assembling parts of a mobile phone as claimed in claim 1 wherein, The cross-domain fusion network determines pixels in the enhanced assembled image of the kth frame and corresponding pixels in the enhanced assembled image of the k+1th frame using a compensation parameter of the kth frame, finds pixel values of neighboring pixels around the corresponding pixels from the enhanced assembled image of the k+1th frame, and updates pixel values in the enhanced assembled image of the kth frame by weighting, generates a motion compensation assembled image of the kth frame, and performs weighted summation of the motion compensation assembled image of the kth frame and a reconstructed assembled image of the kth frame through a preset weight, generates a clear assembled image of the kth frame, and constructs a clear assembled image sequence.

8. The method for visual positioning of mobile phone parts assembly as claimed in any one of the claims 1 to 7 wherein, The visual positioning method further comprises the following steps: The clear assembled image of each frame is input into a pose mapping network together with a camera extrinsic matrix, target feature maps of each frame in the clear assembled image of each frame are extracted and flattened, and the flattened target feature maps are spliced with an encoding result of the camera extrinsic matrix to input a multi-layer perceptron, so as to output a target pose of each frame and construct a target pose sequence; A completed frame is determined, the target pose sequence is input into a time sequence network to generate a predicted target pose of the completed frame, and an intelligent control algorithm is used to control an execution device to adjust the to-be-assembled part to the predicted target pose at the completed frame and assemble the to-be-assembled part.

9. The method for visual positioning of cell phone part assembly of any one of claim 8, wherein, The pose mapping network comprises a residual network, an encoding converter, and a regression network; The residual network performs convolution on the clear assembled image of each frame to generate a global feature map of each frame, superimposes a result of maximum pooling and point-by-point convolution on the global feature map of each frame with a residual error of the global feature map of each frame, generates a feature map of each frame, and removes features in the feature map of each frame that do not belong to a to-be-detected position based on a joint deblurring network to generate a target feature map of each frame; The encoding converter disassembles the camera extrinsic matrix to generate a rotation matrix and a translation vector, and converts the rotation matrix into a quaternion vector using a quaternion; The regression network flattens the target feature map of each frame and splices the flattened target feature map with the quaternion vector and the translation vector to input a multi-layer perceptron, and generates a target pose of each frame through two consecutive linear layers, batch normalization layers, and ReLU functions.

10. The visual positioning method for assembling parts of a handset as claimed in claim 1 wherein, The polarization fusion network comprises a polarization encoding layer and a channel fusion layer; The polarization encoding layer fuses the polarization degree map and the polarization angle map of each frame using a linear layer of a multi-layer perceptron and generates a polarization feature map of each frame based on a ReLU function mapping; The channel fusion layer normalizes the polarization feature map of each frame into a weight mask of each frame using a Softmax function, and performs channel convolution on a splicing result of the assembled image of each frame in the channel dimension under three polarization angles to generate an anti-glare assembled image of each frame.

Citation Information

Patent Citations

  • Visual image positioning method for modular assembly of mobile phone housing

    CN117372528B

  • Video super-resolution reconstruction method and system

    CN120013766A

  • Polarization image fusion method and system based on multi-scale brightness perception

    CN120235775A