A dynamic tracking optical scanning system and method

Through the dynamic tracking optical scanning system, image stitching and three-dimensional reconstruction technology are used to solve the problem of difficulty in dealing with the position changes of the inspected object in real time in the prior art, and high-precision position estimation and three-dimensional model reconstruction are achieved.

CN119515982BActive Publication Date: 2025-05-27SHANGHAI MOGOAI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510067453.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-27
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The existing three-dimensional scanning technology is difficult to deal with the pose changes of the object being inspected in real time, and depends on visual images or depth information, and is easily challenged by poor lighting conditions, low contrast and scene occlusion, resulting in deviations in pose estimation.

Method used

The dynamic tracking optical scanning system is adopted, including a central control subsystem, multiple tracking devices and scanning equipment, and the field of view of the main tracking device is expanded through the image stitching module. The three-dimensional reconstruction module generates sparse three-dimensional models and local dense three-dimensional models based on stereo matching and semantic segmentation algorithms, and accurately obtains the scanning external parameter matrix through the dual feature pose estimation network.

Benefits of technology

Real-time tracking and precise position estimation of the position changes of the inspected object are realized, and the matching accuracy and production efficiency of the three-dimensional scan data are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119515982B_ABST
    Figure CN119515982B_ABST
Patent Text Reader

Abstract

The present invention discloses a dynamic tracking optical scanning system and method, which relates to the field of three-dimensional scanning technology. The central control subsystem controls the tracking device and the scanning device to complete three-dimensional model reconstruction. The central control subsystem includes a calibration and correction module, an image stitching module, a three-dimensional reconstruction module, and a display module; the calibration and correction module feeds back correction instructions through the binocular calibration and correction method; the image stitching module stitches all binocular correction image sequences based on the transformation stitching algorithm to generate a global image; the three-dimensional reconstruction module combines the stereo matching algorithm and the semantic segmentation annotation method to generate an object disparity map and a target disparity map, and uses the reprojection transformation method to obtain the sparse three-dimensional model of the current time frame, and combines the dual feature pose estimation network and the coordinate transformation method to generate the local dense three-dimensional model of the current time frame; the display module performs real-time display based on the pose detection and updated display strategy, realizing the synchronous dynamic tracking of the pose transformation of the object to be inspected and the scanning device during the scanning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional scanning technology, and in particular to a dynamic tracking optical scanning system and method. Background Art

[0002] With the development of modern manufacturing, the complexity of workpieces and working environments has increased day by day, which has led to the widespread application of 3D scanning systems in the manufacturing industry. With the help of advanced 3D scanning control systems, it is easy to achieve precise scanning of a variety of workpiece objects. In the current technical system, 3D scanning often relies on traditional three-dimensional coordinate measuring machines, which use probes to measure various hole positions and complex surface features on workpieces. However, the use and production efficiency of three-dimensional coordinate measuring machines are low and the cost is high.

[0003] The existing invention patent with publication number CN117870543A proposes a tracking scanning system and a tracking scanning processing method, including: a scanning device, a tracking device and a control device; the scanning device is used to scan the workpiece to be inspected and send the scanning information to the control device; the tracking device is used to track the scanning device and send the tracking information to the control device; the control device is used to perform data synchronization processing on the scanning information and the tracking information in real time, and determine the coordinate information of the workpiece to be inspected based on the result of the synchronization processing, so as to realize the tracking scanning of the workpiece to be inspected, thereby improving the use and production efficiency.

[0004] Since the field of view provided by a single tracking device is limited, the existing technology often only tracks and determines the posture of the scanning device and restores the three-dimensional model of the inspected object based on the scanning data. This makes it impossible to respond to the posture changes of the inspected object in a timely manner, making it impossible for the real-time scanning data to match the previously generated local three-dimensional model in a timely manner. At the same time, the existing technology only uses visual images or depth information to determine the posture of the scanning device. The use of visual images will be challenged by poor lighting conditions, low contrast and scene occlusion. Using only depth information faces the problem of difficulty in data structuring processing, which makes the posture estimation biased. Summary of the invention

[0005] In view of the deficiencies in the prior art, the present invention proposes a dynamic tracking optical scanning system and method to provide a dynamic tracking scanning based on an extended field of view, which can effectively cope with the posture changes of the inspected object during the scanning process and provide more accurate posture estimation based on visual and geometric features.

[0006] The technical solution to achieve the purpose of the present invention is:

[0007] A dynamic tracking optical scanning system, comprising a central control subsystem, a plurality of tracking devices and a scanning device;

[0008] The central control subsystem receives the start command and controls the tracking device to perform calibration correction, receives the scanning start command and controls the tracking device and the scanning device to perform equal-interval image acquisition and scanning respectively, and analyzes and processes all binocular correction image sequences and local point cloud data in the current time frame. , restore and generate a sparse 3D model and a local dense 3D model and display them;

[0009] The multiple tracking devices include a master tracking device and a slave tracking device, each tracking device selects to perform calibration correction and image acquisition based on a command-driven strategy;

[0010] The scanning device is equipped with a tracking target, receives the frame synchronization instruction of the current time frame, and obtains the local point cloud data of the current time frame And feed back to the central control subsystem.

[0011] Furthermore, the central control subsystem includes a signal generation module, a frame synchronization module, a calibration correction module, an image stitching module, a three-dimensional reconstruction module and a display module;

[0012] The signal generation module receives the start instruction, sends an initialization signal to the calibration correction module, receives the scan start instruction, and sends a scan start signal to the frame synchronization module;

[0013] The frame synchronization module receives the scanning start signal and broadcasts the frame synchronization instruction, sets the time frames at equal intervals and broadcasts the frame synchronization instruction in each time frame;

[0014] The calibration and correction module receives the initialization signal and sends the initialization command to all tracking devices, generates the correction command of a single tracking device through the dual-target calibration method and feeds it back, calculates the main fixed parameter table and sends it to the 3D reconstruction module;

[0015] The image stitching module receives all binocular correction image sequences of the current time frame and generates a left-eye global image based on the transformation stitching algorithm. and the right eye global image And send it to the 3D reconstruction module;

[0016] The 3D reconstruction module receives and stores the main fixed parameter table and analyzes the left eye global image based on the stereo matching algorithm. and the right eye global image Generate global disparity map , using semantic segmentation and annotation to generate object disparity maps and target disparity map , the sparse 3D model of the current time frame is obtained by reprojection transformation method and sent to the display module, and the scanning external parameter matrix of the current time frame is obtained through the dual feature pose estimation network , based on the coordinate transformation method, the local point cloud data of the current time frame Convert it into a local dense three-dimensional model and send it to the display module;

[0017] The display module updates the display strategy based on pose detection to display the sparse 3D model of the current time frame and all local dense 3D models.

[0018] Furthermore, the calibration correction module generates a calibration instruction for a single tracking device through a dual-target calibration method, including the following specific steps:

[0019] The Zhang Zhengyou calibration method was used to obtain the Tracking devices The left camera external parameters, right camera external parameters, left eye center pixel coordinates 、Right eye center pixel coordinates and camera focal length , the left camera external parameters include the left rotation matrix and the left eye translation vector , the right camera external parameters include the right rotation matrix and the right eye translation vector , To track the device number, , To track the total number of devices;

[0020] Combine the left camera external parameters and the right camera external parameters to establish the world coordinates With the left camera coordinates and the right camera coordinates The coordinate transformation equation is derived by the elimination method, and the self-rotation matrix is ​​obtained based on the coordinate transformation equation. and the self-translation vector ;

[0021] The rotation matrix Partition into semi-positive rotation matrix and the semi-inverse rotation matrix , will Tracking devices The pole of is translated to the infinity of the world coordinate system, generating the self-translation vector Transformation matrix of visual lines in the same direction ;

[0022] Combined with the semi-positive rotation matrix , semi-inverse rotation matrix and the visual line transformation matrix Get the left eye correction matrix separately and the right eye correction matrix , construct the correction instruction And send to Tracking devices .

[0023] Furthermore, the calibration and correction module generates the main fixed parameter table including the following specific steps:

[0024] Temporary storage of master tracking devices The pixel coordinates of the left eye center obtained by Zhang Zhengyou's calibration method , left eye translation vector 、Right eye center pixel coordinates , right eye translation vector , and the camera focal length ;

[0025] Combined correction instructions , left eye translation vector and the right eye translation vector Get the corrected self-translation vector of Axis Value ;

[0026] The left camera is used as the reference camera, based on the left eye center pixel coordinates 、Right eye center pixel coordinates , Camera focal length and the corrected self-translation vector Compute reference camera pose Distance from baseline ;

[0027] Packing reference camera poses , Camera focal length Distance from baseline , generate the main fixed parameter table.

[0028] Furthermore, the image stitching module generates the left eye global image based on the transformation stitching algorithm and the right eye global image The specific steps include:

[0029] Select the unconnected slave tracking device with the smallest tracking device number in order, assuming it is the Slave tracking devices , the corrected image As the image to be stitched, , , To track the total number of devices;

[0030] Before acquisition The first splicing generated Sub-global image , using Harris corner detection algorithm to pair the Sub-global image With the correction image Get the corner point in Secondary paired corner point set;

[0031] The random sampling iterative algorithm is used to process the The second pairing corner point set generates the Subhomography matrix , and according to Subhomography matrix Correct the image Transformed into auxiliary correction image ;

[0032] Obtain auxiliary correction images of all slave tracking devices, and stitch the auxiliary correction images into the correction image in sequence using the hat weighted average fusion algorithm according to the tracking device number sequence. , generate a global image .

[0033] Furthermore, a random sampling iterative algorithm is used to process the The second pairing corner point set generates the Subhomography matrix The specific steps include:

[0034] Perform a random draw estimate, starting from Four paired corner points are randomly selected from the second paired corner point set to construct the The first internal point set is solved by the least squares method to obtain the homography equations Subhomography matrix ;

[0035] Adopt the Subhomography matrix Go to test The remaining paired corner points in the second paired corner point set are calculated, and the estimation errors of all the remaining paired corner points are calculated. The paired corner points with estimation errors less than the estimation threshold are placed in the second paired corner point set. Secondary interior point set;

[0036] Judgement Is the total number of paired corner points in the second inner point set greater than the corner point number threshold? If so, output the number of paired corner points in this round of iteration. Subhomography matrix , if it is less than or equal to, obtain the number of iteration rounds and determine whether the number of iteration rounds is greater than the iteration round number threshold;

[0037] If it is less than or equal to, the first All paired corner points in the second internal point set are put back into the Pair the corner point set again, jump again to perform random extraction estimation, if it is greater than, output the first round of iterative calculation Subhomography matrix .

[0038] Furthermore, a hat-weighted average fusion algorithm is used to generate the left eye global image and the right eye global image The specific steps include:

[0039] Defining the calibration image The first global image , assuming that the Sub-global image , For the Sub-global image The level of vision, ;

[0040] Confirm Sub-global image With auxiliary correction image The overlapping area and all overlapping pixels;

[0041] For Overlapping pixels, The number of overlapping pixels Sub-global image weights The first The ratio of the sum of the sub-global image weights is taken as the Sub-global image weights , will The auxiliary correction image weights of overlapping pixels The ratio of the total weight of the auxiliary correction image of all overlapping pixels is taken as the auxiliary correction image weight , , is the total number of overlapping pixels;

[0042] Get the sub-global image weight vector and auxiliary correction image weight vector , will Sub-global image Non-overlapping areas, auxiliary correction images The non-overlapping area and the weighted sum of all overlapping pixels are concatenated to generate the Sub-global image ;

[0043] Judgement Sub-global image Vision level Is it equal to the total number of tracked devices? , if equal, then Sub-global image The global image ;

[0044] If it is less than, continue splicing Auxiliary correction image of slave tracking device , until the global image is obtained .

[0045] Furthermore, the 3D reconstruction module includes a disparity acquisition unit, a semantic segmentation unit, a sparse reconstruction unit, a pose estimation unit, a dense reconstruction unit and a storage unit;

[0046] The disparity acquisition unit compares the left eye global image based on the stereo matching algorithm and the right eye global image The image patches of the same size in the image are used to generate the global disparity map , the left global image and global disparity map Send to the semantic segmentation unit;

[0047] The semantic segmentation unit uses semantic segmentation annotation method to analyze and process the left eye global image and global disparity map , generate object disparity map and target disparity map , send the left eye global image and object disparity map To the sparse reconstruction unit, send the left eye global image and target disparity map To the pose estimation unit;

[0048] The sparse reconstruction unit uses the reprojection transformation method combined with the left global image and object disparity map Obtain the sparse three-dimensional model of the current time frame and send it to the display module;

[0049] The pose estimation unit uses a dual feature pose estimation network to process the left eye global image and target disparity map , get the scan extrinsic parameter matrix of the current time frame And sent to the dense reconstruction unit;

[0050] The dense reconstruction unit uses the coordinate transformation method combined with the scanned extrinsic matrix of the current time frame and local point cloud data , generate a local dense 3D model of the current time frame and send it to the display module;

[0051] The storage unit receives and stores the main fixed parameter table, and responds to the calls of the sparse reconstruction unit and the pose estimation unit.

[0052] Furthermore, the semantic segmentation unit uses semantic segmentation annotation to generate object disparity maps and target disparity map The specific steps include:

[0053] Normalized left eye global image , generating global feature maps through residual neural network mining ;

[0054] is the global feature map Generate a candidate box for each feature point in the image, perform binary classification through the region selection network, and delete the candidate boxes whose feature points completely belong to the background;

[0055] The left eye global image is aligned using the region of interest method Pixel points and global feature map Align the feature points and candidate boxes;

[0056] The retained candidate boxes are combined with the global feature map The feature points in the candidate frame are input into the classifier and the mask marker respectively. The classifier uses a combination of a fully connected layer and a Softmax function to classify the feature points in the candidate frame as belonging to the object being inspected or the tracking target. The mask marker generates a marking mask corresponding to the candidate frame through two convolutional neural networks.

[0057] Combine the classifier and mask marker to get the global image of the left eye. Add object tags and target tags to each pixel point, and process the global disparity map based on the object tags and target tags respectively. , generate object disparity map and target disparity map .

[0058] Furthermore, the sparse reconstruction unit uses a reprojection transformation method to obtain a sparse three-dimensional model, including the following specific steps:

[0059] Call the main fixed parameter table and extract the reference camera pose ;

[0060] Based on the reference camera pose Create a global image of the left eye The pixel coordinates of each pixel in With world coordinates The reprojection equation between ;

[0061] Obtaining the left eye global image based on the reprojection equation The world coordinates corresponding to all pixels in are used to reconstruct the sparse 3D model of the current time frame using the greedy triangulation method.

[0062] Furthermore, the pose estimation unit uses a dual feature pose estimation network to obtain the scan extrinsic parameter matrix of the current time frame The specific steps include:

[0063] Call the main fixed parameter table and extract the camera focal length Distance from baseline , based on the principle of triangulation, the target disparity map Transformation generates target depth map ;

[0064] The left global image is convolutionally processed through a convolutional layer with a convolution kernel size of 7×7, a convolution step size of 2, and a padding of 3. Perform feature extraction and further generate visual feature maps through dimension reduction through normalization layer and ReLU activation function ;

[0065] The target depth map is obtained through the point cloud network Converted into geometric point data, nonlinearly transformed through a multi-layer perceptron, and generated a target geometric feature map through linear layer dimensionality reduction ;

[0066] Convolutional feature fusion module for visual feature maps and target geometry Perform 3×3 convolution and point multiplication to achieve fusion transformation, and then perform 3×3 convolution again and add them to the visual feature map respectively. and target geometry In the above example, the first fusion feature map is generated. and the second fusion feature map ;

[0067] The first fusion feature map is processed using the channel attention mechanism combined with average pooling , the second fusion feature map is processed by the channel attention mechanism combined with the maximum pooling The output results of the two channels are accumulated by element-wise addition, and processed through a one-dimensional convolution layer, a normalization layer, and a ReLU activation function to generate a dual fusion feature map ;

[0068] The dual fusion feature map Further input 1 maximum pooling layer and 4 fully connected layers for regression, obtain the 6-dimensional target pose sequence and construct the scanning external parameter matrix ,in, and are the scanning rotation matrix and scanning translation vector respectively.

[0069] Furthermore, the dense reconstruction unit generates a local dense 3D model of the current time frame using a coordinate transformation method, including the following specific steps:

[0070] Get the scan extrinsic matrix of the current time frame ;

[0071] Scanning extrinsic matrix based on the current time frame Build coordinate transformation equations to relate local point cloud data The scanning device coordinates of each point cloud in The corresponding world coordinates ;

[0072] Obtain local point cloud data based on coordinate transformation equation derivation The world coordinates corresponding to all point clouds in ;

[0073] The local dense 3D model of the current time frame is reconstructed based on the greedy triangulation method.

[0074] Furthermore, the display module updates the display strategy based on the pose detection to display the sparse 3D model of the current time frame and all local dense 3D models, including the following specific steps:

[0075] Receive the sparse 3D model and the local dense 3D model of the current time frame and attach color labels to determine whether there is a sparse 3D model of the most recent frame;

[0076] If not, the sparse 3D model and the local dense 3D model of the current time frame are stored and displayed according to the color label;

[0077] If it exists, the sparse 3D model of the most recent frame is called, and the key points of the sparse 3D model of the current time frame and the sparse 3D model of the most recent frame are matched by the scale-invariant feature transformation method and feature matching algorithm;

[0078] Calculate the average Euclidean distance of all key points and determine whether the average Euclidean distance is greater than the Euclidean distance threshold;

[0079] If it is less than or equal to, the sparse 3D model of the current time frame replaces the sparse 3D model of the most recent frame and stores the sparse 3D model of the current time frame, and displays the sparse 3D model of the current time frame and all local dense 3D models according to color labels;

[0080] If it is greater than that, construct 6 key point pairs and calculate the object rotation matrix and object translation vector from the previous time frame to the current time frame by the least squares method;

[0081] All stored historical local dense 3D models are rotated and translated accordingly according to the object rotation matrix and the object translation vector. The sparse 3D model of the current time frame replaces the sparse 3D model of the most recent frame and stores the sparse 3D model of the current time frame. The sparse 3D model of the current time frame and all local dense 3D models are displayed according to color labels.

[0082] Furthermore, tracking devices Taiwan, respectively , where the first tracking device The main tracking device, the rest Each tracking device acts as a slave tracking device, and each tracking device selects to perform calibration correction and image acquisition based on the command-driven strategy.

[0083] Furthermore, a single tracking device may perform calibration and image acquisition based on a command-driven strategy including:

[0084] If the received command is an initialization command, the calibration and correction of a single tracking device is performed;

[0085] If the received instruction is a frame synchronization instruction, the image acquisition of a single tracking device is executed, the binocular correction image sequence is obtained and fed back to the central control subsystem.

[0086] Furthermore, the calibration of a single tracking device includes the following specific steps:

[0087] First Tracking devices For example, , To track the total number of devices, capture a sequence of binocular positioning images And fed back to the calibration correction module, where: and They are the left target positioning image and the right target positioning image respectively;

[0088] Receive the Tracking devices Correction instructions , calibrate the left and right cameras of a single tracking device and fix their poses.

[0089] Furthermore, the scanning device includes a scanner and a tracking target, and the scanner receives a frame synchronization instruction to obtain local point cloud data of the current time frame. And feed back to the central control subsystem.

[0090] A dynamic tracking optical scanning method is implemented based on a dynamic tracking optical scanning system, and includes the following specific steps:

[0091] S1, receiving a start command and sending an initialization command to all tracking devices, obtaining a binocular positioning image sequence of a single tracking device, generating a correction command through a binocular positioning correction method and feeding it back to the single tracking device, storing a main fixed parameter table, and receiving a scan start command;

[0092] S2, broadcast the frame synchronization instruction in the current time frame and set the next time frame;

[0093] S3: Obtain the binocular correction image sequence of all tracking devices in the current time frame, and generate the left eye global image based on the transformation stitching algorithm and the right eye global image ;

[0094] S4. Analyze the left eye global image based on stereo matching algorithm and the right eye global image Generate global disparity map , semantic segmentation and annotation method is used to further process the global disparity map Generate object disparity map and target disparity map , obtain the sparse 3D model of the current time frame through the reprojection transformation method;

[0095] S5. Obtain the scan extrinsic parameter matrix of the current time frame through the dual feature pose estimation network , receive the local point cloud data of the current time frame , converted into a local dense three-dimensional model of the current time frame through the coordinate transformation method;

[0096] S6, updating the display strategy based on the pose detection to display the sparse 3D model of the current time frame and all local dense 3D models;

[0097] S7, waiting for the system time to advance to the next time frame, and jumping to step S2, until the scanning of the inspected object is completed.

[0098] Compared with the prior art, the present invention has the following significant advantages:

[0099] 1. Set up a main tracking device and multiple slave tracking devices, design an image stitching module specifically, transform the binocular correction image sequence collected by all slave tracking devices into an auxiliary correction image with the same viewing angle as the main tracking device based on the transformation stitching algorithm, and stitch together to generate the left eye global image and the right eye global image to achieve the field of view expansion of the main tracking device, so that the tracking device can collect the complete image of the inspected object and the scanning device, laying a good foundation for posture determination;

[0100] 2. Design a 3D reconstruction module and a display module. The 3D reconstruction module compares the left-eye global image and the right-eye global image based on a stereo matching algorithm to obtain a global disparity map. The semantic segmentation and annotation method is used to add labels to the left-eye global image to distinguish the background, the inspected object and the tracking target in the left-eye global image, and further adjusts and generates the object disparity map and the target disparity map. The reprojection transformation method is used to obtain the sparse 3D model of the current time frame. The dual feature pose estimation network is used to combine the visual features of the left-eye global image and the geometric features of the target disparity map to accurately obtain the scanning extrinsic parameter matrix of the scanning device corresponding to the current time frame. The local point cloud data of the current time frame received is converted into a local dense 3D model based on the coordinate transformation method. The display strategy is updated based on the pose detection to analyze in real time whether the pose of the inspected object has changed, and adaptively adjust to correctly display the sparse 3D model of the current time frame and all local dense 3D models, so as to realize the synchronous dynamic tracking scanning of the inspected object and the scanning device. BRIEF DESCRIPTION OF THE DRAWINGS

[0101] Figure 1 A model diagram of a dynamic tracking optical scanning system in the present invention;

[0102] Figure 2 This is a flow chart of the random sampling iterative algorithm in the present invention;

[0103] Figure 3 This is a dual feature pose estimation network model diagram in the present invention;

[0104] Figure 4 The figure is a flow chart of a dynamic tracking optical scanning method in the present invention. DETAILED DESCRIPTION

[0105] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0106] Example 1

[0107] like Figure 1 As shown, a specific embodiment of the present invention discloses a dynamic tracking optical scanning system, including a central control subsystem, a plurality of tracking devices and a scanning device;

[0108] The central control subsystem controls all tracking devices to perform calibration and correction when receiving the start command, broadcasts frame synchronization commands at equal intervals after receiving the scanning start command, controls all tracking devices to perform image acquisition and controls scanning devices to perform scanning, and analyzes and processes the binocular correction image sequences collected by all tracking devices in the current time frame and the local point cloud data obtained by scanning devices. , restore and generate a sparse 3D model and a local dense 3D model of the inspected object and display them;

[0109] The multiple tracking devices include a master tracking device and multiple slave tracking devices, each tracking device performs calibration correction and image acquisition based on the instruction-driven strategy selection;

[0110] The scanning device is equipped with a tracking target, receives the frame synchronization instruction of the current time frame and activates the scanning function to obtain the local point cloud data of the current time frame And feed back to the central control subsystem.

[0111] Furthermore, the central control subsystem includes a signal generation module, a frame synchronization module, a calibration correction module, an image stitching module, a three-dimensional reconstruction module and a display module;

[0112] The signal generation module receives a start instruction from the outside, sends an initialization signal to the calibration correction module, receives a scan start instruction from the outside, and sends a scan start signal to the frame synchronization module;

[0113] The frame synchronization module receives the scanning start signal and immediately broadcasts a frame synchronization instruction. During the operation of the system, the time frame is continuously set according to the frame interval, and the frame synchronization instruction is broadcast once in each time frame.

[0114] The calibration correction module receives the initialization signal and sends the initialization command to all tracking devices. It processes the dual-target calibration image sequence fed back by a single tracking device through the dual-target calibration method, generates the calibration command of a single tracking device and feeds it back to the single tracking device, and obtains the master tracking device. The main fixed parameter table is sent to the 3D reconstruction module;

[0115] The image stitching module receives the binocular correction image sequence collected by a single tracking device in the current time frame and generates a left-eye global image based on the transformation stitching algorithm. and the right eye global image And send it to the 3D reconstruction module;

[0116] The 3D reconstruction module receives and stores the main fixed parameter table and receives the left eye global image. and the right eye global image , obtain the global disparity map based on the stereo matching algorithm , using semantic segmentation and annotation to generate object disparity maps and target disparity map , the reprojection transformation method is used to obtain the sparse 3D model of the current time frame and send it to the display module, and the dual feature pose estimation network is used to obtain the scanning extrinsic parameter matrix corresponding to the current time frame , based on the coordinate transformation method, the local point cloud data of the current time frame received Convert it into a local dense three-dimensional model and send it to the display module;

[0117] The display module has a built-in world coordinate system, and receives the sparse 3D model and local dense 3D model of the current time frame based on the posture detection update display strategy, and displays the sparse 3D model of the current time frame and all local dense 3D models. All local dense 3D models include the local dense 3D model of the current time frame and all historical local dense 3D models. The historical local dense 3D model refers to the local dense 3D model of each time frame before the current time frame.

[0118] Furthermore, the calibration correction module generates a calibration instruction for a single tracking device through a dual-target calibration method, including the following specific steps:

[0119] Based on Tracking devices Binary positioning image sequence Left target positioning image in and right target image , the left camera and the right camera are calibrated by Zhang Zhengyou calibration method, and the left camera extrinsic parameters, the right camera extrinsic parameters, and the left eye center pixel coordinates are obtained respectively. 、Right eye center pixel coordinates and camera focal length , where the left camera external parameters include the left rotation matrix and the left eye translation vector , the right camera external parameters include the right rotation matrix and the right eye translation vector , left eye rotation matrix and the right rotation matrix Respectively represent the rotation transformation of the left camera coordinate system and the right camera coordinate system with respect to the world coordinate system in the world coordinate system, and the left camera translation vector and the right eye translation vector They represent the translation transformation of the origin of the left camera coordinate system and the origin of the right camera coordinate system with respect to the origin of the world coordinate system in the world coordinate system. Since the left camera and the right camera are of the same model, the focal length of the left camera is Equal to the focal length of the right camera Equal to the focal length of the camera , left eye centroid pixel coordinates is the pixel coordinate of the left camera optical center in the left camera pixel coordinate system, and the pixel coordinate of the right camera optical center is is the pixel coordinate of the optical center of the right camera in the right pixel coordinate system. The Zhang Zhengyou calibration method is an existing single-target calibration method. To track the device number, , To track the total number of devices;

[0120] World coordinates in the world coordinate system In the Tracking devices The left camera coordinates in the left camera coordinate system And the right camera coordinates in the right camera coordinate system The following coordinate unification equation is satisfied, and the coordinate unification equation is as follows:

[0121] ,

[0122] ,

[0123] The coordinate unification equation is processed by elimination method to eliminate the single coordinate in the world coordinate system , the coordinate transformation equation between the left camera coordinate system and the right camera coordinate system can be derived:

[0124] ,

[0125] Based on the coordinate transformation equation, we can obtain Tracking devices The rotation matrix and the self-translation vector ,Right now , , the rotation matrix and the self-translation vector They respectively represent the rotation transformation and translation transformation of the left camera coordinate system with respect to the right camera coordinate system in the world coordinate system;

[0126] In order to unify the image planes of the left camera and the right camera in the world coordinate system, the rotation matrix Partition into semi-positive rotation matrix and the semi-inverse rotation matrix , that is, the implementation method of unifying the image planes of the left camera and the right camera is to replace the full forward rotation of the left camera with the forward half rotation of the left camera and the reverse half rotation of the right camera;

[0127] To achieve horizontal calibration of the left and right cameras, Tracking devices The pole is translated to the infinity of the world coordinate system. The pole is the intersection of the visual line of the left camera and the right camera, and the self-translation vector is generated. Transformation matrix of visual lines in the same direction , the specific formula is as follows:

[0128] ,

[0129] ,

[0130] ,

[0131] Among them, the self-translation vector , , and are the self-translation vectors of Axis value, Axis value and Axis value, is the self-translation matrix The first visual vector in the same direction, is the second visual vector parallel to the unified image plane, is the third visual vector perpendicular to the unified image plane, Represents the transpose of a matrix or vector;

[0132] Combined with the semi-positive rotation matrix , semi-inverse rotation matrix and the visual line transformation matrix Get the left eye correction matrix separately and the right eye correction matrix , the specific formula is as follows:

[0133] ,

[0134] ,

[0135] Build Correction Instructions And send to Tracking devices .

[0136] Furthermore, the calibration module calculates and generates the main tracking device The main fixed parameter table includes the following specific steps:

[0137] Temporary storage of master tracking devices The pixel coordinates of the left eye center obtained by Zhang Zhengyou's calibration method , left eye translation vector 、Right eye center pixel coordinates , right eye translation vector and camera focal length ;

[0138] Get the corrected instructions Adjusted correction self-translation vector , the specific calculation formula is as follows:

[0139] ,

[0140] Further obtain the corrected self-translation vector In the world coordinate system Axis Value ;

[0141] Select the left camera as the reference camera based on the left eye centroid pixel coordinates 、Right eye center pixel coordinates , Camera focal length and the corrected self-translation vector Compute reference camera pose , the specific formula is as follows:

[0142] ,

[0143] in, and They are the pixel coordinates of the left eye center respectively. of Axis value and Axis value, The pixel coordinates of the right eye center of Axis value, Main tracking device Baseline distance , reference camera pose Used to convert the two-dimensional pixel coordinates in the left eye pixel coordinate system to the corresponding three-dimensional coordinates in the world coordinate system;

[0144] Packing the Master Tracking Device The reference camera pose , Camera focal length Distance from baseline , generate the main tracking device Main fixed parameters table.

[0145] Furthermore, the image stitching module generates the left eye global image based on the transformation stitching algorithm and the right eye global image The specific steps include:

[0146] According to the tracking device sequence, select the unconnected slave tracking device with the smallest tracking device number, assuming it is the Slave tracking devices , get the Slave tracking devices The binocular correction image sequence , No. Slave tracking devices The corresponding corrected image As the image to be stitched, , correct the image exist Time and and , respectively, represent the left eye correction image and right eye correction image , , To track the total number of devices;

[0147] Before acquisition The first splicing generated Sub-global image , using Harris corner detection algorithm to pair the Sub-global image With the correction image Get the corner point in The Harris corner point detection algorithm is an existing algorithm and will not be elaborated in detail;

[0148] The random sampling iterative algorithm is used to process the The second pairing corner point set generates the Subhomography matrix , and according to Subhomography matrix Generating Corrected Images On the main tracking device The reference camera pose The corresponding auxiliary correction image , ;

[0149] Obtain the auxiliary correction images corresponding to all slave tracking devices, and use the hat weighted average fusion algorithm to sequentially fusion the auxiliary correction images of a single slave tracking device with the main tracking device according to the tracking device number sequence. The corresponding corrected image Stitching to generate a global image , , global image exist Time and and represent the left eye global image respectively. and the right eye global image , and the global image Viewing angle and main tracking device The corresponding corrected image Same, subsequent processing and analysis can be directly combined with the main tracking device Main fixed parameters table.

[0150] like Figure 2 As shown, further, a random sampling iterative algorithm is used to process the The second pairing corner point set generates the Subhomography matrix The specific steps include:

[0151] Perform a random draw estimate, starting from Four paired corner points are randomly selected from the second paired corner point set to construct the The second internal point set is solved by the least squares method to obtain the homography equations used to describe the Sub-global image With the correction image The spatial transformation relationship of Subhomography matrix , the homography equations are as follows:

[0152] ,

[0153] in, represents the first Paired corner points, and Respectively represent The paired corner points are Sub-global image The actual pixel coordinates of Axis value and Axis value, and Respectively represent Paired corner points in the corrected image The actual pixel coordinates in of Axis value and Axis value, represents the transpose of a matrix or vector, ;

[0154] Adopt the Subhomography matrix Go to test The remaining paired corner points of the secondary paired corner point set, that is, the remaining paired corner points in the corrected image The actual pixel coordinates in Subhomography matrix Convert to Sub-global image The estimated pixel coordinates in Sub-global image The actual pixel coordinates in the calculation of the estimation error are used to calculate the estimation error, and the paired corner points whose estimation error is less than the estimation threshold are placed in the The secondary interior point set, ;

[0155] Judgement Whether the total number of paired corner points in the second inner point set is greater than the corner point number threshold, if the total number of paired corner points is greater than the corner point number threshold, then output the first Subhomography matrix ;

[0156] If the total number of paired corner points is less than or equal to the corner point number threshold, obtain the number of iteration rounds, and determine whether the number of iteration rounds is greater than the iteration round number threshold;

[0157] If the number of iterations is less than or equal to the iteration threshold, the first All paired corner points in the second internal point set are put back to the Pair the corner point set again and jump again to perform random sampling estimation;

[0158] If the number of iterations is greater than the iteration threshold, the output is the number of Subhomography matrix .

[0159] Furthermore, a hat-weighted average fusion algorithm is used to generate the left eye global image and the right eye global image The specific steps include:

[0160] Define the calibration image corresponding to the main tracking device The first global image , assuming that the second slave tracking device has been The auxiliary correction images of the slave tracking devices are all stitched into the correction image , get the Sub-global image , For the Sub-global image The level of vision, ;

[0161] Confirm Sub-global image With auxiliary correction image Overlapping area , get the overlapping area by counting All overlapping pixels of

[0162] For overlapping areas Middle Overlapping pixels, calculate the The number of overlapping pixels Sub-global image weights and auxiliary correction image weights , the specific formula is as follows:

[0163] ,

[0164] ,

[0165] in, and Respectively Overlapping pixels in Sub-global image The pixel coordinates in of Axis value and Axis value, and Respectively Overlapping pixels in the auxiliary correction image The pixel coordinates in of Axis value and Axis value, and Respectively Sub-global image The image width and image height, and Auxiliary correction images The image width and image height, , is the total number of overlapping pixels, ;

[0166] Calculate overlapping area No. The number of overlapping pixels Sub-global image weights and auxiliary correction image weights , the specific formula is as follows:

[0167] ,

[0168] ,

[0169] Combine them according to the order of overlapping pixels The number of overlapping pixels Sub-global image weights and auxiliary correction image weights , generating sub-global image weight vector and auxiliary correction image weight vector ;

[0170] The first Sub-global image With auxiliary correction image Splice to generate Sub-global image , the specific calculation formula is as follows:

[0171]

[0172] in, For the Sub-global image The pixel coordinates in Sub-global image That is the Sub-global image With auxiliary correction image Non-overlapping and overlapping areas The weighted sum of overlapping pixels in , ;

[0173] Judgement Sub-global image Vision level Is it equal to the total number of tracked devices? , if the vision level Equal to the total number of tracked devices , then Sub-global image The global image , ;

[0174] If the vision level Less than the total number of tracked devices , then continue to merge Auxiliary correction image of slave tracking device , until Stop after all auxiliary correction images of slave tracking devices are stitched together to obtain the global image , .

[0175] Furthermore, the 3D reconstruction module includes a disparity acquisition unit, a semantic segmentation unit, a sparse reconstruction unit, a pose estimation unit, a dense reconstruction unit and a storage unit;

[0176] The parallax acquisition unit receives the left eye global image and the right eye global image , based on the stereo matching algorithm to compare the left global image and the right eye global image The image patches of the same size in the image are used to generate the global disparity map , the left global image and global disparity map Send to the semantic segmentation unit. In this embodiment, the stereo matching algorithms include SAD, SSD, ZSAD, ZNCC and Rank. In this embodiment, ZSAD is used, which is a prior art and will not be elaborated on in detail.

[0177] The semantic segmentation unit uses the semantic segmentation annotation method to annotate the left global image. Add object tags and target tags, and process the global disparity map based on the object tags and target tags respectively , generate object disparity map and target disparity map , send the left eye global image and object disparity map To the sparse reconstruction unit, send the left eye global image and target disparity map To the pose estimation unit;

[0178] The sparse reconstruction unit receives the left eye global image and object disparity map , a reprojection transformation method is used to obtain a sparse three-dimensional model of the inspected object in the world coordinate system in the current time frame and send it to the display module;

[0179] The pose estimation unit receives the left eye global image and target disparity map , a dual feature pose estimation network is used to obtain the scanning extrinsic parameter matrix corresponding to the scanning device in the current time frame And sent to the dense reconstruction unit;

[0180] The dense reconstruction unit receives the scan extrinsic parameter matrix corresponding to the current time frame Local point cloud data transmitted by scanning equipment , using coordinate transformation method to generate local point cloud data The corresponding local dense three-dimensional model of the local object under inspection is sent to the display module;

[0181] The storage unit receives and stores the master tracking device The main fixed parameter table of , and responds to the calls of the sparse reconstruction unit and the pose estimation unit.

[0182] Furthermore, the semantic segmentation unit uses semantic segmentation annotation to generate object disparity maps and target disparity map The specific steps include:

[0183] The left global image Normalized and passed into the pre-trained residual neural network. As a classic convolutional neural network model, the residual neural network can effectively mine the left global image The latent features in generate the global feature map ;

[0184] For the global feature map For each feature point in the image, a candidate box of a fixed size is generated. All feature points in each candidate box are input into the region selection network for binary classification. It is distinguished whether the feature points in the candidate box belong to the foreground or the background. The candidate boxes whose feature points completely belong to the background are deleted. The foreground refers to the detected object or the tracking target, and the background specifically refers to the background excluding the detected object and the tracking target.

[0185] The region of interest alignment method is used to process the candidate frame and the left eye global image Pixel points and global feature map Align the feature points and candidate boxes to avoid the quantization error introduced by the traditional RoI-Pool operation;

[0186] The retained candidate boxes are combined with the global feature map The feature points in the image are input into the classifier and the mask marker respectively. The classifier is a combination of a fully connected layer and a Softmax function, which is used to classify the feature points in the candidate frame as belonging to the inspected object or the tracking target. The mask marker is composed of two convolutional neural networks, which is used to generate the marking mask corresponding to the candidate frame. The marking mask is a binary code, which represents the global image of the left eye. Which pixels aligned with the candidate box belong to the foreground?

[0187] Combine the classifier and mask marker to get the global image of the left eye. Add object tags and target tags to each pixel point, because the global disparity map The left eye global image Each pixel in the image corresponds one by one, and the global disparity map Find all the points corresponding to the pixels with object markers, adjust the disparity of the remaining points to 0, and generate the object disparity map ;

[0188] From the global disparity map Find all the points corresponding to the pixels with target marks, adjust the disparity of the remaining points to 0, and generate the target disparity map .

[0189] Furthermore, the sparse reconstruction unit uses a reprojection transformation method to obtain a sparse three-dimensional model, including the following specific steps:

[0190] Call the main fixed parameter table and extract the reference camera pose ;

[0191] The left global image The pixel coordinates of each pixel in the It indicates that, and Represents pixel coordinates In the left pixel coordinate system Axis value and Axis value;

[0192] From the object disparity map Get the left eye global image The object disparity corresponding to each pixel in the image is unified with express;

[0193] The left global image The world coordinates corresponding to each pixel point in the world coordinate system are uniformly expressed as It indicates that, , and Represents world coordinates In the world coordinate system Axis value, Axis value and Axis value;

[0194] Based on the reference camera pose The pixel coordinates of each pixel can be established With world coordinates The reprojection equation between is as follows:

[0195] ,

[0196] in, is the scale factor, which can be calculated based on the reprojection equation The world coordinates corresponding to each pixel ;

[0197] Get the left eye global image The world coordinates corresponding to all the pixels in the image are obtained, and a sparse three-dimensional model of the inspected object in the world coordinate system is reconstructed based on the greedy triangulation method. The greedy triangulation method is an existing algorithm and will not be elaborated on. When the inspected object moves and rotates due to external force in the current time frame, the sparse three-dimensional model of the current time frame will also move and rotate in the same manner, thereby realizing dynamic tracking of the inspected object.

[0198] like Figure 3 As shown, further, the pose estimation unit uses a dual feature pose estimation network to obtain the scanning extrinsic parameter matrix of the scanning device in the current time frame The specific steps include:

[0199] Call the main fixed parameter table and extract the camera focal length Distance from baseline , based on the principle of triangulation, the target disparity map The parallax of any point in Corresponding conversion to depth , generate target depth map ;

[0200] Left eye global image Mainly color and texture features, which can highlight visual boundaries. Feature extraction is performed through a convolution layer with a convolution kernel size of 7×7, a convolution step of 2, and a padding of 3. The output features of the convolution layer are further reduced to half of the original dimension through a normalization layer and a ReLU activation function to generate a visual feature map. ;

[0201] Target Depth Map The main feature is the spatial three-dimensional feature, which can highlight the geometric boundary of the tracking target and convert the target depth map into a point cloud network. The geometric point data is converted into geometric point data, and then the geometric point data is nonlinearly transformed through a multi-layer perceptron to extract geometric features. The linear layer reduces the dimension to half of the original dimension to generate a target geometric feature map. ;

[0202] The convolution feature fusion module uses 3×3 convolution to process the visual feature map. and target geometry , and perform fusion transformation through dot multiplication, and then use a 3×3 convolution to restore the original channel and add it to the visual feature map respectively and target geometry In the above example, the first fusion feature map is generated. and the second fusion feature map , the specific calculation formula is as follows:

[0203] ,

[0204] ,

[0205] in, , and They are three 3×3 convolutions involved in the convolutional feature fusion module;

[0206] The first fusion feature map is processed using the channel attention mechanism combined with average pooling , smooth attention to the complete first fusion feature map , the second fusion feature map is processed by the channel attention mechanism combined with the maximum pooling , focusing on the second fusion feature map The output results of the two channels are accumulated by element-wise addition and further processed by a one-dimensional convolution layer, a normalization layer, and a ReLU activation function to generate a dual fusion feature map. ;

[0207] Dual fusion feature map The dual fusion feature map takes into account both the visual and geometric features of the tracking target. Further input 1 maximum pooling layer and 4 fully connected layers for regression to obtain a 6-dimensional target pose sequence ,in, , and are the scanning device coordinate system of the scanning device axis, Axis and Axis with respect to the world coordinate system axis, Axis and The rotation angle of the axis, , and The origin of the scanning device coordinate system is about the origin of the world coordinate system. axis, Axis and The translation value of the axis;

[0208] Constructing the scanning extrinsic matrix ,in, and They are the scanning rotation matrix and the scanning translation vector, which respectively represent the rotation transformation of the scanning device coordinate system with respect to the world coordinate system and the translation transformation of the origin of the scanning device coordinate system with respect to the origin of the world coordinate system. The specific calculation formulas are as follows:

[0209] ,

[0210] ,

[0211] in, Represents the transpose of a vector.

[0212] Furthermore, the dense reconstruction unit uses coordinate transformation to generate local point cloud data of the current time frame. The corresponding local dense 3D model includes the following specific steps:

[0213] Get the scan extrinsic parameter matrix corresponding to the current time frame ,in, and are the scanning rotation matrix and scanning translation vector respectively;

[0214] The local point cloud data of the current time frame The scanning device coordinates of each point cloud are unified using Indicates that, the scanning device coordinates refer to the three-dimensional coordinates of each point cloud in the scanning device coordinate system. , and They are the scanning device coordinates in the scanning device coordinate system. Axis value, Axis value and Axis value;

[0215] The world coordinates corresponding to each point cloud in the world coordinate system are unified using express, , and The world coordinates are In the world coordinate system Axis value, Axis value and Axis value;

[0216] Based on scanning extrinsic matrix The scanning device coordinates of each point cloud can be established The corresponding world coordinates The coordinate transformation equation between them is as follows:

[0217] ,

[0218] Based on the coordinate transformation equation, the world coordinates corresponding to each point cloud can be calculated ;

[0219] Get local point cloud data The world coordinates corresponding to all point clouds are obtained, and local point cloud data is reconstructed based on the greedy triangulation method. The local dense 3D model corresponding to the world coordinate system, where the greedy triangulation method is an existing algorithm and will not be elaborated on in detail. The scanning extrinsic matrix is ​​calculated in real time in the current time frame. Combined with local point cloud data acquired in real time , the local dense 3D model of the current time frame can be generated in real time to achieve dynamic tracking and scanning.

[0220] Furthermore, in the display module, the sparse 3D model of the current time frame and all local dense 3D models are displayed in real time based on the posture detection updating display strategy, including the following specific steps:

[0221] Receive the sparse 3D model and the local dense 3D model of the current time frame and attach color labels, the color labels are used to indicate the display colors of the sparse 3D model and the local dense 3D model, and determine whether there is a sparse 3D model of the latest frame in the display memory, the sparse 3D model of the latest frame specifically refers to the sparse 3D model of the previous time frame of the current time frame;

[0222] If there is no sparse 3D model of the most recent frame, it indicates that the sparse 3D model and the local dense 3D model of the current time frame are received for the first time, and are directly stored in the display memory and displayed according to the color label;

[0223] If there is a sparse 3D model of the latest frame, the sparse 3D model of the latest frame is called;

[0224] The key points of the sparse 3D model of the current time frame and the sparse 3D model of the most recent frame are extracted respectively by the scale-invariant feature transformation method, and the key point matching is further assisted by the feature matching algorithm. The scale-invariant feature transformation method and the feature matching algorithm are both existing technologies and will not be elaborated in detail;

[0225] Calculate the average Euclidean distance of all key points in the sparse 3D model of the current time frame and the sparse 3D model of the most recent frame, and determine whether the average Euclidean distance is greater than the Euclidean distance threshold;

[0226] If the average Euclidean distance is less than or equal to the Euclidean distance threshold, it indicates that the position and posture of the detected object has not changed from the previous time frame to the current time frame, and the sparse 3D model of the current time frame replaces the sparse 3D model of the most recent frame, and the sparse 3D model of the current time frame is stored in the display memory, and the sparse 3D model of the current time frame and all local dense 3D models are displayed according to the color label;

[0227] If the average Euclidean distance is greater than the Euclidean distance threshold, it indicates that the posture of the object under inspection has changed from the previous time frame to the current time frame. Select any 6 matched key points from the sparse 3D model of the current time frame and the sparse 3D model of the most recent frame to construct 6 key point pairs.

[0228] The object rotation matrix and the object translation vector from the previous time frame to the current time frame are calculated by fitting the world coordinates of the six key point pairs using the least squares method. The object rotation matrix and the object translation vector are used to represent the rotation relationship and translation relationship between the sparse 3D model of the most recent frame and the sparse 3D model of the current time frame.

[0229] All historical local dense 3D models in the display memory are subjected to corresponding rotation and translation transformations according to the object rotation matrix and the object translation vector, the sparse 3D model of the current time frame replaces the sparse 3D model of the most recent frame, the sparse 3D model of the current time frame is stored in the display memory, and the sparse 3D model of the current time frame and all local dense 3D models are displayed according to color labels.

[0230] Furthermore, tracking devices Taiwan, respectively , where the first tracking device The main tracking device, the rest As a slave tracking device, it is used to expand the field of view of the master tracking device. Each tracking device selects to perform calibration correction and image acquisition based on the instruction-driven strategy. In this embodiment, Tracking Devices They are all binocular cameras of the same model. The binocular cameras need to be calibrated to achieve baseline horizontal calibration when the system is started, and the calibrated posture is always fixed during system operation.

[0231] Furthermore, a single tracking device may perform calibration and image acquisition based on a command-driven strategy including:

[0232] If the received command is an initialization command, the calibration and correction of a single tracking device is performed;

[0233] If the received instruction is a frame synchronization instruction, the image acquisition of a single tracking device is executed, the inspected object is photographed in a fixed posture after calibration and correction, and a binocular correction image sequence is obtained and fed back to the central control subsystem, wherein the binocular correction image sequence includes the left eye correction image and the right eye correction image of the single tracking device.

[0234] Furthermore, the calibration of a single tracking device includes the following specific steps:

[0235] First Tracking devices For example, , To track the total number of devices, set the device status to Startup;

[0236] Shoot the preset calibration target to obtain the Tracking devices Binary positioning image sequence And fed back to the calibration correction module, where: and They are the left target positioning image and the right target positioning image respectively;

[0237] Receive the Tracking devices Correction instructions , through the left eye correction matrix Multiply the left camera external parameters and the right camera correction matrix The left and right cameras are calibrated by multiplying the external parameters of the right camera. The postures of the left and right cameras after correction are fixed so that the baselines of the left and right cameras are calibrated horizontally.

[0238] Furthermore, the scanning device includes a scanner and a tracking target. The scanner receives a frame synchronization instruction and activates a scanning function to obtain local point cloud data of the current time frame. And feedback is given to the central control subsystem. The tracking target is a regular hexahedron stereoscopic target commonly used in the prior art. It is fixed on the scanner through a connecting rod support to provide a stable visual tracking point to assist the central control subsystem in determining the scanning posture of the scanner.

[0239] Example 2

[0240] like Figure 4 As shown, the present invention also discloses a dynamic tracking optical scanning method, which is implemented based on a dynamic tracking optical scanning system and includes the following specific steps:

[0241] S1, receiving a start command and sending an initialization command to all tracking devices, obtaining a binocular positioning image sequence of a single tracking device, generating a correction command through a binocular positioning correction method and feeding it back to the single tracking device, storing a main fixed parameter table, and receiving a scan start command;

[0242] S2, broadcast the frame synchronization instruction in the current time frame, and set the next time frame according to the frame interval;

[0243] S3: Obtain the binocular correction image sequence of all tracking devices in the current time frame, and generate the left eye global image based on the transformation stitching algorithm and the right eye global image ;

[0244] S4. Analyze the left eye global image based on stereo matching algorithm and the right eye global image Generate global disparity map , semantic segmentation and annotation method is used to further process the global disparity map Generate object disparity map and target disparity map , obtain the sparse 3D model of the current time frame through the reprojection transformation method;

[0245] S5. Obtain the scanning extrinsic parameter matrix of the scanning device of the current time frame through the dual feature pose estimation network , receive the local point cloud data of the current time frame , and is transformed into a local dense three-dimensional model of the current time frame through the coordinate transformation method;

[0246] S6, updating the display strategy based on the pose detection to display the sparse 3D model of the current time frame and all local dense 3D models;

[0247] S7, waiting for the system time to advance to the next time frame, and jumping to step S2, until the scanning of the inspected object is completed.

[0248] The present invention discloses a dynamic tracking optical scanning system and method, which controls a plurality of tracking devices to perform calibration correction and equally spaced image acquisition through a central control subsystem, controls a scanning device to perform equally spaced scanning, analyzes and processes all binocular correction image sequences and local point cloud data of the current time frame, restores and generates a sparse 3D model and a local dense 3D model and displays them, wherein the central control subsystem comprises a signal generation module, a frame synchronization module, a calibration correction module, an image stitching module, a 3D reconstruction module and a display module; the signal generation module receives a start instruction and sends an initialization signal, receives a scan start instruction and decides whether to send a scan start signal based on a correction response strategy; the frame synchronization module receives a scan start signal and broadcasts a frame synchronization instruction, sets time frames at equal intervals and broadcasts the frame synchronization instruction in each time frame, thereby realizing equally spaced image acquisition and scanning of a plurality of tracking devices and scanning devices; the calibration correction module receives an initialization signal and sends an initialization instruction, generates a correction instruction for a single tracking device through a binocular calibration correction method and feeds back the correction instruction, thereby realizing tracking The baseline level calibration of the tracking device is performed, the main fixed parameter table is calculated and the correction completion signal is sent; the image stitching module receives all binocular correction image sequences of the current time frame, generates the left eye global image and the right eye global image based on the transformation stitching algorithm, and realizes the field of view expansion of multiple tracking devices; the 3D reconstruction module stores the main fixed parameter table, analyzes the left eye global image and the right eye global image based on the stereo matching algorithm to generate a global disparity map, uses the semantic segmentation and annotation method to generate the object disparity map and the target disparity map, uses the reprojection transformation method to obtain the sparse 3D model of the current time frame, obtains the scanning extrinsic parameter matrix of the current time frame from the dual perspectives of vision and geometry through the dual feature pose estimation network, and converts the local point cloud data of the current time frame into a local dense 3D model based on the coordinate transformation method, realizing the reconstruction of the local 3D model of the current time frame; the display module updates the display strategy based on the pose detection to display the sparse 3D model of the current time frame and all local dense 3D models, ensuring that they can be displayed in time when the poses of the inspected object and the scanning device change.

[0249] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A dynamic tracking optical scanning system, comprising a central control subsystem, a tracking device and a scanning device, characterized in that: The central control subsystem includes a frame synchronization module, a calibration correction module, an image stitching module, a three-dimensional reconstruction module and a display module; The frame synchronization module sets a time frame and broadcasts a frame synchronization instruction in each time frame; The calibration and correction module sends an initialization instruction, feeds back a correction instruction through a dual-target calibration and correction method, calculates a main fixed parameter table and sends it to a three-dimensional reconstruction module; The image stitching module receives all binocular correction image sequences of the current time frame, generates a left-eye global image and a right-eye global image based on a transformation stitching algorithm, and sends them to the 3D reconstruction module; The 3D reconstruction module stores a main fixed parameter table, uses a stereo matching algorithm to analyze the left eye global image and the right eye global image to generate a global disparity map, uses a semantic segmentation and annotation method to process the global disparity map to generate an object disparity map and a target disparity map, uses a reprojection transformation method to obtain a sparse 3D model of the current time frame and sends it to the display module, obtains a scanning extrinsic parameter matrix of the current time frame through a dual feature pose estimation network, receives local point cloud data of the current time frame, converts it into a local dense 3D model through a coordinate transformation method, and sends it to the display module; The display module updates the display strategy based on the posture detection to display the sparse three-dimensional model of the current time frame and all local dense three-dimensional models.

2. A dynamic tracking optical scanning system as claimed in claim 1, characterized in that: The calibration correction module feeds back the correction instruction through the dual-target calibration correction method, including the following specific steps: Get the left camera extrinsics and the right camera extrinsics of a single tracking device; Establish a coordinate unification equation, derive the coordinate transformation equation through elimination method, and obtain the self-rotation matrix and self-translation vector based on the coordinate transformation equation; The self-rotation matrix is ​​divided into a half-forward rotation matrix and a half-inverse rotation matrix, the pole of a single tracking device is translated to infinity, and the visual line transformation matrix is ​​calculated; The left eye correction matrix and the right eye correction matrix are obtained by combining the half-forward rotation matrix, the half-inverse rotation matrix and the visual line transformation matrix, respectively, and the correction instructions are constructed and sent to a single tracking device.

3. A dynamic tracking optical scanning system as claimed in claim 1, characterized in that: The image stitching module generates a left-eye global image and a right-eye global image based on a transformation stitching algorithm, and the specific steps are as follows: Get the first The rectified images corresponding to the slave tracking devices are used as the images to be stitched. , To track the total number of devices; Before acquisition The first splicing generated The second global image is paired with the Harris corner detection algorithm. The corner points in the global image and the image to be stitched are obtained. Secondary paired corner point set; The random sampling iterative algorithm is used to process the The second pairing of corner points generates the The homography matrix is The secondary homography matrix transforms the image to be stitched into an auxiliary corrected image; The auxiliary correction images of all slave tracking devices are obtained, and the auxiliary correction images are sequentially spliced ​​to the correction images corresponding to the main tracking device using the hat weighted average fusion algorithm to generate a global image, which includes the left eye global image and the right eye global image.

4. A dynamic tracking optical scanning system as claimed in claim 3, characterized in that: The random sampling iterative algorithm is used to generate The subhomography matrix includes the following specific steps: Perform a random draw estimate, starting from The second pairing corner points are randomly selected to construct the The second internal point set is solved by the least squares method to obtain the The secondary homography matrix, , To track the total number of devices; Based on The homography matrix is ​​calculated The estimation errors of all the remaining paired corner points in the second paired corner point set are calculated, and the paired corner points with estimation errors less than the estimation threshold are placed in the second paired corner point set. Secondary interior point set; Determine whether the total number of paired corner points is greater than the corner point number threshold. If so, output the If the secondary homography matrix is ​​less than or equal to the corner point number threshold, obtain the number of iterations and determine whether the number of iterations is greater than the iteration number threshold; If less than or equal to, All paired corner points in the second internal point set are put back into the Pair the corner point set again and jump again to perform random sampling estimation; If it is greater, output the value of the previous iteration. Subhomography matrix.

5. A dynamic tracking optical scanning system as claimed in claim 3, characterized in that: The use of the hat weighted average fusion algorithm to generate a global image includes the following specific steps: Get the Sub-global image, is the level of vision; Confirm The second global image and the The overlapping areas and overlapping pixels of the auxiliary correction images corresponding to the slave tracking devices; For a single overlapping pixel, The sub-global image weight accounts for the first The ratio of the sum of the sub-global image weights is taken as the Sub-global image weight, the ratio of the auxiliary correction image weight of a single overlapping pixel to the total auxiliary correction image weights of all overlapping pixels is taken as the auxiliary correction image weight; Get the The global image weight vector and the auxiliary correction image weight vector are The non-overlapping area of ​​the secondary global image, the non-overlapping area of ​​the auxiliary correction image, and the weighted sum of all overlapping pixels are spliced ​​to generate the first Sub-global image; Judgement Field of view level of sub-global image Is it equal to the total number of tracking devices? If so, then The sub-global image is the global image; If it is less than, continue splicing The auxiliary correction images of the slave tracking devices are used until the global image is obtained.

6. A dynamic tracking optical scanning system as claimed in claim 1, characterized in that: The three-dimensional reconstruction module includes a disparity acquisition unit, a semantic segmentation unit, a sparse reconstruction unit, a pose estimation unit, a dense reconstruction unit and a storage unit; The disparity acquisition unit compares the left eye global image and the right eye global image based on a stereo matching algorithm to generate a global disparity map, and sends the left eye global image and the global disparity map to the semantic segmentation unit; The semantic segmentation unit uses the semantic segmentation annotation method to analyze and process the left-eye global image and the global disparity map, generates an object disparity map and a target disparity map, sends the left-eye global image and the object disparity map to the sparse reconstruction unit, and sends the left-eye global image and the target disparity map to the pose estimation unit; The sparse reconstruction unit uses a reprojection transformation method to obtain a sparse three-dimensional model of the current time frame and sends it to the display module; The pose estimation unit uses a dual-feature pose estimation network to process the left-eye global image and the target disparity map, obtains the scanning extrinsic parameter matrix of the current time frame and sends it to the dense reconstruction unit; The dense reconstruction unit uses a coordinate transformation method to combine the scanned external parameter matrix and local point cloud data of the current time frame to generate a local dense three-dimensional model of the current time frame and send it to the display module; The storage unit receives and stores the main fixed parameter table and responds to the call.

7. A dynamic tracking optical scanning system as claimed in claim 6, characterized in that: The semantic segmentation unit generates an object disparity map and a target disparity map by using a semantic segmentation annotation method, and the specific steps include: Normalize the left eye global image and generate a global feature map through residual neural network mining; Generate several candidate boxes, and perform binary classification through the region selection network to retain some candidate boxes; The region of interest alignment method is used to align the pixels of the left global image with the feature points and candidate boxes of the global feature map; The retained candidate boxes and the feature points in the global feature map are input into the classifier and mask marker respectively. The classifier uses a combination of a fully connected layer and a Softmax function to classify the feature points in the candidate boxes. The mask marker generates a marking mask corresponding to the candidate boxes through two convolutional neural networks. Combined with the classifier and the mask marker, object tags and target tags are added to each pixel of the left-eye global image. The global disparity map is processed based on the object tags and target tags to generate the object disparity map and the target disparity map.

8. A dynamic tracking optical scanning system as claimed in claim 6, characterized in that: The pose estimation unit uses a dual feature pose estimation network to obtain a scanning external parameter matrix of the current time frame, including the following specific steps: Call the main fixed parameter table and extract the camera focal length and baseline distance, and convert the target disparity map into a target depth map based on the triangulation principle; The left eye global image is processed through convolutional layers, normalization layers and ReLU activation functions to generate a visual feature map; The target depth map is converted into geometric point data through the point cloud network, nonlinear transformation is performed through the multi-layer perceptron, and the target geometric feature map is generated through linear layer dimensionality reduction; The convolution feature fusion module performs convolution and dot multiplication on the visual feature map and the target geometric feature map to achieve fusion transformation, and then performs convolution again and adds them to the visual feature map and the target geometric feature map respectively to generate a first fusion feature map and a second fusion feature map; The first fused feature map is processed by the channel attention mechanism combined with average pooling, and the second fused feature map is processed by the channel attention mechanism combined with maximum pooling. The output results of the two channels are accumulated by element-wise addition, and processed through a one-dimensional convolution layer, a normalization layer, and a ReLU activation function to generate a double fused feature map. The dual fusion feature map is further input into the maximum pooling layer and multiple fully connected layers for regression to obtain the target pose sequence and construct the scanning extrinsic parameter matrix.

9. A dynamic tracking optical scanning system as claimed in claim 1, characterized in that: The display module updates the display strategy based on the posture detection to display the sparse 3D model of the current time frame and all local dense 3D models, including the following specific steps: Receive the sparse 3D model and the local dense 3D model of the current time frame, and determine whether there is a sparse 3D model of the most recent frame; If not, the sparse 3D model and the local dense 3D model of the current time frame are stored and displayed; If it exists, the sparse 3D model of the nearest frame is called to match the key points through the scale-invariant feature transformation method and feature matching algorithm; Calculate the average Euclidean distance of all key points and determine whether the average Euclidean distance is greater than the Euclidean distance threshold; If it is less than or equal to, the sparse 3D model of the current time frame replaces the sparse 3D model of the most recent frame and stores the sparse 3D model of the current time frame, and displays the sparse 3D model of the current time frame and all local dense 3D models; If it is greater than that, construct 6 key point pairs and calculate the object rotation matrix and object translation vector from the previous time frame to the current time frame by the least squares method; All stored historical local dense 3D models are subjected to corresponding rotation and translation transformations according to the object rotation matrix and the object translation vector, the sparse 3D model of the current time frame replaces the sparse 3D model of the most recent frame and the sparse 3D model of the current time frame is stored, and the sparse 3D model of the current time frame and all local dense 3D models are displayed.

10. A dynamic tracking optical scanning method, characterized in that: The specific steps include: S1, receiving a start command and sending an initialization command to a tracking device, obtaining a binocular positioning image sequence of a single tracking device and feeding back a correction command through a binocular positioning correction method, storing a main fixed parameter table, and receiving a scan start command; S2, broadcast the frame synchronization instruction in the current time frame and set the next time frame; S3, obtaining all binocular correction image sequences of the current time frame, and generating a left-eye global image and a right-eye global image based on a transformation stitching algorithm; S4, using a stereo matching algorithm to analyze the left eye global image and the right eye global image to generate a global disparity map, using a semantic segmentation and annotation method to process the global disparity map to generate an object disparity map and a target disparity map, and obtaining a sparse three-dimensional model of the current time frame through a reprojection transformation method; S5, obtaining the scanning external parameter matrix of the current time frame through the dual feature pose estimation network, receiving the local point cloud data of the current time frame, and converting it into a local dense three-dimensional model through a coordinate transformation method; S6, updating the display strategy based on the pose detection to display the sparse 3D model of the current time frame and all local dense 3D models; S7, waiting for the system time to advance to the next time frame, and jumping to step S2, until the scanning of the inspected object is completed.

Citation Information

Patent Citations

  • Tracking and scanning system and tracking and scanning processing method

    CN117870543A

  • SLAM method based on tight coupling of 2D laser radar and binocular camera

    CN112785702A

  • Endoscopic image three-dimensional reconstruction method combining SfM and binocular matching

    CN112967330A