Six-degree-of-freedom visual feedback real-time motion tracking method

By using a visual feedback dual-path parallel computing module and DenseNet initialization method, the problems of dependence on computer-aided design models and time-consuming training databases in existing technologies are solved. This achieves real-time accuracy and efficiency in 6D object pose and motion tracking for intelligent robots, improving the application effect in digital factory scenarios.

CN115187633BActive Publication Date: 2026-08-04WUXI DONGRU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUXI DONGRU TECH CO LTD
Filing Date
2022-07-12
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing six-DOF visual feedback real-time motion tracking methods require a pre-defined computer-aided design model, which limits their application scope. Furthermore, training the database is time-consuming and labor-intensive, making it difficult to cover the diverse object categories in digital factory scenarios.

Method used

A visual feedback dual-path parallel computing module is adopted, and the model is optimized through an image segmentation framework and constraint functions. Combined with a dual-path parallel network for image segmentation and key point detection, it achieves a significant improvement over existing methods. DenseNet is used for initialization, and semantic segmentation with sparse annotations and high-speed CUDA programming are used to achieve efficient real-time motion tracking.

Benefits of technology

In digital factory scenarios, the system achieves real-time and accurate 6D object pose and motion tracking for intelligent robots, significantly improving runtime and efficiency, reducing reliance on computer-aided design models, and enhancing accuracy and speed, achieving a real-time output frame rate of 26Hz.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_9
    Figure QLYQS_9
  • Figure QLYQS_19
    Figure QLYQS_19
  • Figure QLYQS_22
    Figure QLYQS_22
Patent Text Reader

Abstract

The application discloses a kind of six degrees of freedom visual feedback real-time motion tracking methods, comprising the following operating steps: constructing image segmentation framework and constraint function, and realizing model optimization by minimizing function;Target object segmentation real-time video two-way two-group data processing;Video object image segmentation method;Image segmentation result two-way output method design and two-way parallel network key point detection matching, realize visual feedback;According to the initial pose of object output by the above two-way parallel, realize object motion tracking.The six degrees of freedom visual feedback real-time motion tracking method disclosed in the application adopts visual feedback two-way parallel computing module to realize video frame image segmentation network, respectively synchronously input video sequence for image segmentation, realize the significant improvement to existing most advanced method, compared with other similar methods has the running time of significant competitiveness, solve the problem of intelligent robot 6D object pose and motion tracking control real-time demand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent manufacturing and machine vision, and in particular to a six-degree-of-freedom visual feedback real-time motion tracking method. Background Technology

[0002] Six-degree-of-freedom (6DOF) visual feedback real-time motion tracking is a method for intelligent manufacturing and machine vision tracking. It has wide applications in intelligent robots in digital factory scenarios. Robot pose estimation, motion tracking, and real-time precise control are key technologies for the effective implementation of this application. At the same time, the motion tracking of the manipulated object is particularly important. Detecting the target object to be manipulated and accurately estimating its 3D position, orientation, and actual size and motion tracking are critical links in robot technology and 3D scene analysis. Six-degree-of-freedom object pose estimation and motion tracking, also known as 6D object pose estimation and motion tracking, is increasingly demanding in manufacturing processes as technology continues to develop.

[0003] Existing six-DOF visual feedback real-time motion tracking methods have certain drawbacks. Current work has explored a series of methods for six-DOF object pose estimation and motion tracking at the instance level. Some of these methods require a pre-defined computer-aided design model, which severely limits their practical application. This prevents existing technologies from being used for pose estimation and motion tracking of most objects without a computer-aided design model. Other methods achieve pose estimation and motion tracking by training computer-aided design models of the same type of object, but these still have limitations. This is because these methods are limited by the various categories in the training database, and therefore are still insufficient to cover the diverse object categories in intelligent robot applications in digital factory scenarios. At the same time, 3D model databases usually require a large amount of manual work and collaboration from experts with specialized domain knowledge to build, which is time-consuming and labor-intensive, and has a certain adverse impact on the user experience. To address these issues, we propose a six-DOF visual feedback real-time motion tracking method. Summary of the Invention

[0004] (a) Technical problems to be solved

[0005] To address the shortcomings of existing technologies, this invention provides a six-degree-of-freedom visual feedback real-time motion tracking method. It employs a dual-path parallel computing module for visual feedback to implement a video frame image segmentation network, simultaneously inputting video sequences for image segmentation. This represents a significant improvement over existing state-of-the-art methods, offering significantly faster runtime compared to other similar methods. It effectively solves the real-time requirements for 6D object pose and motion tracking control in intelligent robots, thus addressing the problems mentioned in the background technology.

[0006] (II) Technical Solution

[0007] To achieve the above objectives, the technical solution adopted by the present invention is: a six-degree-of-freedom visual feedback real-time motion tracking method, comprising the following operational steps:

[0008] S1: Construct an image segmentation framework and constraint functions, and optimize the model by minimizing the functions;

[0009] S2: Real-time dual-channel binary grouping data processing for target object segmentation;

[0010] S3: Video object image segmentation method;

[0011] S4: Design of a dual-output method for image segmentation results and key point detection and matching using a dual-parallel network to achieve visual feedback;

[0012] S5: Based on the initial pose of the object output by the dual parallel output above, the pose graph is associated with the finely optimized 6D pose of the object at the current moment to achieve object motion tracking.

[0013] As a preferred technical solution of this application, step S1 specifically includes the following implementation steps:

[0014] A1: Define a time-series image segmentation dataset from a video series. The video dataset consists of a large number of consecutive image frames arranged in temporal order, represented as follows: This represents a video dataset, where the subscript I indicates that the set elements consist of image frames, and g i Let i = 1, 2, ..., n represent n consecutive frames in the video. These represent the background and foreground object segmentation labels for the corresponding image frames. The dataset contains... 1 tag data pair and One unlabeled data point;

[0015] A2: Based on all provided image segmentation datasets Inferring unlabeled image data Object segmentation labels The process of inferring segmentation labels for unlabeled data can be summarized as operator representation. Where tanh is the hyperbolic tangent function, expressed as tanh(x) = (e^(x-1) / (x-1)). x -e -x ) / (e x +e -x Exponential linear units (ELU) c is a positive constant. This is called the segmentation label inference operator, where w ij Represents pixel data points (g) in two different images of a encoded video sequence. i g j The similarity between ) where ε i =∑ j w ij This represents the normalized scaling factor at image pixel i, where the hyperparameter γ is used as an operator. The weighted balance factor between the two terms;

[0016] A3: Based on a continuous image frame dataset Perform smooth constraint operations on the segmentation label inference operator. It contains two sums, where the first term is... For dataset All terms (1, 2, ..., n terms, i.e. n frames of time series images) contain labeled and unlabeled image data. Running the first term implements the smoothing constraint corresponding to the pixel. The calculation result of this run realizes that similar pixels of the gradient descent method have close or approximately the same label value.

[0017] A4: Based on dataset Implement cluster constraint operations in the segmentation label inference operator. The second item included Partially targeting video datasets The first l terms, i.e., terms 1, 2, ..., l, are image data containing already labeled segmentation tags. The operation performed on this subset of data is a cluster constraint operation, i.e., calculating... This process enables the model inference values ​​to be converted into... The values ​​are aggregated with the true label value h, and the inferred value is calculated for each corresponding pixel. Compared with the true label value h i Distance correction is achieved by minimizing constraints to optimize the inferred observations that deviate from the target value.

[0018] A5: Iterative optimization of the segmentation label inference operator Run the following optimization process Specifically, this is achieved through a minimization algorithm, making For w ij The constructed normalized similarity matrix, where matrix V and The eigenvalue matrix is ​​obtained through iterative solution. The iterative process continues until convergence, as follows: h(0) is the initial state of the system, h(0) = [h1, h2, ..., h n ] TThese are the initial observations of the labels clamped using supervised labels; Let be the result of the system at step k. Let be the result of the system at step (k+1) as follows. Where τ is an empirical value, typically ranging from 0.960 to 0.981; the advantage of this iterative algorithm is that it constructs a global model on the dense structure of unlabeled data.

[0019] As a preferred technical solution of this application, step S2 specifically includes the following implementation steps:

[0020] B1: The input video sequence is represented in chronological order as follows: Where v i Let i = 1, 2, ..., n represent the frames in the video sequence arranged chronologically from past to present, and v n This represents the latest, current frame of the image, distinct from the frame used in step S1. The expression, This is a general video sequence dataset corresponding to the description of video image object segmentation algorithms, possessing generality, universality, and summarization; the video sequences used in this section are... This indicates that it has the specificity of its application scenario, because in this application scenario, only This single image frame contains labeled data with ground truth values, corresponding to the dataset. The first frame (g1, h1) in the data. Other versions such as v2, v3, ..., v n These are all unsupervised data, similar to typical video sequence datasets. In These items, that is, relative to a typical video sequence dataset, here Application scenarios

[0021] B2: Video sequence The frames are divided into two groups, A and B, alternating according to the order of odd and even frames. and Dataset The first frame v1 in the image is used as the starting frame and is the only image frame with ground truth labels. Therefore, it is shared by both groups A and B. The first frame v1 is... and Both are used as the starting frame of their dataset, specifically represented as as well as v1 corresponds to a general video sequence dataset The first term in the array is (g1, h1), v3, v5, ..., v 2t-1 and v2, v4, v6, ..., v 2t Corresponding to general video sequence datasets of Unsupervised data, where t∈N, For the sake of uniformity in the following explanation, let n be a positive even number, and when n is a positive odd number, the last frame of the image is aligned by interpolation;

[0022] B3: An image segmentation network implemented using a dual-GPU parallel computing module. These two computing modules are represented as follows: and Video sequences with real-time parallel alternating frame input and Image segmentation is performed synchronously, corresponding to... and The parallel output includes two sets of split results.

[0023] As a preferred technical solution of this application, step S3 specifically includes the following implementation steps:

[0024] C1: The video image segmentation network consists of two computational modules, A and B, with identical structures. These modules are computed in parallel in real-time on two GPUs with identical parameter configurations. These two computational modules are represented as follows: and Input video sequences separately and Image segmentation is performed because and The operation is exactly the same, so we use module A for calculation. For video sequences Image segmentation follows a standard operating procedure.

[0025] C2: The algorithm runs online, using video sequences. As online input, predictions for all previous frames are determined when the current frame (2t-1) arrives, and then the predictions are approximated over time. Among them U 1:(2t-1)→(2t+1) Let U represent the similarity matrix U constructed only between pixels up to frame (2t-1) and pixels in frame (2t+1). Since no labels are provided outside the first frame, the initial term h(0) is omitted in frame (2t+1). For time (2t+1), the above propagation process equivalently minimizes a set of smooth terms in the spatiotemporal pixel array. Where i indexes the pixel at the target time (2t+1), and j indexes the pixels in all frames before and including time (2t-1), compared to operators in typical video sequence datasets. i and j can take any values ​​1, 2, ..., n. Here, due to the strict cross-temporal nature of video frames, i can only take the latest frame, i.e., the target time (2t+1), and j can take the frame at time (2t-1) and the corresponding frame at previous times.

[0026] C3: Given the start frame of the video The label on Process the remaining frames in sequence. From iterative equations Propagate the tags to every frame;

[0027] C4: Optimization of similarity metrics, because the quality of video object segmentation depends on the similarity metric U and the similarity metric w. ij From the appearance item The spatial term Sigmod(-Canberra(loc(i), loc(j)) / ε 2 Optimize in two aspects

[0028] Where f i f j It is pixel g i g j Through feature embedding via a convolutional neural network, loc(i) represents the spatial location of pixel i, with the spatial term controlled by the local parameter ε, where the function Sigmod(t) = 1 / (1+e^(i-1)). -t Canberra is far from

[0029] C5: Based on a 3D dilated convolutional neural network, this project learns foreground object appearance embeddings in an image in a data-driven manner. The appearance embeddings aim to capture changes caused by motion, scale, and image distortion. They are learned from training data where each frame in the video is labeled with segmented objects and object identifiers, given a target pixel g. i We treat all pixels from the previous frame as a reference and set f i and f j Represented as pixel g i and reference pixel g j If the feature embedding is g, then g i Predicted labels Depend on Given, where reference indices j and k span the time history prior to the current frame, and the loss is calculated using the standard cross-entropy loss across all pixels in the target frame. Optimize the appearance embedding;

[0030] C6: Finally, the image segmentation network outputs a video sequence. The segmentation of objects in the image is represented by the label. and the object segmentation label of the previous frame is represented as As mentioned above, the above-mentioned and Describe the module operation steps, because and The operation is exactly the same for the input video sequence. This module also produces similar output.

[0031] As a preferred technical solution of this application, step S4 specifically includes the following implementation steps:

[0032] D1: The video image segmentation network is implemented in parallel real-time computation on two GPUs with identical parameter configurations, represented as... and Input video sequences separately and Image segmentation, video sequence Input A calculation module Perform image segmentation operations; similarly, video sequences Input B calculation module Perform image segmentation operations;

[0033] D2: The video image segmentation network is parallel, i.e., it includes... and Furthermore, each computing module also contains two structurally identical segmentation modules that process the two immediately preceding and following frames of the input video sequence, respectively, as shown below. and Right now {v 2t-3 v 2t-1}, for the input image frame v 2t-3 Video image segmentation module The output object segmentation prediction label is represented as M. 2t-3 ;against {v 2t-2 v 2t The input image frame v 2t-2 Module Its object segmentation prediction label is represented as M 2t-2 ;

[0034] D3: For the calculation module of branch A Input video sequence The currently observed image frame v 2t-1 and image frame v during the adjacent last timestamp (2t-3) 2t-3 Calculated object segmentation prediction label M 2t-3(Obtained from "Step Two" in the previous step), is forward-propagated to the video segmentation network. To calculate the current object segmentation prediction label M 2t-1 Synchronous, for the B branch computing module Input video sequence The currently observed image frame v 2t and the object segmentation prediction label M calculated during the adjacent last timestamp (2t-2). 2t-2 (Obtained from "Step Two" in the previous step), is forward-propagated to the video segmentation network. To calculate the current object segmentation prediction label M 2t Then output M 2t-3 M 2t-1 and M 2t-2 M 2t ;

[0035] D4: Based on the output of M in the previous step 2t-3 M 2t-1 In consecutive frames v 2t-3 and v 2t-1 Local registration is performed between them to calculate the initial pose. To this end, a correspondence operation is performed between the keyframes detected on each image to achieve visual feedback. In the inference phase, the image frame v with the input timestamp (2t-3) is used. 2t-3 Object segmentation prediction label M 2t-3 And input for the new observation timestamp (2t-1) image frame v 2t-1 Object segmentation prediction label M 2t-1 Then corresponding to M respectively 2t- 3. Output n key points Corresponding to M 2t-1 Output n key points And further give M 2t-3 Feature descriptors and corresponding M 2t-1 Feature descriptors Where n = 300, further through Calculate the initial pose, where It is the best sampling correlation; then output

[0036] D5: Based on the output of step three, M 2t-2 M 2t In consecutive frames v 2t-2 and v 2t Local registration is performed between them to calculate the initial pose. To this end, a correspondence operation is performed between the keyframes detected on each image to achieve visual feedback. In the inference phase, the image frame v with the input timestamp (2t-2) is used. 2t-2 Object segmentation prediction label M 2t-2 And input for the new observation timestamp (2t) image frame v 2t Object segmentation prediction label M 2t Then corresponding to M respectively 2t- 2. Output n key points Corresponding to M 2t Output n key points And further give M 2t-2 Feature descriptors and corresponding M 2t Feature descriptors Where n = 300, further through Calculate the initial pose, where It is the best sampling correlation, and then output.

[0037] As a preferred technical solution of this application, step S5 specifically includes the following implementation steps:

[0038] E1: The initial pose of the object obtained from the dual-path parallel output based on the aforementioned algorithm steps, fused together. and based on For the latest current frame Further adjustments and optimizations yielded the result used to initialize the current node P. 2t As part of the pose graph optimization step, select no more than [number missing] pose graphs from the previously stored pose graph library. Each keyframe participates in the optimization. The choice is made to balance the trade-off between efficiency and accuracy;

[0039] E2: The edges of the pose graph in the pose graph library include feature and geometric correspondences. These correspondences are matched in parallel on the GPU. Based on this information, the pose graph state outputs a refined and optimized current timestamp P online. 2t Improved spatiotemporally consistent pose ∈SO(3), where SO(3) represents a 3D orthonormal space;

[0040] E3: If the latest frame corresponds to a new pose feature (if the similarity threshold in the pose graph library is less than 0.3, it is judged as a new pose feature), then it is also included in the pose graph library. The associated pose graph outputs a finely optimized 6D object pose at the current moment to achieve object motion tracking.

[0041] As a preferred technical solution of this application, in steps S1-S5, the backbone network adopts DenseNet, and the weights are initialized using tree-type energy loss based on sparse annotation semantic segmentation.

[0042] As a preferred technical solution of this application, the S1-S5 steps adopt high-speed CUDA programming, which significantly improves online inference efficiency and greatly shortens the running time. The dual-channel Photoneo 3D camera provides 13Hz grayscale and depth images, and the dual-channel 3080 GPU inference channel achieves efficient real-time optimal performance of 26Hz through parallel computation of 13Hz sampling cross-interpolation of each Photoneo 3D camera by the dual-channel 3080 GPU inference channel. The two Photoneo 3D cameras include A, B, and B camera coordinate systems are transformed to the A camera coordinate system through T coordinate transformation to achieve pose estimation normalization.

[0043] (III) Beneficial Effects

[0044] Compared with existing technologies, this invention provides a six-degree-of-freedom visual feedback real-time motion tracking method, which has the following beneficial effects: This six-degree-of-freedom visual feedback real-time motion tracking method solves the problem of real-time and accurate motion tracking of 6D object poses in intelligent robots. Regarding the consistency and production effectiveness of intelligent robot operations in video sequences from a world coordinate system perspective, most existing methods require a computer-aided design model of the target object, which can be used for offline training or template matching during online sessions. This invention overcomes these limitations, effectively solving the problem of real-time and accurate motion tracking of 6D object poses in intelligent robots in digital factory scenarios. It is implemented using a visual feedback dual-path parallel computing module. This video frame image segmentation network performs image segmentation on synchronously input video sequences, achieving significant improvements over existing state-of-the-art methods. It boasts a significantly competitive runtime compared to other similar methods, addressing the real-time requirements for 6D object pose and motion tracking control in intelligent robots. The innovative algorithm uses DenseNet as its backbone network, initializes weights with tree-based energy loss, and employs semantic segmentation based on sparse annotation. Templates and target video frames are sampled from the original frames of the crawled object dataset. During model training, it takes 50 hours to complete on four NVIDIA 3080 GPUs. Online inference utilizes dual-GPU inference, with each inference object sequence running at a speed of 13Hz. The dual-GPU parallel computation outputs the results. The object sequence runs at a speed of 26Hz, meaning it outputs 26 frames per second in real time, which fully meets the requirements for generation and deployment. This method employs high-speed CUDA programming, significantly improving online inference efficiency and greatly reducing runtime. The new method is significantly superior to other existing advanced class-level 6D object pose estimation and tracking methods. Compared to other state-of-the-art methods that rely on object computer-aided design models, this method reduces the need for auxiliary information while maintaining performance advantages. The new method employs: dual Photoneo 3D cameras providing 13Hz grayscale and depth images, and dual 3080 GPU inference channels. The 13Hz sampling cross-interpolation of each Photoneo 3D camera is handled by the dual 3080 GPU inference channels. Parallel computing across multiple channels achieves high-efficiency real-time performance of 26Hz. Two Photoneo 3D cameras, including camera A and camera B, are transformed from their coordinate systems to camera A's coordinate system using the T-coordinate transformation to normalize pose estimation. This method employs learning-based keypoint detection and matching for initial coarse pose estimation, followed by pose graph optimization, effectively achieving spatiotemporal consistency in pose output. Since long videos can span hundreds or more frames, we sample a small number of frames to observe temporal redundancy in the video. Nine frames are sampled from the 50 frames preceding the current frame in the video sequence, including three frames before the current frame used to model near-field object movement. Then, three frames are uniformly sampled from the preceding 17 frames, and three frames are randomly sampled from the next 30 frames.This method effectively balances efficiency and effectiveness. We train the embedding model using DenseNet, adding a residual network and a 3D convolutional layer, mapping to a 512-dimensional embedding representation. We employ random flipping and cropping in the current input frame. During 6D object pose estimation and tracking, features are extracted at the original image resolution. We establish a dense remote interaction mode for video object segmentation based on single-sample supervised learning of single-frame labeled data. Video frames are streamed sequentially, enabling the model to operate online. Real-time frame inference should not rely on future frames, and we achieve effective similarity measurement between pixels, optimizing the similarity metric. Since the quality of video object segmentation depends on the similarity metric, it should consider both global high-level semantics and local low-level spatial continuity. The similarity metric in this application includes appearance terms and... In the spatial term, calculating the similarity matrix on all previous frames is computationally infeasible. Compared to the aggregated global model, typical keyframes are effectively sampled, achieving dense association between multi-point pairs of data and the latest observed frame, significantly shortening runtime and reducing tracking drift. In a dataset of common grasping objects for intelligent robots in a digital factory scenario, under the "5° 5cm" metric, the best accuracy of existing methods is improved from 85.5% to 86.3%. During the online inference phase, the real-time output reaches a maximum frame rate of 26Hz, significantly outperforming the 10Hz maximum frame rate of other methods. Therefore, the performance of this new method is state-of-the-art, and it also has significant advantages compared to other existing methods that use category-level computer-aided design models for training. The entire six-DOF visual feedback real-time motion tracking method has a simple structure, is easy to operate, and performs better than traditional methods. Detailed Implementation

[0045] A six-DOF visual feedback real-time motion tracking method includes the following steps:

[0046] S1: Construct an image segmentation framework and constraint functions, and optimize the model by minimizing the functions;

[0047] S2: Real-time dual-channel binary grouping data processing for target object segmentation;

[0048] S3: Video object image segmentation method;

[0049] S4: Design of a dual-output method for image segmentation results and key point detection and matching using a dual-parallel network to achieve visual feedback;

[0050] S5: Based on the initial pose of the object output by the dual parallel output above, the pose graph is associated with the finely optimized 6D pose of the object at the current moment to achieve object motion tracking.

[0051] Furthermore, step S1 specifically includes the following implementation steps:

[0052] A1: Define a time-series image segmentation dataset from a video series. The video dataset consists of a large number of consecutive image frames arranged in temporal order, represented as follows: This represents a video dataset, where the subscript I indicates that the set elements consist of image frames, and g i Let i = 1, 2, ..., n represent n consecutive frames in the video. These represent the background and foreground object segmentation labels for the corresponding image frames. The dataset contains... 1 tag data pair and One unlabeled data point;

[0053] A2: Based on all provided image segmentation datasets Inferring unlabeled image data Object segmentation labels The process of inferring segmentation labels for unlabeled data can be summarized as operator representation. Where tanh is the hyperbolic tangent function, expressed as tanh(x) = (e^(x-1) / (x-1)). x -e -x ) / (e x +e -x Exponential linear units (ELU) c is a positive constant. This is called the segmentation label inference operator, where w ij Represents pixel data points (g) in two different images of a encoded video sequence. i g j The similarity between ) where ε i =∑ j w ij This represents the normalized scaling factor at image pixel i, where the hyperparameter γ is used as an operator. The weighted balance factor between the two terms;

[0054] A3: Based on a continuous image frame dataset Perform smooth constraint operations on the segmentation label inference operator. It contains two sums, where the first term is... For dataset All terms (1, 2, ..., n terms, i.e. n frames of time series images) contain labeled and unlabeled image data. Running the first term implements the smoothing constraint corresponding to the pixel. The calculation result of this run realizes that similar pixels of the gradient descent method have close or approximately the same label value.

[0055] A4: Based on dataset Implement cluster constraint operations in the segmentation label inference operator. The second item included Partially targeting video datasets The first l terms, i.e., terms 1, 2, ..., l, are image data containing already labeled segmentation tags. The operation performed on this subset of data is a cluster constraint operation, i.e., calculating... This process enables the model inference values ​​to be converted into... The values ​​are aggregated with the true label value h, and the inferred value is calculated for each corresponding pixel. The distance correction from the true label value hi is optimized by minimizing constraints, and the deviation of the inferred observation value is corrected.

[0056] A5: Iterative optimization of the segmentation label inference operator Run the following optimization process Specifically, this is achieved through a minimization algorithm, making For w ij The constructed normalized similarity matrix, where matrix V and The eigenvalue matrix is ​​obtained through iterative solution. The iterative process continues until convergence, as follows: h(0) is the initial state of the system, h(0) = [h1, h2, ..., h n ] T These are the initial observations of the labels clamped using supervised labels; Let be the result of the system at step k. Let be the result of the system at step (k+1) as follows. Where τ is an empirical value, typically ranging from 0.960 to 0.981; the advantage of this iterative algorithm is that it constructs a global model on the dense structure of unlabeled data.

[0057] Furthermore, step S2 specifically includes the following implementation steps:

[0058] B1: The input video sequence is represented in chronological order as follows: Where v i Let i = 1, 2, ..., n represent the frames in the video sequence arranged chronologically from past to present, and v n This represents the latest, current frame of the image, distinct from the frame used in step S1. The expression, This is a general video sequence dataset corresponding to the description of video image object segmentation algorithms, possessing generality, universality, and summarization; the video sequences used in this section are... This indicates that it has the specificity of its application scenario, because in this application scenario, only This single image frame contains labeled data with ground truth values, corresponding to the dataset. The first frame (g1, h1) in the data. Other versions such as v2, v3, ..., v n These are all unsupervised data, similar to typical video sequence datasets. In These items, that is, relative to a typical video sequence dataset, here Application scenarios

[0059] B2: Video sequence The frames are divided into two groups, A and B, alternating according to the order of odd and even frames. and Dataset The first frame v1 in the image is used as the starting frame and is the only image frame with ground truth labels. Therefore, it is shared by both groups A and B. The first frame v1 is... and Both are used as the starting frame of their dataset, specifically represented as as well as v1 corresponds to a general video sequence dataset The first term in the array is (g1, h1), v3, v5, ..., v 2t-1 and v2, v4, v6, ..., v 2t Corresponding to general video sequence datasets of Unsupervised data, where t∈N, For the sake of uniformity in the following explanation, let n be a positive even number, and when n is a positive odd number, the last frame of the image is aligned by interpolation;

[0060] B3: An image segmentation network implemented using a dual-GPU parallel computing module. These two computing modules are represented as follows: and Video sequences with real-time parallel alternating frame input and Image segmentation is performed synchronously, corresponding to... and The parallel output includes two sets of split results.

[0061] Furthermore, step S3 specifically includes the following implementation steps:

[0062] C1: The video image segmentation network consists of two computational modules, A and B, with identical structures. These modules are computed in parallel in real-time on two GPUs with identical parameter configurations. These two computational modules are represented as follows: and Input video sequences separately and Image segmentation is performed because and The operation is exactly the same, so we use module A for calculation. For video sequences Image segmentation follows a standard operating procedure.

[0063] C2: The algorithm runs online, using video sequences. As online input, predictions for all previous frames are determined when the current frame (2t-1) arrives, and then the predictions are approximated over time. U 1:(2t-1)→(2t+1) Let U represent the similarity matrix U constructed only between pixels up to frame (2t-1) and pixels in frame (2t+1). Since no labels are provided outside the first frame, the initial term h(0) is omitted in frame (2t+1). For time (2t+1), the above propagation process equivalently minimizes a set of smooth terms in the spatiotemporal pixel array. Where i indexes the pixel at the target time (2t+1), and j indexes the pixels in all frames before and including time (2t-1), compared to operators in typical video sequence datasets. In this context, i and j can take any values ​​from 1, 2, ..., n. Due to the strict cross-temporal nature of video frames, i can only take the latest frame, i.e., the target time (2t+1), and j can take the frame at time (2t-1) and the corresponding frame at previous times.

[0064] C3: Given the start frame of the video The label on Process the remaining frames in sequence. From iterative equations Propagate the tags to every frame;

[0065] C4: Optimization of similarity metrics, because the quality of video object segmentation depends on the similarity metric U and the similarity metric w. ij From the appearance item The spatial term Sigmod(-Canberra(loc(i), loc(j)) / ε 2 Optimize in two aspects

[0066] Where f i f j It is pixel g i g jThrough feature embedding via a convolutional neural network, loc(i) represents the spatial location of pixel i, with the spatial term controlled by the local parameter ε, where the function Sigmod(t) = 1 / (1+e^(i-1)). -t Canberra is far from

[0067] C5: Based on a 3D dilated convolutional neural network, this project learns foreground object appearance embeddings in an image in a data-driven manner. The appearance embeddings aim to capture changes caused by motion, scale, and image distortion. They are learned from training data where each frame in the video is labeled with segmented objects and object identifiers, given a target pixel g. i We treat all pixels from the previous frame as a reference and set f i and f j Represented as pixel g i and reference pixel g j If the feature embedding is g, then g i Predicted labels Depend on Given, where reference indices j and k span the time history prior to the current frame, and the loss is calculated using the standard cross-entropy loss across all pixels in the target frame. Optimize the appearance embedding;

[0068] C6: Finally, the image segmentation network outputs a video sequence. The segmentation of objects in the image is represented by the label. and the object segmentation label of the previous frame is represented as As mentioned above and Describe the module operation steps, because and The operation is exactly the same for the input video sequence. This module also produces similar output.

[0069] Furthermore, step S4 specifically includes the following implementation steps:

[0070] D1: The video image segmentation network is implemented in parallel real-time computation on two GPUs with identical parameter configurations, represented as... and Input video sequences separately kind Image segmentation, video sequence Input A calculation module Perform image segmentation operations; similarly, video sequences Input B calculation module Perform image segmentation operations;

[0071] D2: The video image segmentation network is parallel, i.e., it includes... and Furthermore, each computing module also contains two structurally identical segmentation modules that process the two immediately preceding and following frames of the input video sequence, respectively, as shown below. and Right now {v 2t-3 v 2t-1}, for the input image frame v 2t-3 Video image segmentation module The output object segmentation prediction label is represented as M. 2t-3 ;against {v 2t-2 v 2t The input image frame v 2t-2 Module Its object segmentation prediction label is represented as M 2t-2 ;

[0072] D3: For the calculation module of branch A Input video sequence The currently observed image frame v 2t-1 and image frame v during the adjacent last timestamp (2t-3) 2t-3 Calculated object segmentation prediction label M 2t-3 (Obtained from "Step Two" in the previous step), is forward-propagated to the video segmentation network. To calculate the current object segmentation prediction label M 2t-1 Synchronous, for the B branch computing module Input video sequence The currently observed image frame v 2t and the object segmentation prediction label M calculated during the adjacent last timestamp (2t-2). 2t-2 (Obtained from "Step Two" in the previous step), is forward-propagated to the video segmentation network. To calculate the current object segmentation prediction label M 2t Then output M 2t-3 M 2t-1 and M 2t-2 M 2t ;

[0073] D4: Based on the output of M in the previous step 2t-3 M 2t-1 In consecutive frames v 2t-3 and v 2t-1 Local registration is performed between them to calculate the initial pose. To this end, a correspondence operation is performed between the keyframes detected on each image to achieve visual feedback. In the inference phase, the image frame v with the input timestamp (2t-3) is used. 2t-3 Object segmentation prediction label M 2t-3 And input for the new observation timestamp (2t-1) image frame v 2t-1 Object segmentation prediction label M 2t-1 Then corresponding to M respectively 2t-3 Output n key points Corresponding to M 2t-1 Output n key points And further give M 2t-3 Feature descriptors and corresponding M 2t-1 Feature descriptors Where n = 300, further through Calculate the initial pose, where It is the best sampling correlation; then output

[0074] D5: Based on the output of step three, M 2t-2 M 2t In consecutive frames v 2t-2 and v 2t Local registration is performed between them to calculate the initial pose. To this end, a correspondence operation is performed between the keyframes detected on each image to achieve visual feedback. In the inference phase, the image frame v with the input timestamp (2t-2) is used. 2t-2 Object segmentation prediction label M 2t-2 And input for the new observation timestamp (2t) image frame v 2t Object segmentation prediction label M 2t Then corresponding to M respectively 2t-2 Output n key points Corresponding to M 2t Output n key points And further give M 2t-2 Feature descriptors and corresponding M 2t Feature descriptors Where n = 300, further through Calculate the initial pose, where It is the best sampling correlation, and then output.

[0075] Furthermore, step S5 specifically includes the following implementation steps:

[0076] E1: The initial pose of the object obtained from the dual-path parallel output based on the aforementioned algorithm steps, fused together. and based on For the latest current frame Further tuning and optimization yielded a value used to initialize the current node P2t as part of the pose graph optimization step. This value was selected from the previously stored pose graph library, with a maximum of [number missing] nodes selected. Each keyframe participates in the optimization. The choice is made to balance the trade-off between efficiency and accuracy;

[0077] E2: The edges of the pose graph in the pose graph library include feature and geometric correspondences. These correspondences are matched in parallel on the GPU. Based on this information, the pose graph state outputs a refined and optimized current timestamp P online. 2t Improved spatiotemporally consistent pose ∈SO(3), where SO(3) represents a 3D orthonormal space;

[0078] E3: If the latest frame corresponds to a new pose feature (if the similarity threshold in the pose graph library is less than 0.3, it is judged as a new pose feature), then it is also included in the pose graph library. The associated pose graph outputs a finely optimized 6D object pose at the current moment to achieve object motion tracking.

[0079] Furthermore, in steps S1-S5, the backbone network adopts DenseNet, and the weights are initialized using tree-based energy loss for semantic segmentation based on sparse annotations.

[0080] Furthermore, high-speed CUDA programming is used in steps S1-S5, which significantly improves online inference efficiency and greatly shortens runtime. The Photoneo 3D camera provides 13Hz grayscale and depth images, and the dual 3080 GPU inference channels provide 13Hz sampling cross-interpolation for each Photoneo 3D camera. The parallel computation of the dual 3080 GPU inference channels achieves efficient real-time optimal performance of 26Hz. The coordinate systems of the two Photoneo 3D cameras, including A and B, are transformed to the A camera coordinate system through the T coordinate system to achieve pose estimation normalization.

[0081] Comparative example:

[0082] Performance comparison of the proposed method with other state-of-the-art methods: OnAVOS, an online adaptation method for convolutional neural networks used in video object segmentation, enables online fine-tuning and is widely used for online introduction of target information; OSVOS achieves single-sample video object segmentation based on a fully convolutional neural network architecture, capable of sequentially transferring general semantic information learned on ImageNet to the foreground segmentation task, and finally to learning the appearance of the image; PReMVOS constructs a proposal generation, refinement, and merging process for video object segmentation; SiamRCNN performs visual tracking through re-detection; STMVOS uses spatiotemporal memory for video object segmentation; EGMN uses a plot graph memory network for video object segmentation; and KMNVOS constructs a kernelized memory network for video object segmentation. We evaluate the algorithm on a dataset of common grasping objects of intelligent robots in a constructed digital factory scenario, following standard protocols. The scoring measures the regional similarity of the algorithm's performance. The score represents the boundary accuracy. The expression represents the average of the two methods, and t / s represents the inference time per frame in seconds. A performance comparison of the methods is shown in the table below.

[0083]

[0084]

[0085] The algorithm performance comparison shows that the method in this application is superior to other existing methods, with a J&F score of 86.3%. Although the performance of EGMN is close to that of our method, EGMN relies on complicated online fine-tuning optimization. It can be seen that the method in this application runs about 30 times faster. Moreover, methods such as STMVOS, KMNVOS and PReMVOS rely heavily on additional synthetic data for pre-training, while this application does not require such cumbersome operations.

[0086] It should be noted that, in this document, relational terms such as first and second (number one, number two), etc., are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0087] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A six-degree-of-freedom visual feedback real-time motion tracking method, characterized in that: The following steps are included: S1: Construct an image segmentation framework and constraint functions, and optimize the model by minimizing the functions; S2: Real-time dual-channel binary grouping data processing for target object segmentation; S3: Video object image segmentation method; S4: Design of a dual-output method for image segmentation results and key point detection and matching using a dual-parallel network to achieve visual feedback; S5: Based on the initial object pose output by the dual parallel outputs, the pose graph is associated with and finely optimized to output the current 6D object pose, thereby achieving object motion tracking; step S1 specifically includes the following implementation steps: A1: Define a time-series image segmentation dataset from a video series. The video dataset consists of a large number of consecutive image frames arranged in temporal order, represented as follows: , Represents a video dataset, with indexes. The set of elements consists of image frames, where Representative video Frame-by-frame images These represent the background and foreground object segmentation labels for the corresponding image frames. The dataset contains... 1 tag data pair and One unlabeled data point; A2: Based on all provided image segmentation datasets Inferring unlabeled image data Object segmentation labels The process of inferring segmentation labels for unlabeled data can be summarized as operator representation. ,in That is, the hyperbolic tangent function, expressed as: Exponential linear units (ELU) , For positive integers, This is called the segmentation label inference operator, where Represents pixel data points in two different images of a encoded video sequence. The similarity between them, among which This indicates that for image pixels Normalized scaling factor, hyperparameter as an operator The weighted balance factor between the two terms; A3: Based on a continuous image frame dataset Perform smooth constraint operations on the segmentation label inference operator. It contains two sums, where the first term is... For dataset All items in Item, that is A time-series image of frames, containing labeled and unlabeled image data, is used to run the first term to implement the smoothing constraint corresponding to the pixels. The calculation result of this run realizes that similar pixels have close or approximately the same label values ​​according to the gradient descent method. A4: Based on dataset Implement cluster constraint operations in the segmentation label inference operator. The second item included Partially targeting video datasets The former Item, that is These items are image data containing already labeled segmentation tags. The operation performed on this subset of data is a cluster constraint operation, which calculates... This process enables the model to infer values. Compared with the actual label value Aggregation calculations are performed, which, specifically for each corresponding pixel, result in the inferred value. Compared with the actual label value Distance correction is achieved by minimizing constraints to optimize the inferred observations that deviate from the target value. A5: Iterative Optimization of Segmentation Label Inference Operator Run the following optimization process Specifically, this is achieved through a minimization algorithm, making For the reason The constructed normalized similarity matrix, where the matrix and The eigenvalue matrix; Iterative solution The iterative process continues until convergence, as follows: , This is the initial state of the system. These are the initial observations of the labels clamped using supervised labels; For the system in the first The system's iteration results in the 1st step are as follows: The result of the step iteration is represented as ;in The value range is 0.960~0.981; step S2 specifically includes the following implementation steps: B1: The input video sequence is represented in chronological order as follows: ,in This represents the frames of an image in a video sequence arranged chronologically from the past to the present. Represents the latest, current frame of the image; only This single image frame contains labeled data with ground truth values, corresponding to the dataset. The first frame in , Others such as This is unsupervised data; B2: Video sequence The frames are divided into two groups, A and B, alternating according to the order of odd and even frames. and Dataset The first frame in As the starting frame, it is the only image frame with ground truth labels, and it is shared by both groups A and B. This is the first frame. quilt and Both are used as the starting frame of their dataset, specifically represented as ,as well as , Corresponding to general video sequence datasets The first item , and Corresponding to general video sequence datasets of Unsupervised data, among which For the sake of the following unified explanation, let's assume... For positive even numbers, when When the value is a positive odd number, alignment is achieved by interpolating the last frame of the image. B3: An image segmentation network implemented using a dual-GPU parallel computing module. These two computing modules are represented as follows: and For video sequences with real-time parallel alternating frame input and Image segmentation is performed synchronously, corresponding to... and The parallel output includes two sets of split results; The S3 step specifically includes the following implementation steps: C1: The video image segmentation network consists of two identical computational modules, A and B, which are implemented in parallel real-time computation on two GPUs with completely identical parameter configurations. These two computational modules are represented as follows: and The video sequences were input synchronously respectively. and Perform image segmentation. and The operation is exactly the same, using module A for computation. For video sequences Image segmentation follows a standard operating procedure. C2: The algorithm runs online, using video sequences. As online input, the current frame Upon arrival, predictions for all previous frames have been determined, and then approximations are made over time. ,in This indicates that it is only built up to the first Pixels of the frame and the first Similarity matrix between pixels in a frame Since no tags are provided outside the first frame, therefore in the frame The initial term is omitted. Regarding time The above propagation process effectively minimizes a set of smoothing terms in the spatiotemporal pixel array. ,in Index target time pixels, Index in time Before and including time Pixels in all frames; C3: Given the start frame of the video The label on Process the remaining frames in order. From the iterative equation Propagate the tags to every frame; C4: Optimization of similarity metrics; the quality of video object segmentation depends on the similarity metric. Similarity measurement From the appearance item and spatial items Optimize both aspects ,in It is a pixel Feature embedding through convolutional neural networks It is a pixel Spatial location, spatial terms are determined by local parameters Control, where functions Canberra is far ; C5: Learns image foreground object appearance embeddings in a data-driven manner based on 3D dilated convolutional neural networks, where each frame in the video is labeled with segmented objects and object tags, given a target pixel. Treating all pixels from the previous frame as a reference, and Represented as pixels and reference pixel The feature embedding, then Predicted labels ,Depend on Provided, with reference index Spanning the temporal history prior to the current frame, using the standard cross-entropy loss across all pixels in the target frame. Optimize the appearance embedding; C6: Image segmentation network outputs video sequence The segmentation of objects in the image is represented by the label. And the object segmentation label of the previous frame is represented as As mentioned above, the above-mentioned and Describe the module operation steps, because and The operation is exactly the same for the input video sequence. This module also produces a similar output; step S4 specifically includes the following implementation steps: D1: The video image segmentation network is implemented in parallel real-time computation on two GPUs with identical parameter configurations, represented as... and The video sequences were input synchronously respectively. and Image segmentation, video sequence Input A calculation module Perform image segmentation operations; similarly, video sequences Input B calculation module Perform image segmentation operations; D2: The video image segmentation network is parallel, i.e., it includes... and Each computing module also contains two structurally identical segmentation modules that process the two immediately preceding and following frames of the input video sequence, respectively, as shown below. and ,Right now In For the input image frame Video image segmentation module The output object segmentation prediction label is represented as ;against In Input image frame Module ; D3: For the calculation module of branch A Input video sequence The currently observed image frame and at the last adjacent timestamp Image frame Calculated object segmentation and predicted labels It is forward-propagated to the video segmentation network. To calculate the predicted label for current object segmentation Synchronous, for the B branch computing module Input video sequence The currently observed image frame and at the last adjacent timestamp Object segmentation prediction labels calculated during the period It is forward-propagated to the video segmentation network. To calculate the predicted label for current object segmentation Then output as well as ; D4: Based on the output of the previous step In consecutive frames and Local registration is performed between them to calculate the initial pose. It performs mapping operations between keyframes detected on each image to achieve visual feedback. During the inference phase, it takes a timestamp as input. Image frame Object segmentation prediction labels And input the new observation timestamp Image frame Object segmentation prediction labels Then corresponding to Output n key points ,correspond Output n key points And further give Feature descriptors , and corresponding Feature descriptors ,in Further through Calculate the initial pose, where It is the best sampling correlation; then output ; D5: Based on the output of step D3 In consecutive frames and Local registration is performed between them to calculate the initial pose. It performs mapping operations between keyframes detected on each image to achieve visual feedback. During the inference phase, it takes a timestamp as input. Image frame Object segmentation prediction labels And input the new observation timestamp Image frame Object segmentation prediction labels Then corresponding to Output n key points ,correspond Output n key points And further give Feature descriptors , and corresponding Feature descriptors ,in Further through Calculate the initial pose, where It is the best sampling correlation, and then output. .

2. The six-degree-of-freedom visual feedback real-time motion tracking method according to claim 1, characterized in that: The S5 step specifically includes the following implementation steps: E1: The initial pose of the object obtained from the dual-path parallel output based on the aforementioned algorithm steps, fused together. and ,based on For the latest current frame Further adjustments and optimizations yielded the result used to initialize the current node. As part of the pose graph optimization step, select no more than [number missing] pose graphs from the previously stored pose graph library. Each keyframe is involved in the optimization; E2: The edges of the pose graph in the pose graph library include feature and geometric correspondences. These correspondences are matched in parallel on the GPU. The pose graph state outputs a refined and optimized current timestamp online. Improved spatiotemporally consistent pose, Represents a 3-dimensional orthogonal space; E3: If the latest frame corresponds to a new pose feature, it is also included in the pose graph library. The associated pose graph outputs a finely optimized 6D object pose at the current moment, realizing object motion tracking.

3. The six-degree-of-freedom visual feedback real-time motion tracking method according to claim 1, characterized in that: In steps S1-S5, the backbone network uses DenseNet, and the weights are initialized using tree-based energy loss and semantic segmentation based on sparse annotations.

4. The six-degree-of-freedom visual feedback real-time motion tracking method according to claim 1, characterized in that: In steps S1-S5, high-speed CUDA programming is used, with dual Photoneo 3D cameras providing 13Hz grayscale and depth images and dual 3080 GPU inference channels. The 13Hz sampling cross-interpolation of each Photoneo 3D camera is computed in parallel through the dual 3080 GPU inference channels to achieve efficient real-time optimal performance of 26Hz. The two Photoneo 3D cameras are A and B. The coordinate system of camera B is transformed to the coordinate system of camera A through the T coordinate transformation to achieve pose estimation normalization.