Self-supervised end-to-end visual reconstruction method and system

Through the self-supervised end-to-end visual reconstruction method, the camera position and pixel depth data are iteratively trained using reprojection errors, and the problem of inaccurate determination of three-dimensional reconstruction in the prior art relying on manual annotation and depth information is solved, achieving efficient and accurate three-dimensional scene reconstruction.

CN120182503AActive Publication Date: 2025-06-20DOMINANT INTELLIGENT TECH (SUZHOU) CO LTD

Patent Information

Application Number
CN202510641272.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-20
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

In the prior art, three-dimensional reconstruction relies on manual annotation, which is inefficient and inaccurate in-depth information determination, limiting the application of three-dimensional reconstruction in large-scale scenarios.

Method used

The self-supervised end-to-end visual reconstruction method is adopted to build an end-to-end training model by acquiring multiple camera parameters and image data from different viewing angles, and iteratively train camera position, pixel correspondence and pixel depth data using reprojection errors to achieve accurate reconstruction of three-dimensional scenes.

Benefits of technology

There is no need for a large amount of manual data labeling, which reduces the workload and cost of data labeling, improves the efficiency and autonomy of model training, enhances the accuracy and stability of three-dimensional reconstruction, and realizes accurate reconstruction from multi-view image data to three-dimensional scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182503A_ABST
    Figure CN120182503A_ABST
Patent Text Reader

Abstract

The invention provides a self-supervised end-to-end vision reconstruction method and system, and relates to the technical field of computer vision processing. The method comprises the following steps: firstly, acquiring multi-camera parameters and different-view-angle image data, including focal length, lens distortion and main coordinate point data of each camera, and pixel point coordinates of a current frame of a first camera and a reference frame of a second camera, then constructing an end-to-end training model, and calculating a re-projection error of pixel points of the current frame and the reference frame to obtain a multi-view-angle image; and solving the parameter update quantity by using a Gaussian Newton iteration method, iteratively optimizing the camera pose, the pixel corresponding relation and the depth data, and reconstructing a three-dimensional coordinate by combining the obtained data set after the re-projection error is converged, and converting and splicing to realize three-dimensional scene reconstruction. By implementing the scheme, end-to-end self-supervised training can be realized under the condition of not depending on manual annotation, so that three-dimensional visual reproduction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision processing technology, and particularly to a self-supervised end-to-end visual reconstruction method and system. Background Art

[0002] In today's digital age, visual reconstruction technology plays an important role in multiple fields. Visual reconstruction technology is dedicated to restoring three-dimensional scene information from two-dimensional images. During the visual reconstruction process, the accuracy requirements for the dataset are getting higher and higher, and the demand for manual annotation supervision is also increasing. Therefore, there is an urgent need for an end-to-end self-supervised method.

[0003] In recent years, traditional end-to-end visual reconstruction methods have received more and more research attention. This method directly outputs three-dimensional scene information after inputting an image into a deep learning model for processing. The advantage of this method is that the operation can be accelerated and run in real time through a GPU, and it has a certain generalization performance for low-brightness or weak-texture regions.

[0004] However, traditional visual reconstruction methods usually need to rely on a large amount of manually annotated data for supervised learning to determine the correspondence between images and the three-dimensional structure of the scene. This method is not only time-consuming and laborious, but also the accuracy of annotation is difficult to guarantee, which limits the application of visual reconstruction technology in large-scale scenes. Summary of the Invention

[0005] This application provides a self-supervised end-to-end visual reconstruction method and system, which is used to achieve end-to-end self-supervised training without relying on manual annotation, and thus achieve three-dimensional visual reproduction.

[0006] In a first aspect, the present application provides a self-supervised end-to-end visual reconstruction method, which is applied to a visual reconstruction system. The method includes: S1. Obtain a plurality of camera parameters and image data captured from different perspectives covering a target image. The camera parameters include focal length data, lens distortion parameters, and principal coordinate point data. The image data at least includes current frame pixel points captured by a first camera and reference frame pixel points captured by a second camera; S2. Obtain the coordinates of the current frame pixel points of the target object captured by the first camera, and the coordinates of the reference frame pixel points corresponding to the current frame pixel points of the target object captured by the second camera from different directions; S3. Construct an end-to-end training model, which is a model that iteratively trains the pose data of multiple cameras, the correspondence between pixel points on the image data, and the initial pixel depth data by calculating the reprojection error between the current frame pixel points and the reference frame pixel points; S4. Repeat step S3 until the reprojection error tends to converge, and obtain a final pixel point correspondence data set, a final camera pose data set, and a final pixel depth data set; S5. Reconstruct three-dimensional coordinates by combining the final pixel point correspondence data set and the final pixel depth data set; S6. Transform and splice the three-dimensional coordinates by combining the final camera pose data set to reconstruct a three-dimensional scene.

[0007] By adopting the above technical solution, a self-supervised end-to-end training method is adopted, which does not require a large amount of manually labeled data, reduces the workload and cost of data labeling, and at the same time can directly learn the features and relationships required for visual reconstruction from the original image data, improving the efficiency and autonomy of model training. Make full use of the parameters of multiple cameras, including focal length data, lens distortion parameters, and principal coordinate point data, as well as image data from different perspectives, comprehensively consider various factors of camera imaging, and improve the accuracy and stability of visual reconstruction. By obtaining the coordinates of the current frame pixel points and the corresponding reference frame pixel points, and using the reprojection error to iteratively train the pose data of multiple cameras, the correspondence between pixel points, and the initial pixel depth data, the correspondence between pixel points between different frames and the depth information of each pixel point can be accurately calculated, providing an accurate data basis for three-dimensional reconstruction. Reconstruct three-dimensional coordinates based on the final pixel point correspondence data set and the final pixel depth data set, and then transform and splice the three-dimensional coordinates by combining the final camera pose data set, which can completely reconstruct a three-dimensional scene and achieve accurate reconstruction from multi-perspective image data to a three-dimensional scene.

[0008] Combined with some embodiments of the first aspect, in some embodiments, before the step of obtaining a plurality of camera parameters and image data captured from different perspectives covering a target image, it further includes: constructing a pose mapping model through a plurality of historical image data with camera pose annotations by a deep learning algorithm; inputting the image data into the pose mapping model to obtain the camera pose data corresponding to the image data.

[0009] By adopting the above technical solution, the law of pose information in historical data is mined by means of a deep learning algorithm, providing an initial and relatively accurate camera pose reference for subsequent end-to-end training models.

[0010] Combined with some embodiments of the first aspect, in some embodiments, after the steps of obtaining a plurality of camera parameters and image data taken from different perspectives covering the target image, the method further includes: after extracting visual feature information from the image data by a feature extraction algorithm, determining a high-dimensional feature matrix corresponding to the image data according to the visual feature information; inputting a plurality of image data into a correspondence determination model to obtain an initial pixel correspondence, where the correspondence determination model is pre-trained by machine learning using an image set with annotations of the correspondence between pixel points, and the initial pixel correspondence is the position correspondence of pixel points of the same target part on the current frame and the reference frame.

[0011] By adopting the above technical solution, the visual feature information helps to highlight the key parts of the image, the high-dimensional feature matrix can express the image features more comprehensively, and the correspondence determination model can initially determine the pixel correspondence based on machine learning training. These technical features cooperate with each other, laying a foundation for subsequent construction of the relationship matrix and obtaining initial pixel depth data, enhancing the reliability of pixel point correlation analysis, and improving the accuracy of depth information determination in 3D reconstruction.

[0012] Combined with some embodiments of the first aspect, in some embodiments, after the step of determining a high-dimensional feature matrix corresponding to the image data according to the visual feature information after extracting the visual feature information from the image data by a feature extraction algorithm, the method further includes: obtaining a relationship matrix by combining the initial pixel correspondence and the high-dimensional feature matrix; obtaining initial pixel depth data by operating on the relationship matrix through a pixel association network.

[0013] By adopting the above technical solution, the relationship matrix integrates the pixel correspondence and the image feature information, and the pixel association network mines the depth association law between pixels based on this matrix. The two work together to more accurately infer pixel depth information from multi-perspective image data, making up for the deficiencies of traditional methods in depth data acquisition, making the restoration in the depth dimension of 3D reconstruction more accurate, and enhancing the three-dimensional sense and realism of the entire 3D reconstruction scene.

[0014] Combined with some embodiments of the first aspect, in some embodiments, before the step of inputting a plurality of image data into a correspondence determination model to obtain an initial pixel correspondence, the method further includes: preprocessing the image data by Gaussian filtering; calculating the motion information of pixels between adjacent frames in the image data by an optical flow algorithm; obtaining the motion trend information of dynamic objects in the image and a preliminary division of dynamic and static regions through the motion information.

[0015] By adopting the above technical solution, the processed image data can avoid erroneous matching due to noise and dynamic object interference in subsequent steps such as determining the initial pixel correspondence, thereby improving the accuracy of pixel correspondence, thereby ensuring the accuracy of image fusion at each perspective during the three-dimensional reconstruction process, and making the final reconstructed three-dimensional scene clearer, more complete and without obvious defects.

[0016] In combination with some embodiments of the first aspect, in some embodiments, the step of optimizing the pose data, initial pixel correspondences and multiple initial pixel depth data of multiple cameras according to the parameter update amount specifically includes: applying the chain rule to calculate the gradient information of the reprojection error to the pose data and the initial pixel depth data respectively; and updating the pose data and the initial pixel depth data using a gradient descent algorithm in combination with the gradient information.

[0017] By adopting the above technical solutions, the chain rule accurately analyzes the changing relationship between error and data, and provides precise update direction for the gradient descent algorithm. This gradient-based optimization method can efficiently adjust the camera pose and pixel depth data during iterative training, allowing the model to quickly converge to the optimal solution, effectively improving the efficiency and accuracy of 3D reconstruction.

[0018] In combination with some embodiments of the first aspect, in some embodiments, the step of converting and splicing the three-dimensional coordinates in combination with the final camera pose data set to reconstruct the three-dimensional scene specifically includes: converting the three-dimensional positions of pixel points captured by multiple cameras in their respective camera coordinate systems into a common coordinate system in combination with the coordinate transformation matrix and the final camera pose data set to determine common coordinate data of multiple pixel points in the common coordinate system; combining adjacent pixel points in the three-dimensional space in the common coordinate data, and finally splicing out the three-dimensional reconstruction result of the entire scene.

[0019] By adopting the above technical solution, the coordinate transformation matrix realizes the unified transformation of the coordinate system according to the camera pose data, so that the pixels of different perspectives can be accurately integrated in the same coordinate system. This process effectively integrates multi-perspective information and avoids reconstruction errors caused by coordinate system differences. The final three-dimensional reconstructed scene is complete, continuous and spatially reasonably arranged, which highly restores the overall picture of the real scene.

[0020] In a second aspect, the present application provides a visual reconstruction system, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the visual reconstruction system to perform the method described in the first aspect and any possible implementation of the first aspect.

[0021] In a third aspect, the present application provides a computer-readable storage medium including instructions that, when running on a visual reconstruction system, cause the visual reconstruction system to execute the method described in the first aspect and any possible implementation manner of the first aspect.

[0022] In a fourth aspect, the present application provides a computer program product that, when running on a visual reconstruction system, causes the visual reconstruction system to execute the method described in the first aspect and any possible implementation manner of the first aspect.

[0023] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0024] 1. Since the technical means of constructing an end-to-end training model and iteratively training and optimizing the camera pose, pixel correspondence, and pixel depth data using the reprojection error are adopted, the problem of three-dimensional reconstruction relying on manual annotation and low efficiency in the prior art is effectively solved, and thus the technical effects of end-to-end self-supervision and accurately restoring the true structure and form of the target scene are achieved.

[0025] 2. Since the technical means of extracting visual feature information to construct a high-dimensional feature matrix and determining the initial pixel correspondence of the model by combining the correspondence are adopted, the problems of low reliability of pixel association analysis and inaccurate determination of depth information in the prior art are effectively solved, and thus the technical effects of improving the accuracy of depth information determination in three-dimensional reconstruction and enhancing the reliability of pixel point association analysis are achieved.

[0026] 3. Since the technical means of obtaining the initial pixel depth data through the operation of the pixel association network by combining the initial pixel correspondence and the high-dimensional feature matrix are adopted, the problem of inaccurate acquisition of pixel depth data in the prior art is effectively solved, and thus the technical effects of making the three-dimensional reconstruction more accurate in the depth dimension and enhancing the three-dimensional sense and realism of the reconstructed scene are achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a flowchart of a self-supervised end-to-end visual reconstruction method in an embodiment of the present application;

[0028] Figure 2 is a scene diagram of a self-supervised end-to-end visual reconstruction method in an embodiment of the present application;

[0029] Figure 3 is a schematic structural diagram of an entity device of a visual reconstruction system in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more of the listed items.

[0031] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and should not be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0032] For ease of understanding, the method provided in this embodiment will be described in terms of a process below. Please refer to Figure 1 , which is a schematic flowchart of an end-to-end visual reconstruction method with self-supervision in the embodiments of the present application.

[0033] S1. Obtain a plurality of camera parameters and image data captured from different perspectives covering the target image. The camera parameters include focal length data, lens distortion parameters, and principal coordinate point data. The image data includes at least the pixel points of the current frame captured by the first camera and the pixel points of the reference frame captured by the second camera;

[0034] The visual reconstruction system first needs to accurately measure the focal length data of each camera in the horizontal and vertical directions in the system. For common camera models, the focal length is a key parameter that determines the scaling ratio of camera imaging.

[0035] Specifically, the camera to be calibrated can be fixed at a stable position to ensure that the camera does not shake during the shooting process. Then place the checkerboard calibration board within the field of view of the camera. To obtain accurate focal length data, the checkerboard needs to be placed at different positions and angles. For example, it can be placed directly in front of the camera, diagonally above, diagonally below, etc., and the checkerboard can be rotated horizontally by a certain angle (such as 0 degrees, 30 degrees, 60 degrees, etc.) and vertically by a certain angle (such as 0 degrees, 15 degrees, 30 degrees, etc.). A set of images is taken at each position and angle. At each position and angle where the checkerboard is placed, use the camera to take images and analyze the taken images using a professional calibration algorithm. This algorithm first identifies the corner points of the checkerboard. The corner points are the points where the black and white squares of the checkerboard pattern intersect, and these points have obvious features in the image, such as sudden changes in gray values, etc. The calibration algorithm locates the positions of the corner points by detecting the gray value changes in the image and usually uses some image processing techniques, such as edge detection, threshold segmentation, etc. to assist in the identification of the corner points. After identifying the corner points of the checkerboard, the calibration algorithm establishes the correspondence between the corner points in the image plane and the real-world coordinates. Since the size of the squares of the checkerboard is known, for example, the side length of each square is 20 millimeters, then the coordinates of each corner point on the checkerboard in the real world can be calculated according to the layout of the checkerboard. For example, the coordinates of the upper left corner point of the checkerboard can be set as (0, 0), and then the coordinates of other corner points are determined in turn according to the square size. At the same time, in the image plane, each corner point also has corresponding pixel coordinates, such as (x, y), where x represents the column position of the corner point in the image and y represents the row position. Through the established correspondence of the corner points, the calibration algorithm can calculate the internal parameter matrix of the camera according to the mathematical model. The internal parameter matrix of the camera contains multiple parameters, including the focal lengths of the camera in the horizontal and vertical directions. The calculation of the internal parameter matrix is based on the camera imaging model, and through a series of mathematical derivations and optimization algorithms, the parameters of the internal parameter matrix are solved from the correspondence between the image coordinates and the real-world coordinates of the corner points. Extract the focal length data of the camera in the horizontal direction and vertical direction from the calculated internal parameter matrix.

[0036] In addition, distortion inevitably exists in the manufacturing process of camera lenses, mainly including radial distortion and tangential distortion. In an actual lens, radial distortion and tangential distortion often exist simultaneously. Therefore, a comprehensive distortion model needs to be established to accurately describe the distortion characteristics of the lens. The comprehensive distortion model combines the radial distortion and tangential distortion models to comprehensively consider various distortion factors of the lens. It is based on a comprehensive analysis of various error sources in the actual imaging process of the lens, recognizing that the effects of radial and tangential distortion on the image are superimposed on each other. By adding the two distortion models, the actual distortion situation during lens imaging can be more accurately described. Thus, during the calibration process, by substituting the ideal coordinates and actual coordinates of the checkerboard corner points into this comprehensive model and using an optimization algorithm to solve for each parameter, accurate lens distortion parameters can be obtained.

[0037] The principal coordinate point is the coordinate of the center of the camera imaging plane in the pixel coordinate system. It is generally regarded as the origin or reference point of the image coordinate system and is the coordinate position of the intersection of the camera optical axis and the imaging plane on the image plane.

[0038] To obtain image data at different perspectives covering the target image, a visual reconstruction system needs to arrange multiple cameras to ensure that they can capture the target object from different directions. These cameras need to be precisely positioned in space and their angles adjusted to ensure sufficient overlapping areas between different perspectives for subsequent feature matching and 3D reconstruction. Among them, the pixel points at the target position in the image data captured by the first camera are set as the current frame pixel points, and the pixel points corresponding to the target position in the image data captured by the second camera are set as the reference frame pixel points. Here, the second camera and the first camera are two cameras that capture the target object from different directions.

[0039] In some embodiments, before the step of obtaining a plurality of camera parameters and image data captured from different perspectives covering a target image, a plurality of historical image data with camera pose annotations may be collected first. These historical image data cover various pose situations of the camera under different scenarios and shooting conditions, providing rich and diverse samples for model learning. For example, in different indoor and outdoor environments, images captured by the camera from different angles and positions, and each image is accurately annotated with the position (such as three-dimensional space coordinates) and pose (such as rotation angle, pitch angle, etc.) information of the camera at the time of shooting. Then, a deep learning algorithm is used to train these historical image data to construct a pose mapping model. Deep learning algorithms (such as convolutional neural networks, etc.) can automatically learn the complex relationship between the features in the image data and the camera pose. During the training process, the input of the network is the historical image data, and the output is the corresponding camera pose annotation. By continuously adjusting the parameters in the network, the network can accurately predict the camera pose based on the input image features. When new image data needs to be processed, it is input into the trained pose mapping model, and the model outputs the camera pose data corresponding to the image data according to the learned mapping relationship. This process realizes the rapid and automatic acquisition of camera pose information without manual measurement or complex calculations.

[0040] S2. Obtain the current frame pixel point coordinates of the target object captured by the first camera, and the reference frame pixel point coordinates of the target object captured by the second camera from different directions corresponding to the current frame pixel points.

[0041] The first camera captures the current frame image of the target object. Its internal imaging sensor converts the optical signal into an electrical signal, and then obtains digital image data through analog-to-digital conversion. Subsequently, it is stored in the system in a specific image format. This process involves the coordinated work of the camera hardware and the system storage mechanism. The system analyzes the stored image format and obtains key parameters, such as the width, height, and pixel depth of the image. These parameters are the basis for calculating the pixel point coordinates. A pixel coordinate system is constructed on the current frame image, with the upper left corner of the image set as the origin. Since the image data is stored row by row, when calculating the pixel point coordinates, for any pixel point in the image, its axis coordinate is equal to its column position in the image data (counting from 0), axis coordinate is equal to the row position. For example, if the image width is 800 pixels and a pixel point is at the 300th column and 400th row, then its coordinates are (300, 400). In this way, the coordinates of all pixel points in the current frame are obtained.

[0042] The second camera captures a reference frame image of the target object from a different direction. After imaging and analog-to-digital conversion, it is stored. The system performs format parsing on it to obtain relevant parameters. The method for obtaining the pixel coordinates of the reference frame image is the same as that of the current frame, that is, first establish a pixel coordinate system with the same rules, and then calculate the coordinates of each pixel in this coordinate system according to the image storage format and the parsed parameters. For example, for a certain pixel in the reference frame image, if its column position in the image data is 500, the row position is 350, and the image width is 1000 pixels, then the pixel coordinate is (500, 350), and thus the coordinates of all pixels in the reference frame image are obtained.

[0043] S3. Construct an end-to-end training model, which is a model that iteratively trains the pose data of multiple cameras, the correspondence between pixel points on the image data, and the initial pixel depth data by calculating the reprojection error between the pixel points of the current frame and the reference frame pixel points;

[0044] After the visual reconstruction system obtains the relevant camera parameters, image data, and pixel coordinates, it starts to construct an end-to-end training model. The core of this model is to iteratively train the pose data of multiple cameras, the correspondence between pixel points on the image data, and the initial pixel depth data by calculating the reprojection error between the pixel points of the current frame and the reference frame pixel points, so as to achieve accurate reconstruction of the three-dimensional scene.

[0045] The specific training steps are as follows:

[0046] S31. Using the pixel coordinates of the current frame Calculate the inverse projection coordinates of the pixel points on the current frame in the three-dimensional coordinates , and its expression is as follows:

[0047]

[0048]

[0049]

[0050]

[0051]

[0052]

[0053]

[0054]

[0055] where and They are the focal lengths of the first camera in the and directions respectively. The first principal point coordinates of the first camera are , and represent the normalized offsets of the current frame pixel coordinates relative to the first principal coordinates in the and directions respectively. and are the camera lens distortion parameters. is an intermediate variable considering the lens distortion. is a variable that combines , and ;

[0056] S32. Calculate the projection coordinates of the inverse projection coordinates on the reference frame . Its expression is as follows:

[0057]

[0058]

[0059]

[0060] where and are the focal length data of the second camera in the and directions respectively. The second principal point coordinates of the second camera are . The d is an intermediate variable of the inverse projection coordinates considering the lens distortion;

[0061] S33. Calculate the reprojection error :

[0062]

[0063] where is the corresponding point coordinates of the current frame pixel point on the reference frame. is the second projection coordinates of the current frame pixel point on the reference frame. represents the set of valid pixel point correspondence relationships. represents the reliability degree of the correspondence relationship ;

[0064] S34. Use the Gauss-Newton iteration method to solve for the parameter update amount; optimize the pose data of multiple cameras, the initial pixel correspondence relationships, and the multiple initial pixel depth data according to the parameter update amount;

[0065] The principle of the Gauss-Newton iteration method is a non-linear optimization method based on Taylor series expansion, used to solve the least squares problem. By iteratively approaching the optimal solution, the parameter update amount is obtained.

[0066] The pose of the camera describes its position and orientation in space. When the parameter update amount is obtained, a part of it is related to the camera pose. The position of the camera can be imagined as a point in three-dimensional space. The parameter update amount indicates how much distance this point needs to move in the front-back, left-right, and up-down directions. For example, if the update amount indicates that the camera needs to move a certain distance forward, then the position coordinates of the camera in space are adjusted accordingly to make it move forward. This is done to make the position of the camera more in line with the position it should be in the actual shooting scene, so that subsequent calculations and reconstructions based on the images captured by the camera are more accurate. The orientation of the camera determines the direction in which it shoots. The parameter update amount also involves the angles by which the camera rotates around different axes (like three mutually perpendicular lines imagined to pass through the center of the camera). For example, it may be necessary for the camera to rotate a certain angle clockwise around the vertical axis, or tilt upward by a certain angle around the horizontal axis. According to these update amounts, the rotation parameters of the camera are adjusted to change the orientation of the camera. After such adjustment, the orientation of the camera can better capture the target object, providing a basis for accurate visual reconstruction.

[0067] The initial pixel correspondence refers to which pixel points in the images captured by different cameras represent points on the same object in space. When initially determining the pixel correspondence, there may be some inaccuracies, and these correspondences can be fine-tuned according to the parameter update amount. For example, a certain pixel point in one image was initially considered to correspond to a certain pixel point in another image, but the update amount obtained through calculation indicates that it should correspond more accurately to an adjacent pixel point. At this time, the correspondence between these two pixel points is adjusted to make them match more accurately. There may be large deviations in the initial judgment of the correspondence between some pixel points. The parameter update amount can be used to discover and correct these incorrect correspondences. By continuously adjusting the pixel correspondence according to the update amount, the matching of pixel points between different images becomes more precise.

[0068] The initial pixel depth data represents the distance between the surface of the object corresponding to each pixel point and the camera. The parameter update amount gives the value by which each pixel depth value should be increased or decreased. For example, the initial depth estimate of a certain pixel point may be inaccurate. If the update amount indicates that the surface of the object represented by this pixel point is actually closer to the camera, then the depth value of this pixel is correspondingly decreased. By making such adjustments to each pixel depth value, they are made closer to the true object depth. The depth data of all pixels are adjusted according to the parameter update amount, so that the depth information of the entire image more conforms to the actual scene. The optimized depth data in this way can more accurately reflect the positional relationship of the object in three-dimensional space, providing reliable data support for finally reconstructing an accurate three-dimensional scene. By performing the above optimizations on the pose data of multiple cameras, the initial pixel correspondence, and the multiple initial pixel depth data, the accuracy and reliability of the visual reconstruction system can be continuously improved, and a three-dimensional reconstruction result that more conforms to the actual scene can be gradually obtained.

[0069] In some embodiments, in this step S3, the reprojection error is a complex function of multiple variables such as pose data and initial pixel depth data. Through the chain rule, the derivative of the reprojection error with respect to these variables can be decomposed into a product form of multiple intermediate derivatives, which is convenient for calculation. In this way, the gradient information of the reprojection error with respect to each variable (i.e., pose data and initial pixel depth data) is gradually calculated. The purpose of calculating the gradient information is to understand the rate of change of the reprojection error with respect to the pose data and initial pixel depth data. The direction of the gradient indicates the direction in which the error increases fastest, and the opposite direction is the direction in which the error decreases fastest. The gradient descent algorithm is an iterative optimization algorithm. Its basic idea is to gradually update the parameters along the opposite direction of the gradient to minimize the objective function (here it is the reprojection error). In each iteration, the gradient of the objective function is calculated according to the current parameter values, and then the parameters are updated in the opposite direction of the gradient according to a certain learning rate (step size). By continuously iteratively updating the pose data and initial pixel depth data, the reprojection error gradually decreases. In this process, the camera pose will gradually be adjusted to a more accurate position and orientation, so that the projection relationship between images from different perspectives is more accurate. The initial pixel depth data will also be continuously optimized to make it closer to the true object depth value. In this way, after multiple iterations, the entire visual reconstruction model can more accurately restore the three-dimensional scene and improve the accuracy and quality of the reconstruction.

[0070] In some embodiments, after the steps of obtaining multiple camera parameters and image data captured from different perspectives covering the target image, the image data can also be processed by feature extraction algorithms to extract the visual feature information therein. These algorithms can identify key elements in the image, such as edges, textures, shapes, etc. Then, based on the extracted visual feature information, the high-dimensional feature matrix corresponding to the image data is determined. This matrix can comprehensively and detailedly describe the features of the image. It represents the feature information of each pixel point or pixel region in the form of a vector and arranges them into a matrix according to certain rules. After that, multiple image data are input into the correspondence determination model, and this correspondence determination model is pre-trained by machine learning using an image set with annotations of the correspondence between pixel points. Finally, the initial pixel correspondence obtained from the correspondence determination model refers to the position correspondence of pixel points of the same target part on the current frame and the reference frame. That is to say, it determines how the pixel points of the same target part in images from different perspectives correspond to their positions in their respective images. Such a correspondence is very important for subsequent operations such as 3D reconstruction, which can help determine the pixel correspondence between images from different perspectives, thereby better restoring the 3D scene.

[0071] In some embodiments, after obtaining the initial pixel correspondence and the high-dimensional feature matrix, the two are combined to construct a relationship matrix. Specifically, the initial pixel correspondence clarifies the pixel position correspondence of the same target part in different images (such as the current frame and the reference frame), while the high-dimensional feature matrix details the visual feature information of each pixel. By combining the feature information of these corresponding pixels in the feature matrix according to certain rules, a relationship matrix is formed. The relationship matrix integrates the position correspondence and feature information of pixels, and it can more comprehensively reflect the internal connections between pixels. Compared with the individual pixel correspondence or high-dimensional feature matrix, the relationship matrix provides a richer information framework, which helps to explore deeper relationships between pixels in the subsequent process. The pixel association network is a neural network structure specifically designed to process pixel relationship data. Based on the principle of deep learning, by learning the pixel relationship patterns in a large amount of image data, it can automatically discover the hidden association rules between pixels. This network contains multiple neuron layers and is trained through forward propagation and backpropagation algorithms. During the training process, it learns how to extract useful information from the input relationship matrix to predict pixel depth data. The constructed relationship matrix is input into the pixel association network for operation. The network processes the information in the relationship matrix according to the patterns and algorithms it has learned, and then outputs the initial pixel depth data. In this process, the network uses the pixel correspondence and feature information in the relationship matrix, through complex calculations and inferences, to estimate the depth value of the object surface corresponding to each pixel from the camera. In this way, it provides key depth information for 3D reconstruction, making the reconstructed 3D scene more accurate and realistic in the depth dimension.

[0072] In some embodiments, before the step of inputting multiple image data into the correspondence determination model to obtain the initial pixel correspondence, the image data can also be preprocessed by Gaussian filtering, and then the optical flow algorithm is used to calculate the motion information of pixels between adjacent frames in the image data. The optical flow algorithm is based on a basic assumption that the brightness of pixels in the image remains unchanged (or changes slowly) between adjacent frames. By analyzing the change in pixel brightness between adjacent frames, the motion information of pixels is calculated. It usually formulates an optimization problem with the goal of minimizing the pixel brightness difference between adjacent frames, while considering pixel motion constraints (such as smoothness of motion, etc.). According to the pixel motion information calculated by the optical flow algorithm, the motion trend of dynamic objects in the image can be judged. For example, if multiple adjacent pixel points show a displacement trend towards the upper left in consecutive frames and have similar speeds, it can be inferred that the corresponding object is moving towards the upper left. This motion trend information is very important for understanding the behavior of objects in the scene. For example, in video surveillance, it can be used to track the motion trajectory of a target object and predict its future position. Further, the pixel motion information can be used to preliminarily divide the dynamic and static regions in the image. The region composed of pixel points with a motion speed exceeding a certain threshold is divided into the dynamic region, which usually corresponds to the objects moving in the scene; while the region composed of pixel points with a motion speed lower than the threshold or almost no motion is the static region, such as the background part. This division helps to adopt different strategies for different regions in subsequent processing. For example, in visual reconstruction, the pixel correspondence and depth data of the dynamic region may need to be updated more frequently, while the static region can be updated relatively less, thereby improving the processing efficiency and accuracy.

[0073] S4. Repeat step S3 until the reprojection error tends to converge, and obtain the final pixel point correspondence dataset, the final camera pose dataset, and the final pixel depth dataset;

[0074] After each iteration (repeating step S3), the reprojection error will change. When the value of the reprojection error no longer significantly decreases after multiple iterations, or the decrease amplitude is less than a very small threshold set in advance (such as 0.001), it is considered that the reprojection error tends to converge. This means that the model has fully adjusted parameters such as the camera pose, pixel correspondence, and pixel depth, making the projection result close enough to the actual situation, and further iteration will not significantly improve the result. By continuously repeating step S3 until the reprojection error tends to converge, accurate final pixel point correspondence dataset, final camera pose dataset, and final pixel depth dataset can be obtained, and these datasets are the key data basis for subsequent reconstruction of three-dimensional coordinates and three-dimensional scenes.

[0075] S5. Reconstruct the three-dimensional coordinates by combining the final pixel point correspondence dataset and the final pixel depth dataset;

[0076] Reconstructing the three-dimensional coordinates using the final pixel point correspondence dataset and the final pixel depth dataset is mainly achieved through the following key steps: From the final pixel point correspondence dataset, the corresponding information of pixel points representing the same spatial position in different images can be obtained. According to the final pixel depth dataset, the depth value of the object surface corresponding to each pixel point can be obtained. For each camera, its camera model needs to be established, which describes how points in three-dimensional space are projected onto the two-dimensional image plane. Commonly used camera models include the internal parameters of the camera (such as focal length, principal point coordinates, etc.) and external parameters (used to describe the position and orientation of the camera in the world coordinate system). According to the camera model, the projection equation can be obtained. With the pixel point correspondence, pixel depth value, and the projection equation of the camera model, the coordinates of the three-dimensional points can be solved by simultaneously solving the equations. Through the above method, for all pixel points with corresponding relationships, using their corresponding depth values and camera models, the three-dimensional coordinates of each point in space can be gradually reconstructed, thus constructing the coordinate framework of the entire three-dimensional scene.

[0077] S6. Transform and splice the three-dimensional coordinates by combining the final camera pose dataset to reconstruct the three-dimensional scene.

[0078] The final camera pose dataset records the position and orientation angle of each camera in space when taking pictures. The position is the specific location of the camera in the entire space, and the orientation angle determines the direction in which the camera is shooting. First, a large spatial reference standard needs to be determined, just like determining a general direction and origin when drawing a map. The viewing space of one of the cameras can be used as this large reference standard, or a completely new space independent of all camera views can be set as the large reference standard. The three-dimensional coordinates obtained by different cameras are initially based on the views of each camera itself. Now, according to the position and orientation angle information of each camera, these coordinates need to be uniformly transformed into the large reference standard space determined just now. For example, a camera takes pictures from a certain position and angle and obtains some three-dimensional coordinate points. According to the position and angle of this camera, the positions of these points are adjusted to conform to the position in the large reference standard space. In this way, no matter which camera the three-dimensional coordinate points are obtained from, they can have a unified position in this large reference standard space. When the three-dimensional coordinates from different camera views are transformed into a unified space, there may be some overlapping parts because the pictures taken by different cameras may overlap. At this time, some methods are needed to remove these overlapping contents and ensure that all parts can be accurately spliced together, just like splicing Figure 1Similarly, duplicate puzzle pieces need to be removed, and then the remaining ones need to be accurately assembled. After removing duplicates and aligning them, all the three-dimensional coordinate points are integrated. These points together form a set of points representing the three-dimensional scene, just like depicting the entire scene with many small dots. In actual operation, further processing may be required for these points. For example, these points can be turned into a model with a surface, so that the appearance of the three-dimensional scene can be seen more clearly, facilitating observation and subsequent analysis. For example, through some methods, these points are turned into a three-dimensional model with a shape and a surface, and finally a complete three-dimensional scene is successfully reconstructed.

[0079] In some embodiments, for each pixel point captured by a camera, it has a corresponding three-dimensional position in its respective camera coordinate system. Using a preset coordinate transformation matrix, these three-dimensional positions are transformed. The specific calculation process involves mathematical operations such as matrix multiplication. By multiplying the camera coordinate system coordinates of the pixel points with the transformation matrix, the coordinates in the common coordinate system are obtained. In this way, for all pixel points captured by multiple cameras, their common coordinate data in the common coordinate system can be determined through this method. After obtaining the common coordinate data of all pixel points, the stitching of the three-dimensional scene begins. Adjacent pixel points in three-dimensional space are combined based on the common coordinate data. Here, adjacent pixel points can be defined according to a certain spatial distance or geometric relationship. For example, pixel points within a certain threshold range are regarded as adjacent. By connecting these adjacent pixel points, the structure of the three-dimensional scene is gradually constructed. For example, for pixel points representing the surface of an object, the combination of adjacent pixel points can form the outline and surface shape of the object. As more and more adjacent pixel points are combined, the three-dimensional structure of the entire scene becomes gradually clear. Finally, through this method, a complete three-dimensional reconstruction result is stitched together, and this result can accurately reflect the three-dimensional form and spatial layout of the target scene.

[0080] In the embodiments of the present application, due to the adoption of the technical means of constructing an end-to-end training model and iteratively training and optimizing the camera pose, pixel correspondence, and pixel depth data using the reprojection error, the problem of three-dimensional reconstruction relying on manual annotation and low efficiency in the prior art is effectively solved, and thus the technical effects of end-to-end self-supervision and accurately restoring the true structure and form of the target scene are achieved.

[0081] After combining the above content, the following further describes the method provided in this embodiment in a more specific scenario. Please refer to Figure 2 , which is a schematic diagram of a scenario of the self-supervised end-to-end visual reconstruction method in the embodiments of the present application.

[0082] In Figure 2In the scene shown, for the pixel points at specific positions in the current frame, there are corresponding pixel points on the reference frame, that is, there is a corresponding relationship between the pixel points at the same position on the reference frame and the current frame. First, the pixel points on the current frame are inverse projected from the two-dimensional image plane to the three-dimensional coordinate space through a specific algorithm, thereby obtaining a series of three-dimensional space points, and these points together form a three-dimensional point cloud. Then, the three-dimensional point cloud is projected onto the reference frame according to established rules to generate corresponding projection points. After that, the deviation between these projection points and the corresponding pixel points on the reference frame is determined, and this deviation is reflected as the difference in their positions on the reference frame image. Finally, relevant parameters or models are adjusted based on this deviation. The adjustment content covers aspects such as the camera pose and the estimation of the object position, aiming to make the projection points coincide with the corresponding pixel points on the reference frame as much as possible to improve the accuracy of scene understanding and reconstruction.

[0083] In the embodiments of the present application, due to the adoption of Figure 2 the complete technical process shown from determining the pixel correspondence relationship, inverse projecting to construct the three-dimensional point cloud, projecting to calculate the deviation to adjusting the camera pose and object position estimation based on the deviation, therefore, it effectively solves the problems of inaccurate pixel correspondence, large projection deviation, and difficulty in estimating the camera pose and object position in the prior art, and further realizes the technical effects of improving the accuracy of scene understanding, enhancing the three-dimensional reconstruction accuracy, optimizing the camera pose, enhancing the authenticity and reliability of the reconstructed scene, and achieving end-to-end self-supervised efficient visual reconstruction and accurately restoring the structure and form of the target scene.

[0084] The visual reconstruction system in the embodiments of the present invention will be described from the perspective of hardware processing. Please refer to Figure 3 which is a schematic structural diagram of an entity device of the visual reconstruction system in the embodiments of the present application.

[0085] It should be noted that Figure 3 the structure of the visual reconstruction system shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0086] As shown in Figure 3As shown, the visual reconstruction system includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 302 or the program loaded from the storage section 308 into the Random Access Memory (RAM) 303, such as executing the method described in the above embodiments. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.

[0087] The following components are connected to the I / O interface 305: an input section 306 including an audio input device, a button switch, etc.; an output section 307 including a Liquid Crystal Display (LCD), an audio output device, an indicator light, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read from it can be installed into the storage section 308 as needed.

[0088] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the Central Processing Unit (CPU) 301, various functions defined in the present invention are executed.

[0089] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings.

[0091] Specifically, the visual reconstruction system of this embodiment includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, it implements the self-supervised end-to-end visual reconstruction method provided in the above embodiment.

[0092] On the other hand, the present invention also provides a computer-readable storage medium, which may be included in the visual reconstruction system described in the above embodiment; or it may exist separately and not be assembled into the visual reconstruction system. The above storage medium carries one or more computer programs. When the above one or more computer programs are executed by a processor of the visual reconstruction system, the visual reconstruction system is enabled to implement the self-supervised end-to-end visual reconstruction method provided in the above embodiment.

[0093] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present application.

[0094] As used in the foregoing embodiments, depending on the context, the term "when" may be interpreted to mean "if" or "after" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "upon determining" or "if (the stated condition or event) is detected" may be interpreted to mean "if determined" or "in response to determining" or "when (the stated condition or event) is detected" or "in response to detecting (the stated condition or event)".

[0095] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the foregoing embodiments can be implemented by a computer program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the foregoing method embodiments. The foregoing storage medium includes various media that can store program codes, such as ROM or random access memory RAM, magnetic disks, or optical discs.

Claims

1. A self-supervised end-to-end visual reconstruction method, applied to a visual reconstruction system, characterized in that: The method comprises: S1, obtaining multiple camera parameters and capturing image data at different viewing angles covering a target image, wherein the camera parameters include focal length data, lens distortion parameters, and principal coordinate point data, and the image data includes at least current frame pixel points captured by a first camera and reference frame pixel points captured by a second camera; S2, obtaining the coordinates of the pixel points of the current frame photographed by the first camera, and the coordinates of the pixel points of the reference frame corresponding to the pixel points of the current frame photographed by the second camera from different directions; S3, constructing an end-to-end training model, wherein the end-to-end training model is a model that iteratively trains the pose data of multiple cameras, the correspondence between the pixels on the image data, and the initial pixel depth data by calculating the reprojection error between the pixels of the current frame and the pixels of the reference frame; S4, repeating step S3 until the reprojection error tends to converge, and obtaining a final pixel point correspondence dataset, a final camera pose dataset, and a final pixel depth dataset; S5, reconstructing three-dimensional coordinates by combining the final pixel point correspondence relationship dataset and the final pixel depth dataset; S6. Convert and stitch the three-dimensional coordinates in combination with the final camera pose data set to reconstruct a three-dimensional scene.

2. The method according to claim 1, characterized in that Before the steps of acquiring multiple camera parameters and capturing image data at different viewing angles covering the target image, the steps further include: A pose mapping model is constructed through a deep learning algorithm using multiple historical image data with camera pose annotations; The image data is input into the pose mapping model to obtain camera pose data corresponding to the image data.

3. The method according to claim 1, characterized in that After the steps of acquiring multiple camera parameters and capturing image data at different viewing angles covering the target image, the method further includes: After extracting visual feature information from the image data by a feature extraction algorithm, determining a high-dimensional feature matrix corresponding to the image data according to the visual feature information; Multiple image data are input into a correspondence determination model to obtain an initial pixel correspondence. The correspondence determination model is obtained by pre-training a machine learning method using an image set annotated with correspondences between pixel points. The initial pixel correspondence is the position correspondence between pixel points of the same target part in the current frame and the reference frame.

4. The method according to claim 3, characterized in that After extracting visual feature information from the image data by a feature extraction algorithm, and determining a high-dimensional feature matrix corresponding to the image data according to the visual feature information, the method further includes: Combining the initial pixel correspondence relationship with the high-dimensional feature matrix to obtain a relationship matrix; The relationship matrix is ​​operated through a pixel association network to obtain initial pixel depth data.

5. The method according to claim 3, characterized in that: Before the step of inputting the plurality of image data into the correspondence determination model to obtain the initial pixel correspondence, the method further includes: Preprocess the image data by Gaussian filtering; The optical flow algorithm is used to calculate the motion information of pixels between adjacent frames in the image data; The motion trend information of the dynamic object in the image and the preliminary division of the dynamic area and the static area are obtained through the motion information.

6. The method according to claim 1, characterized in that In constructing an end-to-end training model, the end-to-end training model is a step of iteratively training the pose data of multiple cameras, the correspondence between the pixels on the image data and the initial pixel depth data by calculating the reprojection error between the pixels of the current frame and the pixels of the reference frame, specifically including: S31, calculating the inverse projection coordinates of the pixels on the current frame in the three-dimensional coordinates based on the coordinates of the pixels of the current frame , which is expressed as follows: in and The first camera is and The focal length in the direction, the coordinates of the first principal point of the first camera are , and Respectively expressed in and The normalized offset of the pixel coordinates of the current frame relative to the first principal coordinates in the direction, and is the camera lens distortion parameter, It is an intermediate variable considering lens distortion. It is a combination , and Variables; S32, calculating the projection coordinates of the inverse projection coordinates on the reference frame , which is expressed as follows: in and The second camera is and The focal length data of the direction, the coordinates of the second principal point of the second camera are , the d is an intermediate variable of the back-projected coordinates under lens distortion; S33. Calculate reprojection error : in is the pixel of the current frame The coordinates of the corresponding points on the reference frame, is the pixel point of the current frame The second projection coordinates on the reference frame, Represents the set of valid pixel correspondences, Represents the corresponding relationship the degree of reliability; S34, using the Gauss-Newton iteration method to solve and obtain the parameter update amount; The pose data, initial pixel correspondences and multiple initial pixel depth data of multiple cameras are optimized according to the parameter update amount.

7. The method according to claim 6, characterized in that The step of optimizing the pose data, initial pixel correspondences, and multiple initial pixel depth data of multiple cameras according to the parameter update amount specifically includes: Applying the chain rule to calculate the gradient information of the reprojection error to the pose data and the initial pixel depth data respectively; The pose data and the initial pixel depth data are updated using a gradient descent algorithm in combination with the gradient information.

8. A visual reconstruction system, characterized in that: The visual reconstruction system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the visual reconstruction system to execute the method described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on a visual reconstruction system, the visual reconstruction system is caused to execute the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that When the computer program product is run on a visual reconstruction system, the visual reconstruction system is caused to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image processing method and unmanned aerial vehicle

    CN110892354A

  • Three-dimensional reconstruction method and device, equipment and readable storage medium

    CN111784842A

  • Real-time three-dimensional scene reconstruction method and device

    CN115035240A

  • Three-dimensional reconstruction method and system based on multi-view RGBD camera

    CN115115780A

  • Training method of three-dimensional scene reconstruction device for multi-camera system

    CN115619928A

Cited By

  • Robot remote guidance method, device and system

    CN121370394A

  • Robotic remote guidance method, device and system

    CN121370394B

  • A method for calibrating camera intrinsic parameters for portable optical pen coordinate measuring machines.

    CN122574111A

  • A camera intrinsic parameter calibration method for portable light pen three-coordinate measurement

    CN122574111B