Measurement detection optimization method and system based on virtual reality
By combining the processing of video images and depth data in a virtual reality environment, optimizing the basic matrix and calculating object postures, it solves the problem that traditional technology is difficult to cope with dynamic changes and complex scenes, and achieves high-precision and stable object posture detection.
Patent Information
- Application Number
- CN202510215421.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional detection optimization techniques are mostly limited to object pose estimation in static scenes, which are difficult to effectively cope with the needs of dynamic changes and multimodal sensor data fusion. Relying on static calibration and data from a single perspective, it is often difficult to cope with optimization problems in complex scenarios or large-scale scenarios, resulting in the accumulation of errors, which in turn affects the accuracy and stability of the final optimization results.
The measurement detection optimization method based on virtual reality is adopted, and the corner feature is preprocessed by collecting video images and depth data, corner point features are extracted and feature point matching is performed. The basic matrix is constructed and optimized using a random sampling consistency algorithm, and the rotation matrix and translation vector of the object are calculated in combination with the PnP algorithm, global feature extraction and weight coefficient calculation are performed to realize the fusion and conversion of the feature map, obtain the final posture of the object, and the detection results are displayed through the visual interface, and the data is stored and encrypted using the database.
It improves the accuracy and stability of object posture detection, can maintain high accuracy in complex scenarios, improves feature extraction accuracy and robustness of attitude optimization, and ensures accurate estimation of object 3D poses.
Smart Images

Figure CN120147581A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virtual reality, and particularly to a method and system for optimizing metrological detection based on virtual reality. Background Art
[0002] With the continuous development of industrial automation, robotics, and intelligent manufacturing, object pose detection and optimization have become an important part of modern industry, automation control, and precision instruments. Pose detection technology is widely used in fields such as robot positioning, part assembly on automated production lines, real-time navigation in unmanned driving systems, and interactive functions in virtual reality and augmented reality systems. Especially in the fields of electrical and mechanical engineering, accurate detection of the pose of electrical equipment or components is crucial for improving production efficiency and product quality. In these applications, virtual reality and augmented reality technologies have been gradually introduced to construct highly immersive 3D environments to assist in the precise detection of object poses. By combining the advantages of virtual reality technology and depth sensors, the accuracy and real-time performance of three-dimensional object detection can be improved, thereby promoting the development of metrological detection optimization methods.
[0003] However, traditional detection and optimization technologies are mostly limited to the pose estimation of objects in static scenes, and it is difficult to effectively meet the requirements of dynamic changes and multi-modal sensor data fusion. Moreover, relying on static calibration and data from a single perspective, it is often difficult to handle optimization problems in complex or large-scale scenes, resulting in the accumulation of errors and thus affecting the accuracy and stability of the final optimization results. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method and system for optimizing metrological detection based on virtual reality, which solves the problems that traditional detection and optimization technologies are mostly limited to the pose estimation of objects in static scenes, it is difficult to effectively meet the requirements of dynamic changes and multi-modal sensor data fusion, and relying on static calibration and data from a single perspective, it is often difficult to handle optimization problems in complex or large-scale scenes, resulting in the accumulation of errors and thus affecting the accuracy and stability of the final optimization results.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In the first aspect, the present invention provides a method for optimizing metrological detection based on virtual reality, which includes,
[0008] Collecting video images and depth data for preprocessing, calculating the gradient direction by extracting corner features, then performing feature point matching using the ORB algorithm and brute-force matching algorithm, constructing an essential matrix using the random sample consensus algorithm, and optimizing the essential matrix to obtain the optimal matrix;
[0009] Based on the optimal matrix, after calculating the rotation matrix and translation vector of the object using the PnP algorithm, global features are extracted, and after calculating the weight coefficients based on the global features, the fusion and transformation of the feature maps are performed to obtain the final pose of the object;
[0010] Error detection is performed based on the final pose, and the detection results are displayed through a visualization interface;
[0011] The database is used to store the final pose and detection results, and the AES encryption algorithm is used for encryption.
[0012] As a preferred solution of the metrology detection optimization method based on virtual reality according to the present invention, wherein: the acquisition of video images and depth data is preprocessed, and after calculating the gradient direction by extracting corner features, feature point matching is performed using the ORB algorithm and the brute-force matching algorithm, and after constructing the fundamental matrix using the random sample consensus algorithm, the optimization of the fundamental matrix is performed to obtain the optimal matrix, including:
[0013] After constructing a 3D scene to be detected in the virtual reality environment, scene calibration is performed using a VR device and a camera to obtain calibration data;
[0014] The real-time video stream of the object is collected by a camera, and the video stream is denoised;
[0015] Depth data is obtained through an RGB-D sensor, and interpolation operation is performed on the depth data;
[0016] Each frame of video image is detected using the FAST detector to obtain the feature points in the image;
[0017] Based on each feature point in the image, using the image gradient method, the gradient direction of each feature point is calculated and then rotated using the gradient direction to obtain the rotated feature points;
[0018] According to the rotated feature points, after generating a fixed window using the ORB algorithm, sampling point pairs within the window are extracted, and further, the gray value difference of the sampling point pairs is calculated using the ORB algorithm to generate binary values;
[0019] Based on the binary values, the binary values are concatenated using the bit string concatenation method to obtain a binary descriptor;
[0020] Based on the binary descriptor, the feature points of every two frames of images are matched using the brute-force matching algorithm, and the similarity between the matched feature points is calculated using the Hamming distance;
[0021] Based on the similarity metric, z feature points are selected from the similarity metric as a subset using the K-nearest neighbor algorithm, and then the fundamental matrix is constructed;
[0022] Based on the fundamental matrix, the feature points are projected using image reprojection technology to obtain the projected points;
[0023] Based on the calibration data, after extracting the camera parameters, according to the projected points, corner features and local features, the reprojection error is calculated using the Euclidean distance;
[0024] Set the selection threshold as τ, and compare the reprojection error with the threshold τ. When the reprojection error is less than or equal to the threshold τ, the point is regarded as an inlier; when the reprojection error is greater than the threshold τ, the point is regarded as an outlier;
[0025] Normalize the total number of retained inliers and the reprojection error, and divide the total number of inliers by the reprojection error to obtain the normalized ratio;
[0026] Based on the normalized ratio, use the ratio as the initial weight coefficient of the total number of inliers and the reprojection error;
[0027] Perform a linear combination based on the initial weight coefficient, the total number of inliers and the reprojection error;
[0028] Define the objective function to maximize the number of inliers and minimize the reprojection error;
[0029] Initialize the objective function, randomly select matching points as parameters through the Random Sample Consensus (RANSAC) algorithm to calculate the objective function value. During the iteration process, calculate the reprojection error of each parameter, and select the fundamental matrix with the maximum number of inliers. When the decrease value of the objective function value no longer decreases significantly, stop the iteration, and use the fundamental matrix with the maximum number of inliers and the minimum reprojection error as the optimal matrix.
[0030] As a preferred embodiment of the virtual reality-based metrology detection optimization method of the present invention, wherein: based on the optimal matrix, use the Perspective-n-Point (PnP) algorithm to calculate the rotation matrix and translation vector of the object, then extract the global features, and based on the global features, calculate the weight coefficient and perform the fusion and transformation of the feature map to obtain the final pose of the object, including:
[0031] According to the optimal matrix, use the calibration data to transform the optimal matrix to obtain the essential matrix;
[0032] Use the singular value decomposition technique to decompose the essential matrix to obtain the orthogonal matrix and the diagonal matrix;
[0033] Based on the orthogonal matrix for transformation to obtain the rotation matrix and translation vector;
[0034] Based on the rotation matrix and translation vector, use the rotation matrix and translation vector as the preliminary pose;
[0035] Based on the initial pose, the coordinates of the three-dimensional object are projected onto the two-dimensional image plane using the camera coordinate transformation method to obtain the initial feature map;
[0036] The initial feature map is convolved using the hybrid dilated convolution technique to obtain feature maps of different scales;
[0037] The pixel values in the feature maps of different scales are calculated using the histogram equalization technique, and then each pixel value is weighted using the sliding window technique to obtain the pixel value at position (o, p) in each feature map;
[0038] Based on the feature maps of different scales and the pixel values at position (o, p) in each feature map, the global features of each feature map are extracted using the global average pooling technique;
[0039] Based on the global features, a non-linear activation function is used to combine with the global features to obtain the weight coefficients of each feature map;
[0040] According to the weight coefficients, each corresponding feature map is weighted using the weighted average method to obtain weighted maps of different scales;
[0041] The weighted maps of different scales are fused using the bidirectional feature pyramid weighting technique to obtain the fused feature map;
[0042] According to the fused feature map, the fused feature map is converted into a one-dimensional vector using the flattening method;
[0043] After converting the one-dimensional vector into a 3D pose using a multi-layer perceptron, the final rotation matrix and translation vector are obtained;
[0044] The final rotation matrix and translation vector are used as the final pose.
[0045] As a preferred solution of the virtual reality-based metrology detection optimization method described in the present invention, wherein: the error detection based on the final pose includes:
[0046] After measuring the pose of the object using a laser scanner, the measurement data is used as the target pose;
[0047] Based on the target pose and the final pose, the rotation matrix error between the target pose and the final pose is calculated using the Frobenius norm;
[0048] The translation vector error between the target pose and the final pose is calculated using the Euclidean norm;
[0049] The dimension unification of the rotation matrix error and the translation vector error is performed using the scaling method;
[0050] The dimension-unified rotation matrix error and translation vector error are fused using the weighted norm method to obtain the comprehensive error;
[0051] Set the global threshold to φ, compare the comprehensive error with the global threshold φ. When the comprehensive error is greater than or equal to the global threshold φ, optimize the final pose. Otherwise, it means the pose detection passes, and the final pose after the detection passes is used as the positioning basis for the object.
[0052] As a preferred solution of the virtual reality-based metrology detection optimization method described in the present invention, wherein: the display of the detection result through the visualization interface refers to the display of the error between the final pose and the target pose of the object through the visualization interface. When the detection passes, it is prompted in green, and the first column of the rotation matrix in the final pose is used as the direction of the object towards the X-axis, the second column as the direction of the object towards the Y-axis, and the third column as the direction of the object towards the Z-axis. The position of the object in the three-dimensional space is determined by the translation vector.
[0053] As a preferred solution of the virtual reality-based metrology detection optimization method described in the present invention, wherein: the storage of the final pose and the detection result using the database refers to creating a table in the MySQL database, generating a CSV file after corresponding each row in the table to the final pose and the detection result of an object, and importing the data in the CSV file into the database for storage.
[0054] As a preferred solution of the virtual reality-based metrology detection optimization method described in the present invention, wherein: the encryption using the AES encryption algorithm refers to generating an encryption key using the AES encryption algorithm and then randomly generating an initialization vector, exporting the data in the MySQL database, and encrypting the exported CSV file using the initialization vector and the AES encryption algorithm.
[0055] In a second aspect, the present invention provides a virtual reality-based metrology detection optimization system, including
[0056] An acquisition and matching module that preprocesses the collected video images and depth data, calculates the gradient direction by extracting corner features, and performs feature point matching using the ORB algorithm and the brute-force matching algorithm;
[0057] A construction and optimization module that constructs a fundamental matrix using the random sample consensus algorithm and then optimizes the fundamental matrix to obtain an optimal matrix;
[0058] An extraction and fusion module that, based on the optimal matrix, calculates the rotation matrix and translation vector of the object using the PnP algorithm, extracts global features, calculates the weight coefficient based on the global features, and performs fusion and transformation of the feature maps to obtain the final pose of the object;
[0059] A detection and display module that performs error detection based on the final pose and displays the detection result through the visualization interface;
[0060] A storage and encryption module that stores the final pose and detection results using a database and encrypts them using the AES encryption algorithm.
[0061] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the virtual reality-based metrology detection optimization method described in the first aspect of the present invention is implemented.
[0062] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the virtual reality-based metrology detection optimization method described in the first aspect of the present invention is implemented.
[0063] The beneficial effects of the present invention are as follows: By effectively combining images and depth data, the present invention enhances the accuracy and stability of object pose detection. Through the optimization of the fundamental matrix, the optimization results can maintain high accuracy in complex scenarios. Secondly, a multi-level feature map fusion technology is adopted, and the weight coefficients of each feature map are calculated through a non-linear activation function, further improving the feature extraction accuracy. By combining the weighted average method and the bidirectional feature pyramid weighting technology, feature maps of different scales are effectively fused, ensuring the accurate extraction of global features and pose optimization, and obtaining a more accurate 3D pose of the object subsequently. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0065] Figure 1 It is a flowchart of the virtual reality-based metrology detection optimization method in Embodiment 1.
[0066] Figure 2 It is a structural diagram of the virtual reality-based metrology detection optimization system in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the drawings in the specification.
[0068] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways than those specifically described herein, and those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0069] Secondly, as used herein, "an embodiment" or "embodiment" refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The appearances of "in an embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other.
[0070] Embodiment 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides an optimization method for metrology detection based on virtual reality, including the following steps:
[0071] S1. Collect video images and depth data for preprocessing, calculate the gradient direction by extracting corner features, perform feature point matching using the ORB algorithm and the brute-force matching algorithm, construct a fundamental matrix using the random sample consensus algorithm, and then optimize the fundamental matrix to obtain the optimal matrix;
[0072] Specifically, collecting video images and depth data for preprocessing, calculating the gradient direction by extracting corner features, performing feature point matching using the ORB algorithm and the brute-force matching algorithm, constructing a fundamental matrix using the random sample consensus algorithm, and then optimizing the fundamental matrix to obtain the optimal matrix includes:
[0073] Construct a 3D scene to be detected in the virtual reality environment, and use a VR device and a camera for scene calibration to obtain calibration data;
[0074] Collect the real-time video stream of the object through a camera, and perform denoising operations on the video stream;
[0075] Obtain depth data through an RGB-D sensor, and perform interpolation operations on the depth data;
[0076] Use the FAST detector to detect each frame of the video image to obtain the feature points in the image;
[0077] Based on each feature point in the image, use the image gradient method to calculate the gradient direction of each feature point, and then use the gradient direction for rotation to obtain the rotated feature points:
[0078]
[0079] In the formula, D represents the rotated feature point, de cDenote the horizontal gradient of the c-th feature point, dy c Denote the vertical gradient of the c-th feature point, arctan() represents the arctangent function, and u represents the total number of feature points;
[0080] Based on the rotated feature points, use the ORB algorithm to generate a fixed window and then extract the sampling point pairs within the window, and further use the ORB algorithm to calculate the gray value difference of the sampling point pairs to generate binary values;
[0081] Based on the binary values, use the bit string splicing method to splice the binary values to obtain a binary descriptor;
[0082] Based on the binary descriptor, use the brute force matching algorithm to match the feature points of every two frames of images, and use the Hamming distance to calculate the similarity between the matched feature points;
[0083] The similarity calculation formula is:
[0084]
[0085] In the formula, Hamming(q 1 ,q 2 ) represents the similarity between descriptor q 1 and q 2 , β represents the length of the descriptor, q 1 (r) represents the r-th bit in descriptor q 1 , q 2 (r) represents the r-th bit in descriptor q 2 ;
[0086] Based on the similarity measure, use the K-nearest neighbor algorithm to select z feature points from the similarity measure as a subset and then construct a fundamental matrix;
[0087] Based on the fundamental matrix, use the image reprojection technique to project the feature points to obtain projected points;
[0088] Based on the calibration data, extract the camera parameters and then according to the projected points, corner features and local features, use the Euclidean distance to calculate the reprojection error:
[0089]
[0090] In the formula, E(x,x') represents the reprojection error, x represents the corner feature, x' represents the local feature, represents the projected point, represents the camera parameters, ||·|| 2 represents the square of the Euclidean norm;
[0091] According to the accuracy requirements and the quality of the data, set the selection threshold as τ. Compare the reprojection error with the threshold τ. When the reprojection error is less than or equal to the threshold τ, then regard this point as an inlier; when the reprojection error is greater than the threshold τ, then regard this point as an outlier;
[0092] Perform a normalization operation on the total number of retained inliers and the reprojection error, and divide the total number of inliers by the reprojection error to obtain the normalized ratio;
[0093] Based on the normalized ratio, use the ratio as the initial weight coefficients of the total number of inliers and the reprojection error;
[0094] Perform a linear combination based on the initial weight coefficients, the total number of inliers, and the reprojection error;
[0095] Define the objective function to maximize the number of inliers and minimize the reprojection error:
[0096]
[0097] In the formula, θ represents the objective function, arg max represents the maximization operation, Y represents the total number of retained inliers, represents the initial weight coefficient of the total number of inliers, ξ represents the initial weight coefficient of the reprojection error, and E represents the reprojection error;
[0098] Initialize the objective function. After randomly selecting matching points as parameters through the Random Sample Consensus (RANSAC) algorithm, calculate the value of the objective function. During the iteration process, calculate the reprojection error of each parameter, and select the fundamental matrix that maximizes the number of inliers. When the decrease value of the objective function value no longer decreases significantly, stop the iteration, and regard the fundamental matrix that maximizes the number of inliers and minimizes the reprojection error as the optimal matrix.
[0099] Through the preprocessing of image and depth data, noise can be effectively removed and image quality can be improved, ensuring the accurate extraction and matching of subsequent feature points. The combination of the ORB algorithm and the brute-force matching algorithm utilizes the efficient feature detection and descriptor generation of ORB, and then accurately matches feature points through brute-force matching, improving the calculation speed and matching accuracy, making the present invention applicable to application scenarios with high real-time requirements. By optimizing the fundamental matrix, incorrect matching points are effectively removed, ensuring the accuracy of the final fundamental matrix, and further improving the quality of object pose estimation and three-dimensional reconstruction. The calculation of reprojection error and inlier screening are key steps to ensure the accuracy of the optimization process. By comparing the error with the threshold, reliable inliers are selected, reducing the influence of noise and outliers, and further improving the optimization effect. Secondly, the normalized ratio calculation, as the initial weight coefficient of the total number of inliers and reprojection error, ensures the weight balance between the number of inliers and error in the objective function during the optimization process, avoiding the overdominance of a certain index, and enhancing the stability and accuracy of the optimization. Through the application of the K-nearest neighbor algorithm, the most suitable subset of feature points can be selected from the similarity metric, improving the accuracy and robustness of feature point matching.
[0100] S2. Based on the optimal matrix, use the PnP algorithm to calculate the rotation matrix and translation vector of the object, then extract the global features, and based on the global features, calculate the weight coefficient and then perform the fusion and transformation of the feature map to obtain the final pose of the object.
[0101] Specifically, based on the optimal matrix, using the PnP algorithm to calculate the rotation matrix and translation vector of the object, then extract the global features, and based on the global features, calculate the weight coefficient and then perform the fusion and transformation of the feature map to obtain the final pose of the object includes:
[0102] According to the optimal matrix, use the calibration data to transform the optimal matrix to obtain the essential matrix.
[0103] The essential matrix describes the geometric relationship between two cameras and defines the relative geometric constraints between feature points under the perspective.
[0104] The optimal matrix focuses on optimizing the matching accuracy, ensuring that the calculated transformation is optimal by maximizing the number of inliers and minimizing the reprojection error.
[0105] Use the singular value decomposition technique to decompose the essential matrix to obtain an orthogonal matrix and a diagonal matrix.
[0106] The decomposition calculation formula is:
[0107] Q = U·∑·V T
[0108] In the formula, Q represents the essential matrix, U represents the orthogonal matrix, Σ represents the diagonal matrix, and V TRepresents the transpose matrix of an orthogonal matrix;
[0109] Perform a transformation based on the orthogonal matrix to obtain a rotation matrix and a translation vector:
[0110] R = U·O·V T
[0111] In the formula, R represents the rotation matrix, O represents the constraint matrix, which can be set through geometric constraints;
[0112] The calculation formula for the translation vector is:
[0113] t = U·[0,0,1] T
[0114] In the formula, t represents the translation vector, [0,0,1] T Represents the Z-axis direction of the translation vector;
[0115] Based on the rotation matrix and the translation vector, use the rotation matrix and the translation vector as the initial pose;
[0116] Based on the initial pose, use the camera coordinate transformation method to project the coordinates of the three-dimensional object onto the two-dimensional image plane to obtain a preliminary feature map;
[0117] Use the hybrid dilated convolution technique to perform a convolution operation on the preliminary feature map to obtain feature maps of different scales;
[0118] The calculation formula for the hybrid dilated convolution is:
[0119]
[0120] In the formula, H k Represents the feature map of the k-th layer, u i Represents the feature map after convolution by the i-th convolution kernel, K i Represents the i-th convolution kernel, Conv represents the convolution operation, i represents the convolution kernel index, and n represents the total number of convolution kernels;
[0121] Use the histogram equalization technique to calculate the pixel values in the feature maps of different scales and then use the sliding window technique to weight each pixel value to obtain the pixel value at position (o,p) in each feature map;
[0122] According to the feature maps of different scales and the pixel values at position (o,p) in each feature map, use the global average pooling technique to extract the global features of each feature map:
[0123]
[0124] In the formula, A k Represents the global average value of the k-th layer feature map, S kDenote the height of the feature map of the k-th layer as \(H\), and the width as \(W\). k Denote the width of the feature map of the k-th layer as \(W\), and the filter size as \(F\). k (o, p) represents the pixel value of the feature map of the k-th layer at position (o, p);
[0125] Based on the global features, use a non-linear activation function to combine with the global features to obtain the weight coefficient of each feature map:
[0126] \(w\) k = σ(λ 2 ·ReLU(λ 1 ·A k + h 1 ) + h 2 ),
[0127] In the formula, \(w\) k represents the weight coefficient of the feature map of the k-th layer, σ represents the Sigmoid activation function, ReLU represents the ReLU activation function, λ 1 and λ 2 represent weight matrices, h 1 and h 2 represent bias terms;
[0128] According to the weight coefficients, use the weighted average method to weight each corresponding feature map to obtain weighted maps of different scales;
[0129] Use the bidirectional feature pyramid weighting technique to fuse the weighted maps of different scales to obtain a fused feature map;
[0130] The fusion formula of the bidirectional feature pyramid weighting is:
[0131]
[0132] In the formula, \(B\) represents the fused feature map, \(w\) m represents the weight coefficient of the m-th weighted map, \(F\) m represents the m-th weighted map, and \(g\) represents the total number of weighted maps;
[0133] According to the fused feature map, use the flattening method to convert the fused feature map into a one-dimensional vector;
[0134] Use a multi-layer perceptron to convert the one-dimensional vector into a 3D pose, and then obtain the final rotation matrix and translation vector;
[0135] Take the final rotation matrix and translation vector as the final pose.
[0136] Through the combination of the optimal matrix and the PnP algorithm, the rotation matrix and translation vector of an object can be calculated efficiently and accurately. The PnP algorithm optimizes feature point matching, ensures the accuracy of the transformation, and adapts to changes in dynamic scenarios. Secondly, through the calculation of the essential matrix and singular value decomposition technology, the present invention further optimizes the calculation of the object pose. By maximizing the number of inliers and minimizing the reprojection error, the accuracy of the fundamental matrix is optimized, enabling a stable and accurate estimation of the change in the object pose. By adopting the hybrid dilated convolution technology, convolution operations are performed on the preliminary feature map, thereby enhancing the ability to capture image features at different scales. By fusing feature maps at different scales using the weighted average method, important information is retained and the calculation efficiency is improved. The use of the global average pooling technology effectively extracts the global information of each feature map, thus ensuring high-precision object pose estimation. Moreover, through the bidirectional feature pyramid weighting technology, multi-scale features are weighted and fused, enhancing the effective utilization of multi-level features and the robustness of pose estimation. Combining the multi-scale features in the present invention, a more accurate estimation of the final pose of the object is obtained.
[0137] S3. Perform error detection based on the final pose and display the detection results through a visualization interface;
[0138] Specifically, error detection based on the final pose includes:
[0139] After measuring the pose of the object using a laser scanner, use the measurement data as the target pose;
[0140] Based on the target pose and the final pose, use the Frobenius norm to calculate the rotation matrix error between the target pose and the final pose;
[0141] Use the Euclidean norm to calculate the translation vector error between the target pose and the final pose;
[0142] Use the scaling method to unify the dimensions of the rotation matrix error and the translation vector error;
[0143] When the scaling method is used for dimension unification, after adjusting the adjustment factor according to the specific application environment and requirements, combine the adjustment factor with the data to achieve dimension unification;
[0144] Use the weighted norm method to fuse the rotation matrix error and the translation vector error after dimension unification to obtain the comprehensive error;
[0145] According to the actual application requirements and experimental settings, set the global threshold as φ, compare the comprehensive error with the global threshold φ. When the comprehensive error is greater than or equal to the global threshold φ, optimize the final pose. Otherwise, it indicates that the pose detection is passed, and use the final pose after passing the detection as the positioning basis for the object.
[0146] Optimizing the final pose means recalculating the optimal matrix. By recalculating the optimal matrix, the present invention can gradually approach the true object pose, thereby improving the accuracy and robustness of object positioning.
[0147] By calculating the rotation matrix error and the translation vector error, the present invention can accurately evaluate the deviation of the object pose, and fuse the errors through the weighted norm method to generate a comprehensive error value. This process effectively improves the accuracy of pose detection and ensures the balance of different error terms in the optimization. The scaling method dynamically adjusts the adjustment factor according to the actual application requirements, enabling the optimization process to adapt to different environments and accuracy requirements. By comparing the global threshold with the comprehensive error, the present invention can determine in real time whether the pose optimization has reached the required accuracy, improving the optimization efficiency.
[0148] Further, displaying the detection result through the visualization interface means displaying the error between the final pose and the target pose of the object through the visualization interface. When the detection passes, it is prompted in green, and the first column of the rotation matrix in the final pose is used as the direction of the object towards the X-axis, the second column as the direction towards the Y-axis, and the third column as the direction towards the Z-axis, and the position of the object in the three-dimensional space is determined through the translation vector.
[0149] For example, in the pose detection of a motor rotor, the rotor is regarded as a rigid body, and the rotation matrix is used to represent the rotation of the rotor relative to a fixed reference coordinate system (for example, the coordinate system of the motor housing);
[0150] The first column of the rotation matrix is represented as the direction of the X-axis after the rotor rotates;
[0151] The second column of the rotation matrix is represented as the direction of the Y-axis after the rotor rotates;
[0152] The third column of the rotation matrix is represented as the direction of the Z-axis after the rotor rotates;
[0153] Obtain the current rotation state of the rotor and generate a 3×3 rotation matrix R as follows:
[0154]
[0155] The first column represents the direction of the X-axis of the rotor in the world coordinate system after rotation;
[0156] The first column represents the direction of the Y-axis of the rotor in the world coordinate system after rotation;
[0157] The third column represents the direction of the Z-axis of the rotor in the world coordinate system after rotation;
[0158] After obtaining the direction in the world coordinate system according to the rotation matrix, translation adjustment is performed using the translation vector:
[0159] y' = R·y + t,
[0160] where y' represents the corresponding point of the rotor in the global coordinate system, and y represents the point of the rotor in the local coordinate system;
[0161] The numerical values 0.707 and -0.707 represent the magnitude and direction of the vector. Specifically, 0.707 is the approximate value of, which means that the direction vector forms a certain angle (such as 45 degrees) with the corresponding coordinate axis. Therefore, each numerical value in the matrix represents the component of the rotated coordinate axis, that is, the direction of the axis in the global coordinate system.
[0162] By standard mathematical methods, the column vectors of the rotation matrix directly represent the direction of an object in three-dimensional space, enabling the rotation state of the object to be presented clearly and intuitively. Secondly, the translation vector ensures the precise position of the object in three-dimensional space. By combining the two, the posture of the object can be completely described, which not only improves the object positioning accuracy but also provides feedback through green prompts during real-time detection to help users quickly identify and process errors.
[0163] S4. Store the final posture and detection results using a database and encrypt them using the AES encryption algorithm;
[0164] Specifically, storing the final posture and detection results using a database means creating a table in the MySQL database, corresponding each row in the table to the final posture and detection results of an object, generating a CSV file, and importing the data in the CSV file into the database for storage.
[0165] Through the efficient query and structured storage of the MySQL database, long-term storage and rapid retrieval of object posture data can be ensured, and at the same time, a solid foundation is provided for further processing and visualization of the data. Secondly, the format of the CSV file is convenient for compatibility and integration with other systems and programs.
[0166] Furthermore, encrypting using the AES encryption algorithm means generating an encryption key using the AES encryption algorithm, randomly generating an initialization vector, exporting the data in the MySQL database, and using the initialization vector and the AES encryption algorithm to encrypt the exported CSV file.
[0167] The strong encryption performance of the AES encryption algorithm and the randomness of the initialization vector ensure the high strength and unpredictability of data encryption, prevent the risks of data leakage and tampering, and the present invention encrypts sensitive information during the storage and transmission processes, avoiding unauthorized access and ensuring the confidentiality of data under any circumstances.
[0168] This embodiment also provides a virtual reality-based metrology detection optimization system, including:
[0169] An acquisition and matching module, which acquires video images and depth data for preprocessing, and calculates the gradient direction by extracting corner features, and then uses the ORB algorithm and the brute-force matching algorithm for feature point matching;
[0170] A construction and optimization module, which uses the random sample consensus algorithm to construct a fundamental matrix and then optimizes the fundamental matrix to obtain an optimal matrix;
[0171] An extraction and fusion module, based on the optimal matrix, uses the PnP algorithm to calculate the rotation matrix and translation vector of the object, then extracts the global features, and calculates the weight coefficients based on the global features to perform the fusion and transformation of the feature map to obtain the final pose of the object;
[0172] A detection and display module, which performs error detection based on the final pose and displays the detection results through a visualization interface;
[0173] A storage and encryption module, which stores the final pose and detection results using a database and encrypts them using the AES encryption algorithm.
[0174] This embodiment also provides a computer device applicable to the case of the virtual reality-based metrology detection optimization method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the virtual reality-based metrology detection optimization method proposed in the above embodiment.
[0175] The computer device may be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, carrier networks, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0176] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for optimizing measurement and detection based on virtual reality as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read-Only Memory (EPROM for short), Programmable Read-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, magnetic disks, or optical discs.
[0177] In summary, through the effective combination of images and depth data, the present invention enhances the accuracy and stability of object pose detection. By optimizing the fundamental matrix, the optimization result can maintain a high level of accuracy in complex scenarios. Secondly, the multi-level feature map fusion technology is adopted, and the weight coefficients of each feature map are calculated through a non-linear activation function, further improving the feature extraction accuracy. By combining the weighted average method and the bidirectional feature pyramid weighting technology, feature maps of different scales are effectively fused, ensuring the accurate extraction of global features and pose optimization, and thus obtaining a more accurate 3D pose of the object subsequently.
[0178] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A measurement detection optimization method based on virtual reality, characterized in that: include, Collect video images and depth data for preprocessing, extract corner features and calculate gradient directions, then use ORB algorithm and brute force matching algorithm to match feature points. Use random sampling consistency algorithm to build a basic matrix and then optimize the basic matrix to obtain the optimal matrix. Based on the optimal matrix, the PnP algorithm is used to calculate the rotation matrix and translation vector of the object, and then the global features are extracted. The weight coefficients are calculated based on the global features, and the feature maps are fused and transformed to obtain the final posture of the object. Perform error detection based on the final posture and display the detection results through a visual interface; The final posture and detection results are stored in the database and encrypted using the AES encryption algorithm.
2. The virtual reality-based measurement and detection optimization method according to claim 1, characterized in that: The video image and depth data are collected for preprocessing, and feature point matching is performed using the ORB algorithm and the brute force matching algorithm after calculating the gradient direction by extracting corner point features. The basic matrix is constructed using the random sampling consistency algorithm and then optimized to obtain the optimal matrix including: After constructing the 3D scene to be detected in the virtual reality environment, the scene is calibrated using the VR device and the camera to obtain calibration data; Collect the real-time video stream of the object through the camera and perform denoising on the video stream; Obtain depth data through the RGB-D sensor and perform interpolation on the depth data; Use the FAST detector to detect each frame of video image and obtain the feature points in the image; Based on each feature point in the image, the image gradient method is used to calculate the gradient direction of each feature point and then rotate it using the gradient direction to obtain the rotated feature point; According to the rotated feature points, the ORB algorithm is used to generate a fixed window and then the sampling point pairs in the window are extracted. The ORB algorithm is further used to calculate the gray value difference of the sampling point pairs and then generate a binary value. Based on the binary value, the binary value is concatenated using the bit string concatenation method to obtain a binary descriptor; Based on binary descriptors, a brute force matching algorithm is used to match the feature points of each two frames of images, and the similarity between the matched feature points is calculated using the Hamming distance; Based on the similarity metric, the K-nearest neighbor algorithm is used to select z feature points from the similarity metric as a subset and then construct the basic matrix; Based on the basic matrix, the feature points are projected using image reprojection technology to obtain projection points; Based on the calibration data, after extracting the camera parameters, the reprojection error is calculated using the Euclidean distance according to the projection points, corner features and local features; Set the selection threshold as τ, compare the reprojection error with the threshold τ, when the reprojection error is less than or equal to the threshold τ, the point is regarded as an internal point, when the reprojection error is greater than the threshold τ, the point is regarded as an external point; The total number of retained inliers and the reprojection error are normalized, and the total number of inliers is divided by the reprojection error to obtain a normalized ratio; Based on the normalized ratio, the ratio is used as the initial weight coefficient for the total number of inliers and the reprojection error; A linear combination is performed based on the initial weight coefficient, the total number of inliers, and the reprojection error; Define the objective function to maximize the number of inliers and minimize the reprojection error; Initialize the objective function, randomly select matching points as parameters through the random sampling consistency algorithm, and then calculate the objective function value. During the iteration process, calculate the reprojection error of each parameter, and select the basic matrix that maximizes the number of inliers. When the decrease value of the objective function value no longer decreases significantly, stop the iteration, and take the basic matrix that maximizes the number of inliers and minimizes the reprojection error as the optimal matrix.
3. The virtual reality-based measurement and detection optimization method according to claim 2, characterized in that: Based on the optimal matrix, the PnP algorithm is used to calculate the rotation matrix and translation vector of the object, and then the global features are extracted. The weight coefficients are calculated based on the global features, and then the feature maps are fused and converted to obtain the final posture of the object, including: According to the optimal matrix, the optimal matrix is transformed using the calibration data to obtain the essential matrix; Use singular value decomposition technology to decompose the essential matrix and obtain an orthogonal matrix and a diagonal matrix; Perform transformation based on the orthogonal matrix to obtain the rotation matrix and translation vector; Based on the rotation matrix and translation vector, the rotation matrix and translation vector are used as the preliminary posture; Based on the preliminary pose, the coordinates of the three-dimensional object are projected onto the two-dimensional image plane using the camera coordinate transformation method to obtain a preliminary feature map; The hybrid dilated convolution technique is used to perform convolution operations on the preliminary feature map to obtain feature maps of different scales; The histogram equalization technique is used to calculate the pixel values in the feature maps of different scales, and then the sliding window technique is used to weight each pixel value to obtain the pixel value of the position (o, p) in each feature map; According to the feature maps of different scales and the pixel values at the position (o, p) in each feature map, the global features of each feature map are extracted using the global average pooling technique; Based on the global features, a nonlinear activation function is used to combine with the global features to obtain the weight coefficient of each feature map; According to the weight coefficient, each corresponding feature map is weighted using the weighted average method to obtain weighted maps of different scales; Use bidirectional feature pyramid weighting technology to fuse weighted graphs of different scales to obtain a fused feature graph; According to the fused feature map, a flattening method is used to convert the fused feature map into a one-dimensional vector; After converting the one-dimensional vector into a 3D posture using a multi-layer perceptron, the final rotation matrix and translation vector are obtained; The final rotation matrix and translation vector are taken as the final pose.
4. The virtual reality-based measurement and detection optimization method according to claim 3, characterized in that: The error detection based on the final posture includes: The posture of the object is measured by using a laser scanner and the measurement data is used as the target posture; Based on the target pose and the final pose, the rotation matrix error in the target pose and the final pose is calculated using the Frobenius norm; The translation vector error in the target pose and the final pose is calculated using the Euclidean norm; Use scaling method to normalize the rotation matrix error and translation vector error; The weighted norm method is used to fuse the dimensionally unified rotation matrix error and the translation vector error to obtain the comprehensive error. Set the global threshold as φ, compare the comprehensive error with the global threshold φ, and when the comprehensive error is greater than or equal to the global threshold φ, the final posture is optimized. Otherwise, it means that the posture detection passes, and the final posture after the detection passes is used as the basis for object positioning.
5. The virtual reality-based measurement and detection optimization method according to claim 4, characterized in that: The display of the detection results through a visual interface refers to displaying the error between the final posture of the object and the target posture through a visual interface. When the detection is passed, green is used as a prompt, and the first column of the rotation matrix in the final posture is used as the direction of the object toward the X-axis, the second column is used as the direction of the object toward the Y-axis, and the third column is used as the direction of the object toward the Z-axis, and the position of the object in three-dimensional space is determined by the translation vector.
6. The virtual reality-based measurement and detection optimization method according to claim 5, characterized in that: The use of a database to store the final posture and detection results refers to creating a table in a MySQL database, generating a CSV file after each row in the table corresponds to the final posture and detection result of an object, and importing the data in the CSV file into the database for storage.
7. The virtual reality-based measurement and detection optimization method according to claim 6, characterized in that: The encryption using the AES encryption algorithm refers to using the AES encryption algorithm to generate an encryption key and then randomly generating an initialization vector, and after exporting the data in the MySQL database, using the initialization vector and the AES encryption algorithm to encrypt the exported CSV file.
8. A virtual reality-based metrology detection optimization system, based on the virtual reality-based metrology detection optimization method according to any one of claims 1 to 7, characterized in that: include: The acquisition and matching module acquires video images and depth data for preprocessing, and uses the ORB algorithm and brute force matching algorithm to match feature points after calculating the gradient direction by extracting corner features; The construction and optimization module uses the random sampling consensus algorithm to construct the basic matrix and then optimizes the basic matrix to obtain the optimal matrix; The extraction and fusion module uses the PnP algorithm to calculate the rotation matrix and translation vector of the object based on the optimal matrix, extracts the global features, and then calculates the weight coefficient based on the global features, fuses and transforms the feature maps to obtain the final posture of the object; The detection and display module performs error detection based on the final posture and displays the detection results through a visual interface; The storage and encryption module uses the database to store the final posture and detection results and encrypts them using the AES encryption algorithm.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the virtual reality-based metrology detection optimization method described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the virtual reality-based metrology detection optimization method according to any one of claims 1 to 7 are implemented.