Stereoscopic vision measurement system based on quaternion theory
Through a stereo vision measurement system based on quaternary theory, using technologies such as image semantic segmentation, Fourier transform and Kalman filtering, the time-variability and low-frequency micro-vibration problems in stereo vision measurement are solved, and high-precision and robust posture estimation are achieved, which improves the stability and safety of autonomous driving, robot navigation and virtual reality.
Patent Information
- Application Number
- CN202510344650.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
When existing stereoscopic vision measurement technology faces problems such as time-varying, instability, low-frequency micro-vibration and non-smoothing Euler angle interpolation, it is difficult to achieve high-precision and robust attitude estimation, affecting the stability and safety of autonomous driving, robot navigation and virtual reality.
A stereoscopic visual measurement system based on quaternary theory is adopted, and through technologies such as image semantic segmentation, Fourier transform, Kalman filtering and quaternary interpolation, the pose estimation of real-time video streams is optimized, low-frequency micro-vibration is filtered out and inter-frame poses are smoothed, and quaternary numbers are used to represent the pose changes and correct them through Kalman filters.
It improves the visual quality of the video stream and the accuracy of pose estimation, reduces pose instability, ensures the real-time and security of the system, reduces calculation pressure, and improves the stability and security of the system.
Smart Images

Figure CN120279225A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine vision, and particularly relates to a stereo vision measurement system based on quaternion theory. Background Art
[0002] As an important branch in the field of artificial intelligence (AI), machine vision aims to enable visual sensors to understand and interpret the three-dimensional world like humans. Through stereo vision, the spatial depth information provided by each pixel in the image can be intuitively presented, thereby realizing the perception of three-dimensional scenes. Therefore, this technology has been widely applied in applications such as 3D reconstruction, autonomous driving, virtual reality (VR), and robot navigation. However, the time-varying and unstable nature of these scenarios pose significant challenges to the continuity and consistency of stereo vision measurement.
[0003] 1. Currently, there is a lack of a method to directly extract the time-varying pose from visual information. In the fields of autonomous driving and robot navigation: It has become normal to introduce multi-device and multi-algorithm integration, but this significantly increases the economic and space pressure. Therefore, a new method is needed to reduce the complexity, cost, and resource consumption of the system while ensuring high precision and robustness. Requirements for rapidity and real-time: The movements of intelligent driving and robot navigation are real-time. For effectiveness considerations, the delay in pose estimation has a direct impact on the effectiveness and safety of the control system. If the pose estimation times out or is inaccurate, the control system may make wrong decisions, resulting in untimely system responses and even safety accidents. Designers should continuously explore multiple directions such as lightweight algorithms, data fusion technologies, and hardware acceleration solutions to ensure that the system can complete accurate pose estimation in the shortest possible time.
[0004] 2. Currently, there is a lack of a method for filtering out low-frequency micro-vibrations generated during motion. 3D reconstruction: In the field of real-time 3D reconstruction, depth information provides the necessary distance and contour of the object being measured. If there are low-frequency micro-vibrations on-site, these vibrations will cause fluctuations in the collected depth data, resulting in the accumulation of inter-frame alignment errors, geometric shape distortion, texture blurring, and detail loss, thus leading to the instability of the measured depth data. Autonomous driving field: During the driving process of intelligent vehicles, they are inevitably affected by road unevenness, which will cause the vehicle to vibrate and shake. This not only affects the data acquisition quality of visual sensors but also may interfere with the accurate perception and real-time decision-making ability of the intelligent system regarding the surrounding environment. Effectively dealing with these motion interferences caused by the road is of great significance for improving the stability and safety of the autonomous driving system. Virtual Reality (VR) field: Emphasizing the combination of visual sensors and human body sensations, using visual devices to replace human vision to create a four-dimensional space, thus providing a highly immersive interactive experience. Human behavior and motion are more complex. Therefore, the lack of predictive filtering of random low-frequency micro-vibrations will reduce the perception and safety of VR, thereby reducing the user experience.
[0005] 3. Currently, there is a lack of a method for achieving smoothness between frames in a dynamic video stream using quaternion theory. Error accumulation in time-varying pose perception: Euler angles are composed of three independent angles used to describe rotations in three-dimensional space. When the rotation angle reaches certain specific values, two rotation axes align, resulting in the loss of one degree of freedom. During pose interpolation, the deadlock phenomenon of Euler angles will cause the rotation direction to not change continuously, resulting in a sudden jump in rotation, affecting smoothness. This problem is prone to appear during high-frequency updates or long-term interpolation processes, especially in complex motion situations, such as when the rotation angle is close to the deadlock area. Path planning and obstacle avoidance anomalies: The interpolation of Euler angles is not performed in the rotation space but in the angle space. When the inter-frame pose interpolation is not smooth, the pose update will suddenly change discontinuously. Therefore, it may cause unstable trajectory planning, resulting in problems such as direction jumps, inaccurate obstacle avoidance paths, and decision-making delays. This will not only affect the stability and accuracy of driving but also may increase the risk of collisions and even lead to traffic accidents. Summary of the Invention
[0006] In view of the deficiencies in the prior art, the present invention provides a stereo vision measurement system based on quaternion theory. By quickly constructing and deploying an algorithm model, collecting and processing real-time video stream data, optimizing the pose information of the real-time video stream, and correcting its pixel-level error with the real world, the robustness and accuracy of real-time machine vision measurement and detection are improved.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] A stereoscopic vision measurement system based on quaternion theory, characterized by comprising: an input layer, a calculation module and an output layer; the input layer obtains a real-time video stream, uses image semantic segmentation technology to separate the background of the video stream, evaluates the clarity of each object in the field of view, and selects the object with the highest clarity as the display target, and inputs it into the calculation module; the calculation module performs low-frequency micro-vibration filtering, correction of time-varying posture, and smoothing processing between adjacent frames based on quaternion theory on the video stream with the display target; the output layer connects the video stream processed by the calculation module to a human-computer interaction display screen for display.
[0009] Optionally, the input layer obtains a real-time video stream through a camera with a resolution of 2560×1440.
[0010] Optionally, the input layer evaluates the clarity of each object in the field of view through the standard deviation of blurriness, and the calculation formula of the standard deviation of blurriness is:
[0011]
[0012] In the formula, σ(x, y) represents the standard deviation of blurriness of the target point (x, y), I(x+i, y+j) represents the gray value of the pixel (i, j) near the target point (x, y), μ(x, y) represents the average gray value within the current window, and the size of the window is k×k.
[0013] Optionally, the process of the calculation module performing low-frequency micro-vibration filtering is as follows:
[0014] Perform Fourier transform on the video stream, and the calculation formula is:
[0015]
[0016] In the formula, I(x, y) is the intensity value of the original image in the spatial domain, is the spectrum of the image in the frequency domain, and (x, y) and (u, v) are the coordinates of the spatial domain and the frequency domain respectively;
[0017] Set the low-frequency filtering threshold f threshold , filter out the components with frequencies lower than the threshold, and the calculation formula is:
[0018]
[0019] In the formula, I(u, v) is the frequency domain signal after low-frequency filtering;
[0020] Convert the frequency domain signal after low-frequency filtering back to the time domain through inverse Fourier transform, and the calculation formula is:
[0021]
[0022] Wherein, I′(x, y) is the image obtained after the inverse Fourier transform.
[0023] Optionally, the process of the calculation module for correcting the time-varying attitude is as follows:
[0024] Extract the time-varying attitude of the appearance target, convert the obtained spatial position and rotation information into quaternion representation, calculate the quaternion difference between consecutive frames through spherical linear interpolation, and use a Kalman filter for correction.
[0025] Optionally, the process of the Kalman filter for correcting the quaternion difference between consecutive frames is as follows:
[0026] Use the quaternion of the previous frame and the predicted rotation increment to predict the quaternion of the current frame, and the calculation formula is:
[0027]
[0028] Wherein, represents the estimated quaternion q of the previous frame t the estimated quaternion of the predicted current frame, Δq t→t+1 represents the predicted rotation increment;
[0029] Calculate the Kalman gain K as:
[0030] K = P t+1|t ·H T ·(H·P t+1|t ·H T +R) -1 ;
[0031] Wherein, P t+1|t is the covariance matrix at the predicted time t + 1, representing the error uncertainty between the actual measurement value and the estimated value; H represents the observation matrix, which is used to transform the predicted attitude into the observation space; R is the covariance matrix of the measurement noise;
[0032] Adjust the estimated quaternion of the current frame according to the prediction error through the Kalman gain, and the calculation formula is:
[0033]
[0034] Wherein, q t+1|t+1 is the adjusted estimated quaternion of the current frame.
[0035] Optionally, the update formula of the covariance matrix P t+1|t is:
[0036] P t+1|t+1 =(I - K·H)·P t-+1|t ;
[0037] Wherein, I represents the identity matrix.
[0038] Optionally, the process of the calculation module performing smoothing between adjacent frames based on quaternion theory is as follows:
[0039] Smoothing between adjacent frames is performed using quaternion interpolation, and the smoothed quaternion is calculated using spherical linear interpolation. The calculation formula is:
[0040]
[0041] Wherein, q slerp (α) represents the smoothed quaternion, q t and q t+1 are the quaternions of adjacent frames respectively, α represents the interpolation factor, and θ represents the angle between the two quaternions.
[0042] The beneficial effects of the present invention are:
[0043] 1. By accurately estimating and correcting the dynamic attitude, the present invention can effectively smooth the video stream and reduce the attitude instability caused by factors such as vibration and occlusion. Through technologies such as quaternion interpolation and Kalman filtering, it ensures that the attitude transition between consecutive frames is natural and smooth, avoiding abrupt changes. This processing not only improves the visual quality of the video stream but also provides more accurate data support for subsequent measurement and analysis.
[0044] 2. Considering efficiency in design, the present invention combines technologies such as deep learning object detection and Kalman filtering, and can process a large amount of data in real time in an environment of multiple devices and multiple sensors. In this way, the computational pressure can be greatly reduced, and high-precision attitude estimation in real-time operation can be ensured.
[0045] 3. The present invention uses Kalman filtering and neural network optimization, enabling the system to adaptively adjust the attitude estimation according to the target attitude and environmental characteristics. Through this dynamic correction mechanism, the system can continuously optimize the attitude estimation, reduce noise interference, ensure that the final output result is accurate and reliable, and improve the overall stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is the system design flow chart of the present invention;
[0047] Figure 2 is the visualization signal analysis chart of the present invention;
[0048] Figure 3 is the visualization attitude change chart of the present invention;
[0049] Figure 4 is the quaternion optimization process chart of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0050] The present invention will now be described in further detail with reference to the accompanying drawings.
[0051] The present invention proposes a stereo vision measurement system based on quaternion theory, and its design process is as Figure 1 shown. The system is mainly divided into four parts: input layer 1, input layer 2, calculation module, and output layer. In input layer 1, a real-time video stream is obtained through a camera with a resolution of 2560×1440, which can provide high-quality video stream input; in input layer 2, the system uses image semantic segmentation technology to separate the background, evaluates the clarity of each object in the field of view, selects the object with the highest clarity as the display target, and inputs it into the calculation module; in the calculation module, five specific technologies are adopted to filter out low-frequency micro-vibrations, extract its own time-varying posture, and smooth processing between adjacent frames based on quaternion theory. These technologies specifically include methods such as Fourier transform, quaternion theory, convolutional neural network, Kalman filter, and machine learning. Through these technologies, the output robustness and accuracy in dynamic vision measurement and detection can be effectively improved, and the video stream can be optimized; finally, the output layer connects the picture to the human-computer interaction display screen, displays the optimized measurement data and reconstructed video stream information, not only improves the visual quality, but also provides more accurate real-time feedback in applications such as dynamic object tracking and pose estimation.
[0052] The stage of image semantic segmentation technology includes two key steps: target extraction and segmentation of background objects and dynamic objects. First, the task of target extraction is to distinguish the static background and dominant targets (such as people, vehicles, machinery, etc.) in the field of view from the video; then, the system evaluates the clarity of each object in the field of view through the standard deviation of blurriness. By comparing the standard deviation of blurriness in the local area of the same frame image through a 3×3 window, the clarity of the image can be effectively evaluated, and the object with the highest clarity is selected as the display target, so that the subsequent calculation module can focus on processing and analyzing this dominant target. The calculation formula is:
[0053]
[0054] where, σ(x, y) represents the standard deviation of blurriness of the target point (x, y), k = 3 represents the window size is 3×3, I(x + i, y + j) represents the gray value of the pixel (i, j) near the target point (x, y), and μ(x, y) represents the average gray value within the current window.
[0055] The system evaluates the clarity of each object in the field of view through the standard deviation of blurriness, usually quantifies based on indicators such as the edge sharpness and contrast of the object, establishes a multi-dimensional standard deviation of blurriness evaluation criterion, divides each frame of image into multiple local areas, evaluates the clarity of each area separately, and finally selects the one with the optimal clarity as the display target.
[0056] The preliminary low-frequency micro-vibration filtering method is achieved through Fourier transform, and the signal analysis diagram is as Figure 2 shown. Specifically, first, perform Fourier transform on the captured video stream to convert the signal from the time domain to the frequency domain. Its calculation formula is:
[0057]
[0058] where I(x, y) is the intensity value of the original image in the spatial domain, is the spectrum of the image in the frequency domain, and (x, y) and (u, v) are the coordinates of the spatial domain and the frequency domain respectively.
[0059] In the frequency domain, the low-frequency part usually corresponds to the stationary background or slow-changing motion in the image, while the high-frequency part contains the details of the target object and fast-changing information. Low-frequency micro-vibrations usually manifest as low-frequency components in the spectrum. Therefore, by setting a frequency threshold, the components with frequencies lower than this threshold can be filtered out, thereby effectively removing the background changes or noises caused by vibrations. Its calculation formula is:
[0060]
[0061] where f threshold is the set low-frequency filtering threshold.
[0062] After filtering out the low-frequency signal, the system then converts the processed frequency-domain signal back to the time domain through inverse Fourier transform, thereby obtaining a stable image sequence and removing vibration interference. While maintaining the image details and high-frequency changes, this method significantly reduces the low-frequency noises caused by equipment vibrations, camera jitters, etc., ensuring the accuracy and stability of target detection and time-varying pose estimation. Its calculation formula is:
[0063]
[0064] In the formula, I(u, v) is the frequency-domain signal after low-frequency filtering, and I′(x, y) is the image obtained after inverse Fourier transform.
[0065] The above method significantly reduces the low-frequency noises caused by equipment vibrations, camera jitters, etc., while maintaining the image details and high-frequency changes, ensuring the accuracy and stability of target detection and time-varying pose estimation.
[0066] The correction and feedback optimization of the time-varying pose are achieved through an improved quaternion method, and the pose change diagram is as Figure 3 shown. Through the extraction method of its own time-varying pose, the obtained spatial position and rotation information are converted into quaternion representation, and the rotation difference between two adjacent frames is calculated through spherical linear interpolation.
[0067] A quaternion is a four - dimensional mathematical object, consisting of a real part and three imaginary parts, and its mathematical expression is:
[0068] q = w + xi + yj + zk;
[0069] Among them, w represents the real part of the quaternion, x, y, z represent the imaginary parts of the quaternion, and i, j, k represent the imaginary unit.
[0070] Kalman filter predicts and updates the system state in a recursive manner, compares the difference between the predicted attitude and the actual measured attitude, dynamically adjusts the prediction result, minimizes the final attitude estimation error, can effectively reduce noise interference, and provides a more accurate estimation. In the present invention, Kalman filter is used for quaternion difference correction between consecutive frames, and the quaternion optimization process is as Figure 4 shown, and the specific steps include:
[0071] 1. Prediction step
[0072] In each frame, Kalman filter first uses the attitude estimation of the previous frame and the prediction model to predict the attitude of the current frame. The predicted quaternion is generated based on the quaternion of the previous frame and the predicted rotation change. Its calculation formula is:
[0073]
[0074] where, q t represents the quaternion estimation of the previous frame, Δq t→t+1 represents the rotation increment predicted by the system, represents the attitude of the current frame deduced according to the known estimation value of the previous frame.
[0075] 2. Update step
[0076] The Kalman filter compares the actual measured value and the predicted value, calculates the error, and uses this error to update the attitude estimation of the current frame. The core of this step is the acquisition of the Kalman gain, which can balance the weights of the prediction error and the measurement error. Its calculation formula is:
[0077] K = P t+1|t ·H T ·(H·P t+1|t ·H T + R) -1 ;
[0078] where, P t+1|t is the covariance matrix at the prediction time t + 1, representing the prediction error uncertainty, H represents the observation matrix, which is used to transform the predicted attitude into the observation space, and R is the covariance matrix of the measurement noise, reflecting the noise influence generated during the measurement process.
[0079] Through the Kalman gain, the system adjusts the quaternion estimate according to the prediction error. The updated attitude estimation formula is as follows:
[0080]
[0081] In addition, the Kalman filter also incorporates the update of the covariance matrix to reflect the uncertainty of the updated estimation error. Its calculation formula is as follows:
[0082] P t+1|t+1 =(I - K·H)·P t+1|t ;
[0083] where I represents the identity matrix.
[0084] Smoothing between frames adopts quaternion interpolation to generate a smooth rotation trajectory between two frames, avoiding the gimbal lock problem in Euler angle representation. At the same time, given the quaternions between two frames, spherical linear interpolation is used to calculate the smoothed quaternion. Its calculation formula is as follows:
[0085]
[0086] where α represents the interpolation factor and θ represents the angle between two quaternions.
[0087] At the output layer, the system presents the video reconstruction effect in a visual way according to the analysis results of the calculation module, and superimposes the feature point information to highlight the detection of key points. At the same time, the system displays the pose prediction data, shows the change trend of the camera pose through line charts or heat maps, etc., and combines time series analysis to evaluate the pose deviation under vibration conditions. In addition, the system calculates the error between the predicted pose and the true pose, and conducts accuracy analysis through statistical charts. At the same time, the frames optimized by post-processing are superimposed to intuitively compare the optimization effect, and finally provide comprehensive data support for the application of the visual sensor under vibration conditions.
[0088] The above is only the preferred implementation manner of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A stereo vision measurement system based on quaternion theory, characterized in that, Including: An input layer, a calculation module, and an output layer; The input layer obtains a real-time video stream, separates the background of the video stream using image semantic segmentation technology, evaluates the clarity of each object in the field of view, and selects the object with the highest clarity as the display target, and inputs it into the calculation module; the calculation module filters out low-frequency micro-vibrations, corrects time-varying postures, and performs smoothing processing between adjacent frames based on quaternion theory on the video stream with the display target; the output layer connects the video stream processed by the calculation module to a human-computer interaction display screen for display.
2. The three-dimensional vision measurement system based on quaternion theory according to claim 1, characterized in that: The input layer obtains a real-time video stream through a camera with a resolution of 2560×1440.
3. A stereo vision measurement system based on quaternion theory according to claim 1, characterized in that: The input layer evaluates the clarity of each object in the field of view through the standard deviation of blurriness, and the calculation formula for the standard deviation of blurriness is: Where σ(x, y) represents the standard deviation of blurriness of the target point (x, y), I(x + i, y + j) represents the gray value of the pixel (i, j) near the target point (x, y), μ(x, y) represents the average gray value within the current window, and the size of the window is k×k.
4. A stereovision measurement system based on quaternion theory as claimed in claim 1, wherein: The process of the calculation module filtering out low-frequency micro-vibrations is as follows: Perform Fourier transform on the video stream, and the calculation formula is: Wherein, I(x, y) is the intensity value of the original image in the spatial domain, is the spectrum of the image in the frequency domain, and (x, y) and (u, v) are the coordinates of the spatial domain and the frequency domain, respectively; Set the low-frequency filtering threshold f threshold , and filter out the components with frequencies lower than the threshold. The calculation formula is as follows: Where I(u, v) is the frequency-domain signal after low-frequency filtering; Convert the frequency-domain signal after low-frequency filtering back to the time domain through inverse Fourier transform, and the calculation formula is: Where I′(x, y) is the image obtained after inverse Fourier transform.
5. The three-dimensional vision measurement system based on quaternion theory according to claim 1, wherein: The process of the calculation module correcting time-varying postures is as follows: Extract the time-varying posture of the display target, convert the obtained spatial position and rotation information into quaternion representation, calculate the quaternion difference between consecutive frames through spherical linear interpolation, and use a Kalman filter for correction.
6. The three-dimensional vision measurement system based on quaternion theory according to claim 5, characterized in that: The process of the Kalman filter correcting the quaternion difference between consecutive frames is as follows: Use the quaternion of the previous frame and the predicted rotation increment to predict the quaternion of the current frame, and the calculation formula is: Wherein, represents the quaternion estimated value q of the previous frame t the predicted quaternion estimated value of the current frame, Δq t→t+1 represents the predicted rotation increment; Calculate the Kalman gain K as: K = P t+1|t ·H T ·(H·P t+1|t ·H T + R) -1 ; Where, P t+1|t is the covariance matrix at the prediction time t + 1, representing the error uncertainty between the actual measurement value and the estimated value; H represents the observation matrix, which is used to convert the predicted posture into the observation space; R is the covariance matrix of the measurement noise; Adjust the quaternion estimate value of the current frame according to the prediction error through the Kalman gain, and the calculation formula is: where q t+1|t+1 is the quaternion estimate of the adjusted current frame.
7. A stereoscopic vision measurement system based on quaternion theory according to claim 6, characterized in that: The covariance matrix P t+1|t is updated by the following formula: P t-1|t+1 = (I - K·H)·P t+I|t ; Where I represents the identity matrix.
8. A stereo vision measurement system based on quaternion theory according to claim 1, characterized in that: The process of the calculation module performing smoothing processing between adjacent frames based on quaternion theory is as follows: Perform smoothing between adjacent frames using quaternion interpolation, and calculate the smoothed quaternion using spherical linear interpolation, and the calculation formula is: Where q slerp (α) represents the smoothed quaternion, q t and q t+1 are the quaternions of adjacent frames respectively, α represents the interpolation factor, and θ represents the angle between the two quaternions.