A method and system for scanning intraoral tomographic images based on machine vision
By integrating an active structured light projector and an inertial measurement unit into a miniaturized optical probe, and combining an extended Kalman filter and a Markov random field model, the problems of unstable illumination, trajectory reproduction, and inconsistent slice thickness in intraoral tomographic imaging were solved, achieving high-quality tomographic slice generation and improving the system's practicality and diagnostic efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI ZHENGSHI PHOTOELECTRIC TECH CO LTD
- Filing Date
- 2026-01-22
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies for intraoral tomographic imaging suffer from problems such as unstable illumination leading to fluctuations in image signal-to-noise ratio, difficulty in accurately reproducing probe motion trajectory, depth map voids and distortion, and inconsistent sampling slice thickness, making it impossible to achieve real-time quality assessment and dynamic parameter adjustment.
A miniaturized optical probe integrating an active structured light projector, a high-speed image sensor, an inertial measurement unit, and a precision displacement mechanism is used. Combined with an extended Kalman filter and a Markov random field model, it achieves adaptive illumination compensation, pose estimation, and depth optimization, and real-time closed-loop control of layer thickness to generate high-quality tomographic slices.
It significantly improves the robustness of probe motion trajectory estimation, effectively repairs depth map holes and distortions, ensures the consistency of slice thickness and resolution, shortens clinical operation time, and improves the system's practicality and diagnostic efficiency.
Smart Images

Figure CN121549950B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical imaging technology, and in particular to a scanning method and system for intraoral tomographic images based on machine vision. Background Technology
[0002] With the continuous evolution of digital healthcare technology, medical imaging diagnosis is playing an increasingly prominent role in the clinical diagnosis and treatment system. This is especially true in the field of oral medicine, where high-precision, high-efficiency imaging technology has become a key infrastructure supporting precision treatment and personalized intervention. The oral cavity has a complex structure, including teeth, gums, mucosa, and jawbone, among other tissues. Its spatial geometry is highly irregular and exhibits significant individual differences, placing stringent demands on the resolution, depth-of-field control, anti-interference capabilities, and real-time performance of imaging systems. Traditional methods for acquiring oral images mainly include intraoral X-rays, cone-beam computed tomography (CBCT), and optical coherence tomography (OCT). However, these technologies generally suffer from inherent limitations when faced with the task of acquiring continuous tomographic images in a dynamic oral environment, such as insufficient adaptability, strong operational dependence, radiation exposure risks, or high costs. Against this backdrop, machine vision technology based on non-invasive optical sensing paths, due to its advantages of being radiation-free, having a high frame rate, low cost, and ease of integration, is gradually becoming an important direction in intraoral tomographic imaging research.
[0003] Among these, the machine vision-based intraoral tomographic imaging method focuses on using miniaturized optical probes to perform continuous motion sampling within the oral cavity. By simultaneously capturing surface texture and depth information, it constructs a volumetric dataset with tomographic resolution. The basic principle of this method is to integrate a high-speed camera unit and a precision displacement mechanism into a handheld or catheter-type probe. As the probe moves along a preset trajectory, image acquisition is triggered with a fixed step size. Combining active or passive depth sensing mechanisms such as structured light, laser triangulation, or stereo vision, a two-dimensional intensity image and its corresponding depth map are acquired at each location. Then, through registration, interpolation, and voxelization, a continuous sequence of tomographic slices is generated. Theoretically, this approach can avoid the safety hazards of traditional ionizing radiation imaging and supports real-time intraoperative feedback, providing high spatiotemporal resolution data support for clinical scenarios such as early caries detection, periodontal disease assessment, and implant planning.
[0004] Current technologies still face multiple systemic challenges in achieving intraoral tomographic image scanning based on machine vision: First, the lighting conditions inside the oral cavity are extremely unstable, with saliva reflection, tooth surface highlights, and soft tissue occlusion causing drastic fluctuations in the image signal-to-noise ratio, severely affecting the robustness of feature point extraction and matching; second, the probe's motion trajectory is difficult to reproduce accurately, and the non-uniform translation and rotation disturbances introduced by manual operation cause significant pose deviations between adjacent frames, resulting in inter-layer misalignment and structural distortion in 3D reconstruction; third, existing depth estimation algorithms are mostly based on static scene assumptions, and cannot effectively handle dynamic blur and motion artifacts during continuous probe movement, resulting in large areas of voids or distortion in the depth map at the edges; in addition, the slice thickness consistency of tomographic images lacks a closed-loop control mechanism, and the cumulative error between the sampling step size and the actual physical displacement due to mechanical hysteresis or sliding friction makes the Z-axis resolution of the final volume data uncontrollable; finally, existing systems generally adopt offline post-processing mode, which cannot evaluate image quality and dynamically adjust acquisition parameters in real time during scanning, resulting in an excessively high proportion of invalid data and prolonging clinical operation time. Summary of the Invention
[0005] The purpose of this invention is to provide a scanning method and system for intraoral tomographic images based on machine vision, so as to solve the problems of manual operation disturbance, failure in weak texture areas, depth map holes and distortion, and inconsistent sampling layer thickness in the prior art.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] A scanning method for intraoral tomographic images based on machine vision, the specific steps of which are as follows:
[0008] Step S110: A miniaturized optical probe integrating an active structured light projector, a high-speed image sensor, an inertial measurement unit, and a precision displacement mechanism moves along a preset linear trajectory inside the oral cavity, synchronously acquiring two-dimensional images modulated by structured light patterns at each sampling position, and recording the original motion data of the inertial measurement unit and the encoder feedback of the displacement mechanism.
[0009] Step S120: Adaptive illumination compensation and feature extraction are performed on the acquired two-dimensional image sequence. Combined with the inertial measurement unit data, the precise six-degree-of-freedom pose transformation matrix of the probe between continuous sampling positions is calculated by fusing the data through an extended Kalman filter.
[0010] Step S130: Based on the pose transformation matrix and the initial depth map obtained by structured light decoding, construct and solve the Markov random field energy function that incorporates motion consistency constraints, optimize and repair the holes in the initial depth map, and generate an optimized dense depth map corresponding to each sampling position.
[0011] Step S140: The optimized dense depth maps of all sampling positions are uniformly transformed to the global three-dimensional coordinate system according to their corresponding pose matrices. Through voxelization and equal-interval resampling along the main direction of probe movement, a series of continuous tomographic slice images are generated.
[0012] In step S150, during the generation process of step S140, the error between the preset theoretical layer thickness and the actual physical displacement is calculated in real time, and the motion step size of the precision displacement mechanism in the next sampling cycle is dynamically adjusted based on the error to achieve closed-loop control of the consistency of the tomographic image layer thickness.
[0013] Preferably, the adaptive illumination compensation in step S120 specifically involves: real-time monitoring of the grayscale histogram distribution of the current frame's structured light pattern; when an overexposed area (i.e., the area with a pixel grayscale value greater than 240) accounts for more than 5% or an underexposed area (i.e., the area with a pixel grayscale value less than 15) accounts for more than 10%, it is determined that there is strong reflection or insufficient illumination; the compensation strategy is, if overexposure exists, then according to the formula... Global attenuation is applied to the image, where These are the original image pixel values. The average grayscale value of the image outside the overexposed area. This represents the adjusted grayscale value; if underexposure exists, the active structured light projection unit is triggered to emit a compensation pattern with a 30% increase in intensity in the next frame, and the image sensor gain is simultaneously increased by 8dB.
[0014] Furthermore, the pose estimation in step S120 employs a tightly coupled vision-inertial odometry method. This method uses the angular velocity measured by the inertial measurement unit... and acceleration Pre-integration is performed between adjacent image frames to obtain relative rotation. Translation and speed change Simultaneously, the essential matrix of the hybrid feature points extracted from the image is calculated using epipolar geometry. The rotation matrix is obtained by decomposition. Translation vector The state vector of the extended Kalman filter includes the probe pose, velocity, and inertial measurement unit zero bias. Its observation equation is composed of visual geometric constraints and inertial pre-integration constraints. By minimizing the weighted sum of reprojection error and inertial measurement error, it outputs the optimal probe pose estimate in real time.
[0015] Furthermore, the energy function of the Markov random field constructed in step S130 is defined as:
[0016]
[0017] in, The depth map to be solved is shown below. and For pixels, The set of all pixels. This is a four-neighborhood system. (Data item) It consists of two items:
[0018]
[0019] in, The initial depth value is obtained based on structured light triangulation. The depth value is calculated using epipolar search matching based on the multi-view geometric relationships obtained in step S120. Smoothing term. Using a truncated linear model:
[0020]
[0021] Among them, weight With pixels Image gradient magnitude at [location] and consistency of motion vectors Related, specifically , This is the truncation threshold. Parameter Dynamically adjust based on the image signal-to-noise ratio of the current frame.
[0022] Preferably, the voxelization process in step S140 employs a weighted moving cube algorithm. For each depth point transformed to the global coordinate system... ,in Assign the confidence weights for the depth estimation at that point to the voxel containing it. At the 8 corner points. Voxels scalar value It is obtained by weighted summation of all depth points falling within its influence range:
[0023]
[0024] in, Here are the coordinates of the voxel corner point. It is a trilinear kernel function. This represents the Z-coordinate of the section of the cross-section where the voxel is located. This is a sign function. This method effectively smooths noise and maintains a sharp transition at tissue boundaries.
[0025] Furthermore, the layer thickness closed-loop control in step S150 is specifically implemented as follows: Let the preset theoretical sampling layer thickness be... In the During the second sampling, the actual physical displacement obtained by the fusion calculation of the inertial measurement unit and the encoder is: Calculate the layer thickness error. A proportional-integral controller is used to generate the control input:
[0026]
[0027] in, and For controller parameters, the control quantity Compensation amount converted into the next motion step of the precision displacement mechanism This makes the target displacement for the next cycle be This gradually eliminates accumulated errors and ensures that the final volumetric data has a uniform resolution in the Z-axis direction.
[0028] A machine vision-based intraoral tomographic imaging scanning system, comprising the following components:
[0029] A miniaturized optical probe module is used to move along a preset trajectory within the oral cavity and simultaneously acquire two-dimensional texture images and three-dimensional depth information of the target area. The miniaturized optical probe module integrates an active structured light projection unit, a high-speed complementary metal-oxide-semiconductor (CMOS) image sensor, an inertial measurement unit, and a precision displacement mechanism driven by a micro-stepping motor. The active structured light projection unit projects a specific coded pattern onto the target area; the high-speed CMOS image sensor synchronously captures the deformed pattern modulated by the surface of the target area; the inertial measurement unit measures the probe's triaxial acceleration and triaxial angular velocity in real time; and the precision displacement mechanism drives the probe to perform micron-level precision step-by-step translation along a linear guide rail.
[0030] The multimodal data fusion and pose estimation module receives the raw data stream from the miniaturized optical probe module and performs adaptive illumination compensation, robust feature point extraction and matching, and probe pose fusion estimation based on extended Kalman filtering. This module dynamically adjusts the exposure time and gain of the image sensor by analyzing the intensity distribution of the structured light pattern and the motion data of the inertial measurement unit to suppress saliva reflection and tooth surface highlights. At the same time, this module adopts a hybrid feature descriptor based on scale-invariant feature transformation and acceleration robust features, combines a random sampling consensus algorithm to eliminate mismatched point pairs, and integrates the pre-integration constraints of the inertial measurement unit to construct a six-degree-of-freedom pose transformation matrix of the probe between consecutive frames.
[0031] The real-time depth optimization and hole filling module is used to perform depth optimization and hole filling under motion consistency constraints based on the pose transformation matrix and the initial depth map obtained by structured light decoding. This module establishes an energy function based on a Markov random field, whose data term is composed of the depth measured by structured light triangulation and the depth constrained by multi-view geometry. The smoothing term introduces an adaptive weight based on the consistency between image edges and motion trajectories. The energy function is minimized by a graph cut algorithm, and the optimized dense depth map is output.
[0032] The tomographic slice generation and slice thickness closed-loop control module is used to generate continuous tomographic slice images through voxelization and resampling based on the optimized dense depth map sequence and precise probe pose, and to realize real-time feedback control of the sampled slice thickness. This module constructs a three-dimensional volume data space in a global coordinate system, transforms the depth point cloud of each frame to this space according to its corresponding pose matrix and assigns voxel values, and generates a series of tomographic slices parallel to the imaging plane by performing equal-interval resampling along the main direction of probe movement (Z-axis); at the same time, this module calculates the slice thickness error by comparing the preset theoretical sampling step size with the actual physical displacement fed back by the inertial measurement unit and encoder, and generates control commands to dynamically adjust the next motion step size of the precision displacement mechanism.
[0033] The central processing and display module is used to coordinate the timing operations of the above modules, execute the scheduling and calculation of the algorithm, and display the scanning process, intermediate results and the final generated tomographic slice sequence in real time.
[0034] Compared with the prior art, the beneficial technical effects of the present invention are as follows:
[0035] This invention significantly improves the robustness and accuracy of probe motion trajectory estimation in the complex dynamic environment of the oral cavity by introducing an inertial measurement unit and tightly coupling it with visual features through extended Kalman filter pose estimation. It effectively overcomes the problems of manual operation disturbance and failure of pure vision methods in weak texture areas, and fundamentally reduces interlayer misalignment and structural distortion in 3D reconstruction.
[0036] This invention designs a Markov random field depth optimization model based on multimodal data fusion. This model can effectively repair the holes and distortions in the depth map caused by motion blur, occlusion or reflection by dynamically weighted smoothing terms and dual data terms, and obtain high-quality, high-complete dense 3D point cloud data.
[0037] This invention proposes a real-time closed-loop control mechanism for the slice thickness of tomographic images. By feeding back the actual physical displacement and comparing it with a preset value, the probe movement step size is dynamically adjusted, which solves the problem of inconsistent sample slice thickness caused by mechanical backlash and sliding friction, and ensures that the final generated tomographic slice sequence has controllable and uniform resolution in the Z-axis direction.
[0038] This invention enables real-time or near-real-time processing of the entire process from image acquisition, pose estimation, depth optimization to slice generation, and integrates adaptive illumination compensation. It can evaluate data quality and adjust parameters in real time during scanning, significantly reducing the proportion of invalid data, shortening the overall clinical operation time, and improving the system's practicality and diagnostic efficiency. Attached Figure Description
[0039] Figure 1 This is a flowchart illustrating the specific steps of a machine vision-based intraoral tomographic image scanning method proposed in this invention.
[0040] Figure 2 This is a schematic diagram of the overall technical architecture of a machine vision-based intraoral tomographic imaging scanning system proposed in this invention.
[0041] Figure 3 This is a schematic diagram of the core principle framework of vision-inertial tightly coupled pose estimation and depth optimization proposed in this invention. Detailed Implementation
[0042] The features and exemplary embodiments of various aspects of the present invention will now be described in detail. To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely intended to explain the present invention and not to limit the present invention. For those skilled in the art, the present invention can be practiced without some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present invention by illustrating examples of the invention.
[0043] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0044] Example 1
[0045] In clinical dental settings, such as early diagnosis of proximal caries or root fractures in a single molar, high-precision, high-resolution tomographic images of the tooth and surrounding gingival tissue are required. Traditional oral endoscopes only provide two-dimensional surface views, while computed tomography (CT) scanners are bulky, emit high radiation doses, and cannot be operated in real-time at the dental chair. This invention provides a machine vision-based intraoral tomographic image scanning system and method. This system uses a miniaturized probe to perform linear scanning along the tooth surface within the oral cavity, simultaneously acquiring texture and depth information. After a series of real-time processing steps, a series of continuous tomographic slices with uniform thickness and clear structure are generated, providing dentists with a three-dimensional tomographic view.
[0046] See Figure 1 A scanning method for intraoral tomographic images based on machine vision, the specific steps of which are as follows:
[0047] Step S110: A miniaturized optical probe integrating an active structured light projector, a high-speed image sensor, an inertial measurement unit, and a precision displacement mechanism moves along a preset linear trajectory inside the oral cavity, synchronously acquiring two-dimensional images modulated by structured light patterns at each sampling position, and recording the original motion data of the inertial measurement unit and the encoder feedback of the displacement mechanism.
[0048] Step S120: Adaptive illumination compensation and feature extraction are performed on the acquired two-dimensional image sequence. Combined with the inertial measurement unit data, the precise six-degree-of-freedom pose transformation matrix of the probe between continuous sampling positions is calculated by fusing the data through an extended Kalman filter.
[0049] Step S130: Based on the pose transformation matrix and the initial depth map obtained by structured light decoding, construct and solve the Markov random field energy function that incorporates motion consistency constraints, optimize and repair the holes in the initial depth map, and generate an optimized dense depth map corresponding to each sampling position.
[0050] Step S140: The optimized dense depth maps of all sampling positions are uniformly transformed to the global three-dimensional coordinate system according to their corresponding pose matrices. Through voxelization and equal-interval resampling along the main direction of probe movement, a series of continuous tomographic slice images are generated.
[0051] In step S150, during the generation process of step S140, the error between the preset theoretical layer thickness and the actual physical displacement is calculated in real time, and the motion step size of the precision displacement mechanism in the next sampling cycle is dynamically adjusted based on the error to achieve closed-loop control of the consistency of the tomographic image layer thickness.
[0052] See Figure 2A machine vision-based intraoral tomographic imaging scanning system includes a miniaturized optical probe module, a multimodal data fusion and pose estimation module, a real-time depth optimization and cavity repair module, a tomographic slice generation and slice thickness closed-loop control module, and a central processing and display module. Under the coordination of the central processing and display module, each module works collaboratively according to a strict timing sequence and data flow.
[0053] The miniaturized optical probe module is the front-end data acquisition unit of the system. Its physical dimensions are specially designed to allow it to smoothly enter and stabilize within the target oral cavity with the assistance of a standard dental endoscope. At the core of this module is an integrated probe assembly consisting of an active structured light projection unit, a high-speed complementary metal-oxide-semiconductor (CMOS) image sensor, an inertial measurement unit, and a precision displacement mechanism driven by a micro-stepper motor. The active structured light projection unit uses digital micromirrors or laser diffraction elements to project a specific light spot pattern encoded with a combination of Gray code and sinusoidal phase shift onto the target area of the teeth or gums. The high-speed CMOS image sensor is mounted at a fixed angle to the optical axis of the structured light projection unit, forming the basis of the binocular structured light vision system. Its frame rate is no less than 200 frames per second, used to synchronously capture the coded pattern that deforms after being modulated by the complex curvature of the target area. The inertial measurement unit (IMU) is mounted close to the image sensor and includes a three-axis MEMS gyroscope and a three-axis MEMS accelerometer. It measures the probe's triaxial angular velocity and triaxial acceleration in real time at frequencies above 1000 Hz during probe movement. A precision displacement mechanism uses a miniature stepper motor combined with a precision ball screw or piezoelectric ceramic driver to drive the entire probe assembly to translate along a rigid linear guide. This mechanism incorporates an optical encoder to provide feedback on the probe's actual physical displacement, achieving micron-level displacement resolution and repeatability better than ±5 microns. At the start of the scan, the probe is positioned at the beginning of the target area. The central processing and display module issues a command, and the precision displacement mechanism drives the probe to begin step-by-step movement with a preset initial step size. At each sampling position, the system executes strict synchronous triggering: the active structured light projection unit emits an coded pattern, the high-speed complementary metal-oxide-semiconductor (CMOS) image sensor captures a frame after the pattern stabilizes, and simultaneously, the IMU acquires a set of motion data, while the encoder records the current position. All raw data is timestamped and transmitted in real time to the subsequent processing module via a high-speed serial bus.
[0054] The multimodal data fusion and pose estimation module receives raw image streams and inertial measurement unit (IMU) data streams from the front end. This module first performs adaptive illumination compensation processing on each frame of the 2D image. The processing involves: calculating the grayscale histogram of the structured light pattern region in the current frame in real time, and statistically analyzing the pixel percentages in two key threshold intervals. Specifically, the system statistically analyzes the percentage of overexposed areas (pixel grayscale values greater than 240) and underexposed areas (pixel grayscale values less than 15). If the percentage of overexposed areas exceeds 5%, it is determined that there is strong specular reflection caused by saliva or smooth tooth enamel. The compensation strategy is to calculate the average grayscale value of all pixels outside the overexposed areas, denoted as . Then follow the formula Global grayscale attenuation is applied to the entire image, where Represents the grayscale value of each pixel in the original image. This represents the adjusted grayscale value. The formula adjusts the overall image brightness median to around 128, effectively suppressing highlights. If the underexposed area accounts for more than 10%, it is determined that the ambient light is insufficient or the probe is too far from the target. The compensation strategy is that the module immediately sends a control command to the miniaturized optical probe module, requiring the active structured light projection unit to increase the light intensity of the projected pattern by 30% in the next frame; simultaneously, it instructs the high-speed complementary metal-oxide-semiconductor image sensor to increase its analog gain by 8 dB to enhance signal reception sensitivity. After completing the illumination compensation, the module extracts robust feature points from the image. A hybrid feature descriptor algorithm based on scale-invariant feature transform and accelerated robust features is adopted. First, key points are detected in different scale spaces, and then a hybrid feature vector combining a gradient orientation histogram descriptor based on scale-invariant feature transform and a binary descriptor based on accelerated robust features is calculated for each key point. For two consecutive frames, a fast nearest neighbor search algorithm is used for feature point matching, and a random sampling consensus algorithm is used to iteratively estimate the fundamental matrix, thereby eliminating mismatched point pairs caused by duplicate textures or noise, and obtaining a reliable set of feature point correspondences.
[0055] Simultaneously, this module processes inertial measurement unit (IMU) data in parallel. A tightly coupled vision-inertial odometry method is used for probe pose fusion estimation. The IMU measures the angular velocity. and acceleration Pre-integration is performed over the time interval between adjacent image frames to calculate the relative rotation change of the probe between the two frames. Relative translation changes and speed change The pre-integration process considers the zero bias of the gyroscope and accelerometer, modeling it as a random walk process incorporated into the state vector. In the vision part, the essential matrix E is calculated from the matched feature point pairs through epipolar geometry constraints, and the rotation matrix between the two frames is obtained through singular value decomposition. Translation vector The scale of the translation vector is recovered using scale information provided by pre-integration of the inertial measurement unit (IMU). The state vector of the extended Kalman filter is defined as 15-dimensional, containing the probe's position, attitude, and velocity in the world coordinate system, as well as the zero bias of the gyroscope and accelerometer. The system's state prediction is based on the IMU pre-integration results. Observation updates are derived from both visual and inertial observations. The visual observation equation consists of the reprojection error of the feature points, i.e., the difference between the predicted feature point position and the actual extracted position. The inertial observation equation consists of the error of the pre-integrated quantities. The filter iteratively updates the optimal estimate of the state vector in real time by minimizing the weighted sum of squares of the reprojection error and the inertial measurement error, ultimately outputting the precise six-DOF pose transformation matrix of the probe relative to the starting point at each sampling moment. This matrix describes the probe's rotational and translational motion from the previous frame to the current frame and is the basis for subsequent 3D reconstruction.
[0056] The real-time depth optimization and cavity repair module receives the pose transformation matrix from the pose estimation module, along with an initial depth map obtained through phase calculation and triangulation of the structured light coded pattern. Due to uneven illumination in the oral cavity environment, tissue translucency, motion blur, and local occlusion, the initial depth map often contains noise, outliers, and void regions with missing data. This module optimizes the depth map by constructing a Markov random field model that integrates multi-view geometric constraints and motion consistency. The energy function of this model... Defined as the sum of the data item and the smoothing term. Data item For each pixel p in the depth map, the depth value The constraint is composed of two parts: the first part is the depth value obtained based on single-frame structured light triangulation. and the depth value to be determined The first part is the squared deviation; the second part is the depth value obtained based on multi-view geometric relationships. and the depth value to be determined The squared deviation. Where, This is calculated by projecting the feature points of the current frame onto adjacent keyframes according to the pose transformation matrix, performing sub-pixel precision search and matching along the epipolar lines, and then calculating through triangulation. The two terms are multiplied by weight coefficients α and β respectively and then added together. Smoothing term. The effect applies between adjacent pixels p and q, using a truncated linear model, whose expression is: Where T is the cutoff threshold, used to allow for larger depth differences at depth discontinuities, such as tooth edges. The weight of the smoothing term... It's not fixed, but adaptive. It depends on the difference in image gradient magnitudes at pixels p and q, as well as the consistency of motion vectors between two frames. Dynamic calculation. Motion vector consistency. The consistency between the local motion calculated using the optical flow method and the global motion given by the pose estimation module in that region is obtained. The specific calculation formula is as follows: , where η is the baseline weight parameter. Weight Dynamically adjust based on the signal-to-noise ratio of the current frame image: when the signal-to-noise ratio is low, increase the reliance on direct measurements. Reducing geometric constraints and smoothing strength reduces and Conversely, the signal-to-noise ratio decreases when it is high. The entire energy function is minimized using a graph cut algorithm, ultimately outputting an optimized, dense depth map with holes filled. This depth map is smooth in flat areas, sharp at edges, and effectively repairs missing data caused by reflections or occlusions.
[0057] The core task of the tomographic slice generation and slice thickness closed-loop control module is to fuse a series of optimized dense depth maps over time into a globally consistent 3D volumetric data set and extract tomographic slices from it. The module first constructs a global 3D Cartesian coordinate system volumetric data space in memory, covering the entire scanning area. For the optimized dense depth map obtained in the k-th frame, the module transforms each 3D point cloud in the depth map to the global space using the pose matrix provided by the multimodal data fusion and pose estimation module, which transforms the k-th frame coordinate system to the global coordinate system. Each 3D point carries a confidence weight in addition to its coordinate information. The weight is derived from the estimated variance of the depth value at that point during depth optimization. The voxelization process employs a weighted moving cube algorithm. For a voxel in global space, its scalar value is jointly determined by all 3D points falling within its influence range. Specifically, for each 3D point... The trilinear kernel function φ is used to calculate the weighted contribution of a point to its eight surrounding voxel corner points, and this contribution is multiplied by the confidence weight of that point. and symbolic functions ,in This is the Z-axis coordinate of the current tomographic slice to be extracted. Then, the weighted contributions of all points are accumulated and applied to the corresponding voxel corner points. The final scalar value of each voxel corner point... The result is obtained by weighted summation and division by the weighted sum. The advantage of this method is that it naturally achieves the fusion of data from different perspectives and with different confidence levels, while maintaining a clear transition at tissue boundaries and avoiding excessive smoothing. After completing the voxelization accumulation of all frame data, the module moves along the main direction of probe movement, i.e., the Z-axis direction of the global coordinate system, with a preset theoretical layer thickness. For each interval, perform equal-interval resampling. The location is determined, and all voxel values on the XY plane are extracted to generate a two-dimensional grayscale tomographic slice image. A series of equally spaced slices constitute the final tomographic image sequence.
[0058] To ensure absolute uniformity of the sampling layer thickness, the module synchronously executes closed-loop control of the layer thickness. After the k-th sampling, the module reads the pre-integration result of the inertial measurement unit and the feedback from the encoder of the precision displacement mechanism, and calculates the actual physical displacement of the probe during the period from the (k-1)-th to the k-th sampling using a sensor fusion algorithm. Calculate the layer thickness error for this sampling. The module maintains a cumulative sum of layer thickness errors. A digital proportional-integral controller is used to handle this error. The controller's output u(k) consists of a proportional term and an integral term: the proportional term represents the current error. Multiply by the proportionality factor The integral term is the sum of all errors from the start of the scan to the current time, multiplied by the integral coefficient. The control output u(k) is converted into a compensation amount for the next motion step of the precision displacement mechanism. That is, the next target displacement command issued by the system to the precision displacement mechanism is no longer fixed. , but This real-time feedback and compensation mechanism can gradually offset the accumulation of displacement errors caused by factors such as mechanical transmission backlash, uneven sliding friction of the guide rail, or stepping loss of the motor, ensuring that the final generated three-dimensional volume data has a highly consistent resolution in the Z-axis direction and avoiding stretching or compression distortion of the sliced images.
[0059] The central processing and display module, serving as the system's overall control core, runs on an embedded industrial computer or high-end workstation. This module is responsible for initializing all hardware, configuring algorithm parameters, and establishing communication channels between modules. During the scanning process, it rigorously schedules the pipeline timing of data acquisition, pose estimation, depth optimization, slice generation, and control command issuance to ensure real-time performance. Simultaneously, this module provides a graphical user interface that displays in real-time the raw images acquired by the probe, feature matching results, estimated motion trajectories, depth map comparisons before and after optimization, real-time fusion effects of 3D point clouds, and the final generated tomographic slice sequence. Doctors can interact with the interface to adjust the scanning area, preset slice thickness, start or pause the scan, and perform post-processing operations such as adjusting window width and level, and taking measurements on the generated tomographic slices.
[0060] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape, and principle of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A scanning method for intraoral tomographic images based on machine vision, characterized in that, The specific steps of this method are as follows: Step S110: A miniaturized optical probe, which integrates an inertial measurement unit and a precision displacement mechanism, moves along a preset linear trajectory inside the oral cavity, simultaneously acquiring two-dimensional images and recording raw motion data and encoder feedback. Step S120: Perform adaptive illumination compensation and feature extraction on the acquired two-dimensional image sequence, and calculate the six-degree-of-freedom pose transformation matrix of the probe between continuous sampling positions; The adaptive illumination compensation includes: dynamically adjusting the exposure time and gain of the image sensor by analyzing the intensity distribution of the structured light pattern and the motion data of the inertial measurement unit to suppress saliva reflection and tooth surface highlights; the feature extraction adopts a hybrid feature descriptor based on scale-invariant feature transformation and acceleration robust features, combined with a random sampling consensus algorithm to eliminate mismatched point pairs, and integrates the pre-integration constraints of the inertial measurement unit to construct a six-degree-of-freedom pose transformation matrix of the probe between consecutive frames; the six-degree-of-freedom pose estimation adopts a tightly coupled vision-inertial odometry method, the state vector of the extended Kalman filter includes the probe pose, velocity and inertial measurement unit zero bias, the observation equation is composed of visual geometric constraints and inertial pre-integration constraints, and the probe pose estimate is output by minimizing the weighted sum of reprojection error and inertial measurement error; Step S130: Based on the pose transformation matrix and the initial depth map obtained by structured light decoding, a Markov random field energy function is constructed and solved to optimize and repair holes in the initial depth map, generating an optimized dense depth map. The data term of the Markov random field energy function is composed of the structured light triangulation measurement depth and the multi-view geometric constraint depth. The smoothing term introduces an adaptive weight based on the consistency between image edges and motion trajectories. The energy function is minimized by a graph cut algorithm, and the optimized dense depth map is output. Step S140: The optimized dense depth map is uniformly transformed to the global three-dimensional coordinate system. Through voxelization and equal-interval resampling along the main direction of probe movement, a series of continuous tomographic slice images are generated. In step S150, during the generation process of step S140, the error between the preset theoretical layer thickness and the actual physical displacement is calculated in real time. Based on this error, the motion step size of the precision displacement mechanism in the next sampling cycle is dynamically adjusted to complete the closed-loop control of the tomographic image layer thickness consistency. The closed-loop control is specifically implemented as follows: assuming a preset theoretical sampling layer thickness, the actual physical displacement is obtained by fusing the inertial measurement unit and the encoder at the k-th sampling; the layer thickness error is calculated; a proportional-integral controller is used to generate a control quantity; the control quantity is converted into a compensation quantity for the next motion step size of the precision displacement mechanism to obtain the target displacement of the next cycle.
2. A scanning system for intraoral tomographic images based on machine vision, characterized in that, The system includes the following components: A miniaturized optical probe module is used to move along a preset trajectory within the oral cavity and simultaneously acquire two-dimensional texture images and three-dimensional depth information of the target area; The miniaturized optical probe module integrates an inertial measurement unit and a precision displacement mechanism driven by a micro stepper motor. The multimodal data fusion and pose estimation module receives the raw data stream from the miniaturized optical probe module and performs adaptive illumination compensation, robust feature point extraction and matching, and probe pose fusion estimation based on extended Kalman filtering. The adaptive illumination compensation includes dynamically adjusting the exposure time and gain of the image sensor by analyzing the intensity distribution of the structured light pattern and the motion data of the inertial measurement unit (IMU) to suppress saliva reflection and tooth surface highlights. The feature point extraction and matching uses a hybrid feature descriptor based on scale-invariant feature transformation and accelerated robust features, combined with a random sampling consensus algorithm to eliminate mismatched point pairs, and integrates the pre-integration constraints of the IMU to construct a six-degree-of-freedom pose transformation matrix for the probe across consecutive frames. The pose estimation uses a tightly coupled visual-inertial odometry method. The state vector of the extended Kalman filter includes the probe pose, velocity, and IMU zero bias. Its observation equation is composed of visual geometric constraints and inertial pre-integration constraints. By minimizing the weighted sum of reprojection error and inertial measurement error, the probe pose estimate is output. The real-time depth optimization and hole repair module is used to perform depth optimization and hole filling under motion consistency constraints based on the pose transformation matrix and the initial depth map obtained by structured light decoding. The real-time depth optimization and hole repair module establishes an energy function based on a Markov random field. Its data term is composed of the depth measured by structured light triangulation and the depth constrained by multi-view geometry. The smoothing term introduces an adaptive weight based on the consistency between image edges and motion trajectories. The energy function is minimized by a graph cut algorithm, and the optimized dense depth map is output. The tomographic slice generation and slice thickness closed-loop control module is used to generate continuous tomographic slice images through voxelization and resampling based on the optimized dense depth map sequence and precise probe pose, and to realize real-time feedback control of the sampled slice thickness. The slice thickness closed-loop control is specifically implemented as follows: Given a preset theoretical sampled slice thickness, at the k-th sampling, the actual physical displacement is calculated by fusing the inertial measurement unit and the encoder; the slice thickness error is calculated; a proportional-integral controller is used to generate a control quantity; the control quantity is converted into a compensation quantity for the next motion step of the precision displacement mechanism to obtain the target displacement for the next cycle. The central processing and display module is used to coordinate the timing operations of the above modules, execute the scheduling and calculation of the algorithm, and display the scanning process, intermediate results and the final generated tomographic slice sequence in real time.
3. The intraoral tomographic imaging scanning system based on machine vision according to claim 2, characterized in that: The miniaturized optical probe module also integrates an active structured light projection unit and a high-speed complementary metal-oxide-semiconductor (CMOS) image sensor. The active structured light projection unit is used to project a specific coded pattern onto the target area. The high-speed CMOS image sensor is used to synchronously capture the deformed pattern modulated by the surface of the target area. The inertial measurement unit is used to measure the probe's triaxial acceleration and triaxial angular velocity in real time. The precision displacement mechanism is used to drive the probe to perform micron-level precision step-by-step translation along a linear guide rail.
4. The intraoral tomographic imaging scanning system based on machine vision according to claim 2, characterized in that, The adaptive illumination compensation in the multimodal data fusion and pose estimation module specifically involves: real-time monitoring of the grayscale histogram distribution of the structured light pattern in the current frame; if overexposed, global attenuation of the image; if underexposed, triggering the active structured light projection unit to emit a compensation pattern with increased intensity in the next frame, and simultaneously increasing the image sensor gain.
5. The intraoral tomographic imaging scanning system based on machine vision according to claim 2, characterized in that: The tomographic slice generation and slice thickness closed-loop control module constructs a three-dimensional volume data space in a global coordinate system. It transforms the depth point cloud of each frame to this space according to its corresponding pose matrix and assigns voxel values. By performing equidistant resampling along the main direction of probe movement, a series of tomographic slices parallel to the imaging plane are generated.