Intelligent voice-guided oral scanning device and automatic acquisition method thereof
By using intelligent voice-guided oral scanning equipment, combined with multimodal fusion positioning and automatic triggering algorithms, the problems of expensive equipment and complex operation in existing technologies have been solved, enabling efficient and low-threshold oral scanning and high-quality image acquisition for home users.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI HANYING HEALTH MANAGEMENT CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-06-26
AI Technical Summary
Existing oral scanning equipment is expensive and complex to operate, and cannot achieve adaptive guidance for home users or image quality that meets clinical diagnostic requirements.
The intelligent voice-guided oral scanning device combines a voice interaction module, an image acquisition module, a motion sensing module, and an edge computing module. Through multimodal fusion positioning algorithms and automatic triggering algorithms, it achieves real-time positioning of the device in the oral cavity coordinate system and image quality assessment, and automatically triggers image acquisition.
It enables oral scanning for home users with zero barriers to entry, with image quality meeting clinical diagnostic requirements, a success rate of over 95%, and a rejection rate of less than 5%, offering both high-end medical-grade and economical consumer-grade configurations.
Smart Images

Figure CN122272208A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oral scanning technology, and in particular to an intelligent voice-guided oral scanning device and its automatic data acquisition method. Background Technology
[0002] Oral health is an important component of overall health. Oral diseases such as dental caries and periodontal disease are characterized by high incidence and early insidious occurrence. Regular oral monitoring in the home setting is of great significance for early detection and early intervention.
[0003] Existing technologies, such as the method for color chromaticity matching disclosed in CN112912933B, identify one or more chromaticity values by comparing the angular distribution of a color vector with a set of reference angular distributions of the color vector. Each chromaticity value is associated with an angular distribution of the color vector from the spatially resolved angular distribution of the color vector, and each reference angular distribution in the set is associated with a corresponding chromaticity value. This method displays, stores, or transmits surface data and indications of one or more chromaticity values. However, it has the following drawbacks: it mainly relies on professional intraoral scanners such as 3Shape TRIOS and Shining 3D, employing structured light or confocal microscopy techniques. While offering high precision, the equipment is expensive, bulky, and requires professional operation, limiting its use to medical institutions. Its reliance on professional intraoral scanners results in expensive equipment and a high operational barrier.
[0004] For example, the AR-based root canal monitoring method and system disclosed in announcement number CN110269715B, although CT can clearly capture the internal condition of the root canal, cannot be used for real-time surgery due to the harmfulness of X-rays to the human body, and cannot realize the implementation of displaying the actual root canal shape and precise location in real time. Summary of the Invention
[0005] The purpose of this invention is to solve the problems in the prior art where the guidance strategy cannot be adaptively adjusted according to user behavior and the image quality cannot meet the requirements of clinical diagnosis. Therefore, this invention proposes an intelligent voice-guided oral scanning device and its automatic acquisition method.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] A smart voice-guided oral scanning device includes a grip and a scanning head. The grip and the scanning head are connected by a double-beveled buckle and a metal positioning pin. The scanning head is equipped with:
[0008] The voice interaction module is used to output scanning guidance voice and receive user voice feedback;
[0009] Image acquisition module, the image acquisition module being used to acquire oral cavity images;
[0010] A motion sensing module, which is used to collect device motion data;
[0011] Optical markers are set at the edge of the device's scanning head to provide artificial feature-assisted positioning when visual features are sparse;
[0012] The edge computing module is used to run:
[0013] Oral cavity region detection algorithm, which is used to identify the currently scanned oral cavity region;
[0014] Image quality assessment algorithm, used to evaluate image quality in real time;
[0015] A multimodal fusion localization algorithm is used to fuse the visual data from the image acquisition module, the observation data from the optical markers, and the inertial data from the motion sensing module to obtain the real-time pose of the device in the oral cavity coordinate system, and to map the real-time pose to the corresponding oral cavity region.
[0016] An automatic triggering algorithm is used to automatically trigger the image acquisition module to take pictures when the image quality, the oral cavity region, and the stability of the real-time pose all meet preset conditions.
[0017] Preferably, the scanning head is further provided with an auxiliary device, which includes a camera, a fill light, a microphone, a speaker, and an indicator light.
[0018] Preferably, the optical markers include eight miniature infrared reflective markers arranged on the scanning head.
[0019] Preferably, the eight micro infrared reflective markers are distributed in a non-coplanar spatial manner, with four micro infrared reflective markers located on the front face and the other four located on the side arc face.
[0020] Preferably, the multimodal fusion localization algorithm employs an extended Kalman filter framework, where visual feature points and optical marker points are used as observations, inertial measurement integrals are used as predictions, and the observation covariance weights are dynamically adjusted based on the current visual feature richness.
[0021] When there are enough visual feature points, visual observation is the primary method.
[0022] When visual feature points are insufficient, observation is mainly based on optical markers.
[0023] Preferably, the automatic triggering algorithm adopts a multi-condition joint judgment mechanism, and the preset conditions include: the image quality score is higher than the dynamic threshold, the matching degree between the current oral cavity region and the target region is higher than the preset matching threshold, the device movement speed is lower than the preset speed threshold, and the stable duration of the device at the current position exceeds the preset time threshold.
[0024] The dynamic threshold is adaptively adjusted based on the user's age and / or the number of failures in the current scanning cycle.
[0025] Preferably, the voice interaction module employs a hierarchical finite state machine and extracts user operation behavior features in real time;
[0026] The operational behavior characteristics include movement speed, jitter amplitude, response delay, number of erroneous entries, and number of timeouts.
[0027] User behavior profiles are dynamically constructed based on an online Bayesian classifier. The guidance mode is adaptively selected according to user type, including gamified mode for children, concise mode for adults, assistance mode for the elderly, and assistance mode for unstable users. Different modes have differences in voice style, level of detail, speech rate, and tactile feedback.
[0028] Implement a progressive bootstrapping strategy based on the number of failures.
[0029] Preferably, the image quality assessment algorithm includes the calculation of the following diagnostic-specific features: tooth edge contrast, gingival region color distribution, highlight area ratio, and consistency with dentist annotations;
[0030] The image quality assessment algorithm also includes an illumination adaptive processing module, which first identifies the illumination type through a lightweight CNN and then calls the corresponding enhancement operator.
[0031] An automatic data acquisition method for an intelligent voice-guided oral scanning device includes the following steps:
[0032] Step S1: Guide the user via voice to aim the device at the target oral cavity area;
[0033] Step S2: Real-time acquisition of preview images, device motion data, and optical marker observation data;
[0034] Step S3: Based on the extended Kalman filter, the visual feature points, optical marker points and inertial measurement data are fused to obtain the precise pose of the device in the oral cavity coordinate system and map it to the corresponding oral cavity region.
[0035] Step S4: Calculate the current image quality score based on a multi-dimensional quality assessment model, wherein the quality assessment model includes diagnostic-specific features and is adaptively enhanced according to the illumination type;
[0036] Step S5: When the image quality score, region matching degree, and device pose stability all meet the dynamic threshold, shooting is automatically triggered;
[0037] Step S6: Evaluate the shooting results. If the results are unsatisfactory, perform a progressive voice-guided reshoot based on the number of failures. If the results are satisfactory, update the area coverage matrix.
[0038] Step S7: Check the scan integrity, including area coverage and path compliance rate. When the overall integrity score reaches the threshold, the scan ends, a report is generated, and synchronized to the cloud.
[0039] Compared with the prior art, the present invention has the following advantages:
[0040] 1. This invention adopts a handheld design, which is ergonomic. The grip is cylindrical with a diameter of 30-70mm, and the scanning head is at a 15° angle to the grip, making it easy to penetrate deep into the oral cavity.
[0041] 2. By designing the optical markers to be distributed in a non-coplanar space, this invention ensures that at least three markers are visible under any device posture, which helps to solve the technical problem of positioning failure in scenarios with sparse visual features inside the oral cavity, thereby achieving continuous positioning throughout the entire scanning process.
[0042] 3. This invention uses extended Kalman filtering to fuse multi-source data, dynamically adjusts the fusion weights based on the richness of visual features, and achieves fully automatic acquisition through voice hierarchical guidance and dynamic quality thresholds. This helps to solve the problems of high threshold and high rejection rate in home dental scanning operations, thereby achieving zero-threshold standardized dental image acquisition. It is available in both high-end medical-grade and economical consumer-grade configurations and can be used for home oral health monitoring and remote diagnosis and treatment. Attached Figure Description
[0043] Figure 1 This is a schematic diagram of the structure of an intelligent voice-guided oral scanning device proposed in this invention;
[0044] Figure 2 This is a top view of the structure of an intelligent voice-guided oral scanning device proposed in this invention;
[0045] Figure 3 This is a schematic diagram of a module of an intelligent voice-guided oral scanning device proposed in this invention;
[0046] Figure 4 This is a schematic diagram of the application layer of an intelligent voice-guided oral scanning device proposed in this invention;
[0047] Figure 5 This is a schematic diagram of entering the scanning cycle phase.
[0048] In the diagram: 1. Grip; 2. Scanning head; 3. Double-beveled buckle; 4. Auxiliary device; 41. Camera; 42. Fill light; 43. Microphone; 44. Speaker; 45. Indicator light. Detailed Implementation
[0049] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0050] Reference Figures 1-5 A smart voice-guided oral scanning device includes a grip 1 and a scanning head 2. To ensure stable one-handed operation in a moist and narrow oral environment, the grip 1 is cylindrical with a diameter of 30-70mm. This diameter range is determined based on ergonomic experimental data on the hand grip comfort of the 5th to 95th percentile adults. The aim is to maximize the contact area between the web of the thumb and the fingertips to distribute pressure, while reducing the ulnar torque of the wrist caused by prolonged operation.
[0051] The gripping part 1 and the scanning head 2 are connected by a double-beveled buckle 3 and a metal positioning pin. The two are angled at 15° by the double-beveled buckle 3 and the metal positioning pin, facilitating deep penetration into the oral cavity. This 15° golden working angle is derived from geometrical constraints simulation of the depth of the oral vestibule and the space behind the molars, allowing for easy access into the oral cavity and bypassing bony prominences such as the zygomatic alveolar ridge. Furthermore, the double-beveled buckle 3 is injection molded from medical-grade PEEK polymer material. Its two wedge-shaped mating surfaces have a complementary Mohs taper design (typically 1:10 to 1:20), generating elastic deformation under axial clamping force to eliminate mating gaps and provide high-friction self-locking torque. Together with the implanted 304 stainless steel metal positioning pin, they achieve rigid fixation of the scanning head 2 at a 15° angle relative to the axis of the gripping part 1.
[0052] The entire device is approximately 60-130mm long and weighs less than 150g. This lightweight design helps reduce static muscle load during cantilever operation and meets the safety requirements of IEC 60601-1 for medical electrical equipment regarding mechanical stress on handheld components. It should be noted that the entire device is IPX7 waterproof, the scanning head 2 is detachable and replaceable, and its electrical interface uses a flexible gold-plated pin connector. After 500 repeated disassembly and reassembly cycles, the positioning deviation is <0.1mm. A 0.1mm medical-grade PET film and anti-reflective coating are used, and a disposable transparent protective sleeve with >95% light transmittance protects the scanning head 2. The grip part 1 has a silver ion antibacterial coating.
[0053] The scanning head 2 is also equipped with an auxiliary device 4, which includes a camera 41, a fill light 42, a microphone 43, a speaker 44, and an indicator light 45.
[0054] It is worth noting that:
[0055] The optical axis of the optical module of camera 41 and the axis of the sound outlet of speaker 44 of auxiliary device 4 are at a specific angle in the YZ plane to avoid the micro-shaking interference caused by sound wave vibration on the CMOS rolling shutter. The supplementary light 42 uses four high color rendering index LEDs (CRI>95) arranged in a ring symmetrical arrangement. Its spectral power distribution curve is close to that of the D65 standard light source, which can maximize the restoration of the color saturation of gingival soft tissue and provide an unbiased light source basis for the extraction of gingival color distribution features in subsequent image quality assessment. The microphone 43, speaker 44, and indicator light 45 are all common technical means in this field, so they will not be described in detail.
[0056] The scanning head 2 is equipped with a voice interaction module, an image acquisition module, a motion sensing module, optical markers, a wireless communication module, and a power supply module.
[0057] The voice interaction module outputs scanning guidance voice and receives user voice feedback. In some implementations, a microphone array and speaker are arranged side-by-side, with dual / single microphones, MEMS silicon microphones, and 1W speakers to ensure good sound quality. It should be noted that the voice guidance strategy algorithm uses a hierarchical finite state machine to implement adaptive voice guidance. The top layer of this state machine is divided into five main states: "Idle - Guidance - Acquisition - Diagnosis - Retry". Within each main state, nested regions traverse sub-states (such as the maxillary lateral surface and mandibular occlusal surface). The state transition conditions not only depend on the image recognition results but also integrate user operation behavior features extracted in real time by the motion perception module (including movement speed vector magnitude, jitter amplitude spectral density, command response delay, frequency of entry into erroneous regions, and number of timeouts in non-target regions).
[0058] The voice interaction module uses a hierarchical finite state machine and extracts user operation behavior characteristics in real time. These characteristics include movement speed, jitter amplitude, response latency, number of incorrect entries, and number of timeouts.
[0059] The voice interaction makes operation convenient, with voice guidance throughout the process, analogous to brushing teeth. It is gamified for children, simple for adults, and features a loud, slow voice for the elderly. Moreover, it is automatically triggered without the need for buttons.
[0060] The image acquisition module is used to acquire oral cavity images, mainly using a CMOS sensor, with four specifications to choose from: 4K / 1080P / 720P / 480P. Its 80° wide-angle design facilitates autofocus and fixed focus. The 4TOPS computing power is determined based on constraints of end-to-end latency <100ms (image acquisition 33ms + region detection 15ms + quality assessment 10ms + fusion 5ms + speech 30ms) and total power consumption <2W. This timing constraint, together with the thermal design power of <2W, determines the trade-off between model pruning rate and INT8 quantization accuracy.
[0061] The motion sensing module is used to collect device motion data. In some implementations, a 6-axis IMU is used in conjunction with a gyroscope and an accelerometer, with a sampling rate ≥100Hz. The center of the IMU chip is offset from the optical center of the camera in the X / Y directions by ≤2mm, and coplanar in the Z direction, in order to minimize pseudo-acceleration interference caused by the lever arm effect. The system adopts a hardware-triggered synchronization mechanism: the rising edge signal at the start of CMOS exposure triggers an external interrupt of the MCU.
[0062] By capturing and latching the FIFO data index of the IMU through a hardware timer, the time synchronization accuracy is <1ms, thereby effectively suppressing pose estimation overshoot caused by delay jitter in asynchronous fusion.
[0063] Optical markers are positioned at the edge of the scanning head to provide artificial feature-assisted localization in situations where visual features are sparse. The optical markers consist of eight miniature infrared reflective markers arranged on the scanning head 2. These eight miniature infrared reflective markers are non-coplanarly distributed, with four located on the front face and the other four on the side curved surface. This geometric configuration ensures that at least three markers are visible in any viewing posture within the oral cavity, even if some markers are obscured by saliva or lips / cheeks. This is used for PnP pose solving and helps solve the technical problem of localization failure in scenarios with sparse visual features inside the oral cavity, achieving continuous localization throughout the scanning process.
[0064] In some implementations, a ring-shaped LED, consisting of 4-6 high color rendering LEDs (CRI > 95), with adjustable brightness, is used. Further, when solving the six-DOF pose of the device in the oral cavity coordinate system based on the EPnP algorithm, this configuration reduces the condition number and enhances the robustness of pitch and roll angle observations. This addresses the localization failure problem caused by the loss of visual SLAM features in scenarios with sparse textures in the soft and hard tissues inside the oral cavity. The ring-shaped LED supplementary lighting includes an 850nm infrared component, specifically used to improve the signal-to-noise ratio of marker points on the CMOS sensor. It should be noted that in terms of adaptive lighting processing, a lightweight CNN (3-layer convolution) is used for lighting type classification (underexposure / overexposure / uniform reflection / shadow), and corresponding enhancement operators (gamma correction, CLAHE, multi-scale Retinex) are invoked.
[0065] The wireless communication module mainly adopts WiFi and BLE, dual-mode communication, which enables data synchronization and cloud interaction. Based on the TDMA time slot scheduling mechanism, it ensures the concurrent transmission of high-definition video streams and low-power status maintenance signaling.
[0066] The power module mainly uses lithium polymer batteries with a capacity of 1000-2000mAh and magnetic charging for easy use.
[0067] The edge computing module utilizes an AI chip / NPU, providing 4 TOPS of computing power. In the cost-effective version, an ESP32-S3 AI accelerator is used, with power consumption <0.5W, significantly reducing costs. The edge computing module is used to run oral cavity region detection algorithms, image quality assessment algorithms, multimodal fusion localization algorithms, and automatic triggering algorithms.
[0068] The oral cavity region detection algorithm is used to identify the currently scanned oral cavity region. It adopts a lightweight object detection network (based on YOLO-Nano or MobileNet-SSD improvements), optimized for oral cavity scenes. During inference, it integrates PANet feature pyramids, and the output includes oral cavity bounding boxes, tooth semantic segmentation masks, and region categories (maxillary / mandibular × lateral / medial / occlusal surface). On the ESP32-S3 platform, it is accelerated by a vectorized instruction set to achieve a smooth detection frame rate of >30fps.
[0069] Input: 640×480 RGB image;
[0070] Backbone network: MobileNetV3-Large (after pruning);
[0071] Feature pyramid: PANet multi-scale fusion;
[0072] Output: oral cavity bounding box, tooth segmentation mask, region category (maxillary / mandibular × lateral / medial / occlusal surface);
[0073] Model optimization: INT8 quantization + channel pruning + knowledge distillation, model size <5MB, inference speed >30fps@640×480.
[0074] Image quality assessment algorithms are used to evaluate image quality in real time, construct a multi-dimensional quality assessment model, and output a comprehensive quality score (0-1):
[0075] Evaluation Dimensions Weight Calculation method Clarity 25% Laplace variance, high-frequency energy Exposure 20% Average brightness, histogram distribution Contrast 15% Local contrast, dynamic range Oral integrity 20% Oral cavity region proportion and edge integrity (deep learning) Tooth visibility 15% Teeth exposure ratio and occlusion (deep learning) Motion blur 5% Inter-frame difference and optical flow analysis
[0076] The aforementioned general dimensions constitute the basic quality score, upon which diagnostic-specific features are added for weighted adjustment.
[0077] To meet clinical needs, features such as edge contrast (dental caries screening), gingival color distribution (periodontal screening), and area of highlight region (reflection detection) were introduced. The consistency with the dentist's annotation was verified through ablation experiments (F1=0.91, which is better than the general index of 0.72).
[0078] Specific features introduced to meet dental diagnostic needs include: using the Sobel operator to calculate tooth margin contrast to assess the visibility of caries and leukoplakia; using the HSV spatial Gaussian mixture model to analyze gingival color distribution to aid in the assessment of inflammation; and using threshold segmentation to calculate the area ratio of highlight regions to detect enamel specular reflection.
[0079] The multimodal fusion localization algorithm is used to fuse visual data from the image acquisition module, observation data from optical markers, and inertial data from the motion sensing module to obtain the real-time pose of the device in the oral cavity coordinate system, and then map it to the corresponding oral cavity region based on the real-time pose.
[0080] The automatic triggering algorithm automatically triggers the image acquisition module to capture images when the image quality, oral cavity region, and real-time pose stability all meet preset conditions. In some implementations, the automatic triggering algorithm employs a multi-condition joint judgment mechanism. The preset conditions include: the image quality score is higher than a dynamic threshold, the matching degree between the current oral cavity region and the target region is higher than a preset matching threshold, the device movement speed is lower than a preset speed threshold, and the stable duration of the device in the current position exceeds a preset time threshold. The dynamic threshold is adaptively adjusted based on the user's age and / or the number of failures in the current scanning cycle.
[0081] The multimodal fusion localization algorithm employs an extended Kalman filter framework. The state vector includes position, velocity, and attitude quaternions, as well as IMU zero bias. Visual feature points and optical markers serve as observations. The prediction phase utilizes IMU median integration; the update phase fuses three observations: ORB visual feature point reprojection error, optical marker reprojection error, and ZUPT zero-velocity pseudo-observation. The inertial measurement integral is used as the predicted value, and the observation covariance weights are dynamically adjusted based on the current visual feature richness.
[0082] When there are enough visual feature points, visual observation is the primary method.
[0083] When visual feature points are insufficient, observation is mainly based on optical markers.
[0084] In some implementations, the extended Kalman filter (EKF) fuses three observations:
[0085] Visual feature point observation (30Hz): ORB feature matching, PnP solution;
[0086] Optical marker observation (30Hz): Detect the center of the marker and solve for PnP;
[0087] IMU Zero Speed Correction (ZUPT): The speed of the detection device is set to zero when it is stationary.
[0088] The voice interaction module employs a hierarchical finite state machine and extracts user operation behavior features in real time. These features include movement speed, jitter amplitude, response latency, number of erroneous entries, and number of pause timeouts. In some implementations, these real-time operation behavior features are quantified.
[0089] Movement speed: Moving average of IMU position differential mode (window 0.5s), eliminating high-frequency noise;
[0090] Shaking amplitude: The RMS value of the high-frequency component (10-30Hz) of the accelerometer, reflecting the degree of physiological tremor in the hand;
[0091] Response latency: The time difference between the end of a voice command and the start of movement by the user;
[0092] Error entry count: The number of times the current area has been switched to a non-target area;
[0093] Number of times the stay exceeds the time limit: The number of times the stay in a non-target area exceeds 3 seconds.
[0094] User behavior profiles are dynamically constructed based on an online Bayesian classifier. The guidance mode is adaptively selected according to user type, including a gamified mode for children, a concise mode for adults, an assistance mode for the elderly, and an assistance mode for unstable users. Different modes differ in voice style, level of detail, speech rate, and tactile feedback. In some implementations, the posterior probability of user type (novice / expert / unstable / child) is calculated, updated every 2 seconds, and locked when the confidence level is >0.8.
[0095] A gradual bootstrapping strategy is implemented based on the number of failures, as follows:
[0096] Number of failures Guiding strategy 0 Concise command (only announces the area name) 1 Detailed instructions (with added directional hints) 2 Breaking down the steps ("Stop first, listen to me...") 3 Visual assistance (indicator lights flashing to indicate direction) ≥4 Detailed tutorial (animation played on the mobile app)
[0097] Image quality assessment algorithms include the calculation of the following diagnostic-specific features: tooth margin contrast, gingival region color distribution, highlight area ratio, and consistency with dentist annotations. In some implementations, the image acquisition module has a resolution of at least 1080P.
[0098] The image quality assessment algorithm also includes an illumination adaptive processing module, which first identifies the illumination type through a lightweight CNN and then calls the corresponding enhancement operator.
[0099] The edge computing module also runs a scan integrity detection algorithm. This algorithm maintains a coverage matrix and region topology map of six standard oral regions, tracks scanned and unscanned regions in real time, detects the anatomical rationality of the scan path, and calculates a comprehensive integrity score. The scan is considered complete when the score reaches a threshold. It should be noted that the scan integrity detection algorithm uses a region coverage matrix (6 regions × 2 qualified images), defines reasonable paths based on the topology map (distance between adjacent regions ≤ 2), and detects path violations (forced redirection when distance ≥ 4).
[0100] The scan integrity detection algorithm maintains a six-zone coverage matrix and regional topology map of the oral cavity based on the FDI tooth position marking method. Graph nodes represent the six zones, and edge weights are defined as anatomically adjacent distances. The algorithm verifies the scan path in real time. If a cross-zone jump distance ≥4 is detected (e.g., jumping directly from the left side of the maxilla to the right side of the mandible), it is determined to be an anatomical path violation, triggering a forced redirection voice prompt to ensure the scan logic complies with dental examination standards.
[0101] It should be noted that the specific model and specifications to be adopted need to be determined based on the actual specifications of the device. The specific selection and calculation methods use existing technology in this field, so they will not be elaborated here.
[0102] The above methods can achieve the following:
[0103] Fully automated data acquisition: the success rate has increased from 60% in traditional solutions to over 95%, and the scrap rate has decreased from 40% to below 5%;
[0104] Real-time quality assurance: Multi-dimensional quality assessment + dynamic thresholds ensure that every image meets the requirements for medical analysis;
[0105] Complete area coverage: Visual-inertial-marker fusion accurately tracks scanning progress and avoids omissions;
[0106] Edge intelligence: Lightweight models are deployed on the device side with a response latency of <100ms, protecting privacy.
[0107] In some implementations, high-end configurations are used to suit high-precision scenarios such as medical-grade remote diagnosis and orthodontic monitoring, namely, a 4K+4TOPS NPU is employed.
[0108] Hardware configuration: Sony IMX415 4K sensor, Horizon Journey 2 NPU (4TOPS), 6-axis IMU (ICM-42688), dual microphone array, 2000mAh battery, 8 infrared optical markers;
[0109] Algorithm execution: YOLO-Nano region detection (30fps), EKF triadic fusion localization (error <0.3mm), multi-dimensional quality assessment + diagnostic features, dynamic threshold adjustment, and progressive voice guidance;
[0110] Usage flow: The user picks up the device to wake it up, and a voice prompt says "Please move the device to the upper left outer side." The device detects the position and quality in real time, and automatically triggers shooting when the threshold is reached, with a voice feedback of "Shooting successful." This cycle repeats until all 6 areas are completed, generating a report that is synced to the mobile app. Children's mode features a slower speech rate and gamified incentives; adult mode is simple and efficient.
[0111] In some implementations, an economical configuration is adopted to suit mass-market home oral health screening, significantly reducing costs; this includes the use of a 1080P+ESP32-S3.
[0112] Hardware configuration: SmartSens SC2310 1080P sensor, ESP32-S3 main controller (dual-core 240MHz + AI accelerator), basic 6-axis IMU, single microphone, 1000mAh battery, 4 optical markers (simplified layout), fixed focal length lens.
[0113] Algorithm Adjustment:
[0114] Region recognition: Based on optical marker geometric localization + traditional CV segmentation, replacing YOLO-Nano;
[0115] Quality assessment: Simplified to sharpness (Laplace variance) + exposure + motion detection;
[0116] Positioning: Primarily based on optical marker points (PnP), supplemented by IMU; error <0.8mm.
[0117] Voice guidance: Retains the hierarchical FSM, but the offline vocabulary is simplified;
[0118] Automatic Trigger: Only stability and basic quality are detected. A multi-condition joint judgment mechanism is used, and the trigger threshold is not a fixed value, but is dynamically adjusted based on PID control. If the number of retries increases due to image blurring or pose jitter during the current scanning cycle, the system will automatically lower the sharpness and stability thresholds (relaxing the entry conditions) and trigger a voice prompt to the user to slow down. If the user is elderly and the amplitude of operation jitter continuously exceeds the threshold, the system will temporarily lower the movement speed threshold requirement. This mechanism ensures that users of different skill levels can successfully complete data acquisition.
[0119] Performance metrics: Area recognition accuracy 85-90%, automatic trigger success rate 85-90%, battery life 20-25 scans, retail price can be controlled between 250-400 yuan.
[0120] Use cases: daily oral health monitoring at home, caries risk screening, evaluation of children's brushing effectiveness, etc., with image quality meeting the basic needs of remote initial diagnosis.
[0121] In some implementations, it is applied to assistive scenarios for elderly users:
[0122] The device automatically identifies elderly users (through App settings or voice feature analysis), increases the volume to maximum, reduces the speech rate by 30%, adds haptic feedback (vibration prompts), and adds a confirmation step to the voice guidance: "Did you understand? Please say 'okay' to continue." It also supports voice commands such as "Next," "Replay," and "Pause."
[0123] In some implementations, it is applied to gamified scenarios for children:
[0124] When a 6-year-old child uses the device for the first time, the device plays a cartoon voice: "Hello, little friend! I'm your little tooth assistant. Shall we take a picture of our teeth together?" Encouraging sound effects and applause are played after each area is completed. After completing 3 areas, the voice prompts: "Great job! You've already taken half the picture!" Victory music plays when all areas are completed. The total time is approximately 45 seconds.
[0125] The functional principle of this invention can be explained through the following operational methods:
[0126] The first step involves guiding the user to align the device with the target oral cavity area via voice commands. The guidance content is generated based on the gaps in the area coverage matrix from the previous step, prioritizing guidance for areas that have not yet met the target criteria.
[0127] The second step involves real-time acquisition of preview images, device motion data, and optical marker observation data. The data acquisition process includes a hardware synchronization latching mechanism to ensure that the timestamp alignment error between the visual frame and the inertial data is less than 1ms.
[0128] The third step involves fusing visual feature points, optical markers, and inertial measurement data using extended Kalman filtering, and introducing prior knowledge of oral anatomy as an inequality constraint (e.g., working distance 5-25mm, absolute pitch angle ≤60°). When the predicted state violates the constraints, the covariance is increased through pseudo-observations to pull back into the feasible region and prevent pose drift to the external space of the oral cavity.
[0129] To further explain, the precise pose of the device in the oral cavity coordinate system is obtained and mapped to the corresponding oral cavity region. It should be noted that prior knowledge of oral anatomy is introduced: a region topology map and constraint inequalities (depth constraint 5-25mm, attitude angle constraint |θ|≤60°) are constructed and used as pseudo-observations in the EKF update. When the predicted pose violates the constraints, the observation noise covariance is increased.
[0130] It is worth noting that a three-level working mode switching is adopted:
[0131] Normal mode (feature points ≥ 30): Ternary fusion, error < 0.3mm;
[0132] Degradation mode (10-29 points): optical markers as the primary method + IMU, error 0.3-0.8mm;
[0133] Blind push mode (<10 points): pure IMU integration + motion model, error 0.8-2mm.
[0134] The fourth step is to calculate the current image quality score based on a multi-dimensional quality assessment model. This model includes diagnostic features and is adaptively enhanced according to the lighting type. If the lighting type is classified as underexposed by the CNN, the local contrast enhancement operator based on CLAHE is automatically invoked.
[0135] The fifth step is to automatically trigger shooting when the image quality score, region matching degree, and device pose stability all meet the dynamic threshold.
[0136] Step 6: Evaluate the shooting results. If they are unsatisfactory, perform a progressive voice-guided reshoot based on the number of failures. If they are satisfactory, update the area coverage matrix.
[0137] The seventh step is to check the integrity of the scan, including regional coverage and path compliance. When the overall integrity score reaches the threshold, the scan ends, a report is generated, and the scan is synchronized to the cloud.
[0138] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. An intelligent voice-guided oral scanning device, comprising a grip (1) and a scanning head (2), characterized in that, The grip (1) and the scanning head (2) are connected by a double-beveled buckle (3) and a metal positioning pin, and the scanning head (2) is provided with: The voice interaction module is used to output scanning guidance voice and receive user voice feedback; Image acquisition module, the image acquisition module being used to acquire oral cavity images; A motion sensing module, which is used to collect device motion data; Optical markers are set at the edge of the device's scanning head to provide artificial feature-assisted positioning when visual features are sparse; The edge computing module is used to run: Oral cavity region detection algorithm, which is used to identify the currently scanned oral cavity region; Image quality assessment algorithm, used to evaluate image quality in real time; A multimodal fusion localization algorithm is used to fuse the visual data from the image acquisition module, the observation data from the optical markers, and the inertial data from the motion sensing module to obtain the real-time pose of the device in the oral cavity coordinate system, and to map the real-time pose to the corresponding oral cavity region. An automatic triggering algorithm is used to automatically trigger the image acquisition module to take pictures when the image quality, the oral cavity region, and the stability of the real-time pose all meet preset conditions.
2. The intelligent voice-guided oral scanning device according to claim 1, characterized in that, The scanning head (2) is also equipped with an auxiliary device (4), which includes a camera (41), a fill light (42), a microphone (43), a speaker (44), and an indicator light (45).
3. The intelligent voice-guided oral scanning device according to claim 2, characterized in that, The optical markers include eight miniature infrared reflective markers arranged on the scanning head (2).
4. The intelligent voice-guided oral scanning device according to claim 3, characterized in that, The eight micro infrared reflective markers are distributed in a non-coplanar spatial distribution, with four micro infrared reflective markers located on the front face and the other four located on the side arc face.
5. The intelligent voice-guided oral scanning device according to claim 4, characterized in that, The multimodal fusion localization algorithm employs an extended Kalman filter framework, where visual feature points and optical marker points are used as observations, inertial measurement integrals are used as predictions, and the observation covariance weights are dynamically adjusted based on the current visual feature richness. When there are enough visual feature points, visual observation is the primary method. When visual feature points are insufficient, observation is mainly based on optical markers.
6. The intelligent voice-guided oral scanning device according to claim 5, characterized in that, The automatic triggering algorithm adopts a multi-condition joint judgment mechanism. The preset conditions include: the image quality score is higher than the dynamic threshold, the matching degree between the current oral cavity region and the target region is higher than the preset matching threshold, the device movement speed is lower than the preset speed threshold, and the stable duration of the device at the current position exceeds the preset time threshold. The dynamic threshold is adaptively adjusted based on the user's age and / or the number of failures in the current scanning cycle.
7. The intelligent voice-guided oral scanning device according to claim 6, characterized in that, The voice interaction module uses a hierarchical finite state machine and extracts user operation behavior features in real time. The operational behavior characteristics include movement speed, jitter amplitude, response delay, number of erroneous entries, and number of timeouts. User behavior profiles are dynamically constructed based on an online Bayesian classifier. The guidance mode is adaptively selected according to user type, including gamified mode for children, concise mode for adults, assistance mode for the elderly, and assistance mode for unstable users. Different modes have differences in voice style, level of detail, speech rate, and tactile feedback. Implement a progressive bootstrapping strategy based on the number of failures.
8. The intelligent voice-guided oral scanning device according to claim 7, characterized in that, The image quality assessment algorithm includes the calculation of the following diagnostic-specific features: tooth edge contrast, gingival region color distribution, highlight area ratio, and consistency with dentist annotations. The image quality assessment algorithm also includes an illumination adaptive processing module, which first identifies the illumination type through a lightweight CNN and then calls the corresponding enhancement operator.
9. The intelligent voice-guided oral scanning device according to claim 8, characterized in that, The edge computing module also runs a scan integrity detection algorithm. This algorithm maintains the coverage matrix and region topology map of six standard regions of the oral cavity, tracks the scanned and unscanned regions in real time, detects the anatomical rationality of the scan path, and calculates a comprehensive integrity score. When the score reaches a threshold, the scan is considered complete.
10. An automatic data acquisition method for an intelligent voice-guided oral scanning device as described in claim 9, characterized in that, The automatic data acquisition method includes the following steps: Step S1: Guide the user via voice to aim the device at the target oral cavity area; Step S2: Real-time acquisition of preview images, device motion data, and optical marker observation data; Step S3: Based on the extended Kalman filter, the visual feature points, optical marker points and inertial measurement data are fused to obtain the precise pose of the device in the oral cavity coordinate system and map it to the corresponding oral cavity region. Step S4: Calculate the current image quality score based on a multi-dimensional quality assessment model, wherein the quality assessment model includes diagnostic-specific features and is adaptively enhanced according to the illumination type; Step S5: When the image quality score, region matching degree, and device pose stability all meet the dynamic threshold, shooting is automatically triggered; Step S6: Evaluate the shooting results. If the results are unsatisfactory, perform a progressive voice-guided reshoot based on the number of failures. If the results are satisfactory, update the area coverage matrix. Step S7: Check the scan integrity, including area coverage and path compliance rate. When the overall integrity score reaches the threshold, the scan ends, a report is generated, and synchronized to the cloud.
Citation Information
Patent Citations
An AR-based root canal monitoring method and system
CN110269715B
Method for intraoral 3D scanning and intraoral imaging device
CN112912933B