Handheld all-in-one multimodal diabetic foot screening terminal and cloud-edge collaborative screening system
By combining a handheld integrated multimodal diabetic foot screening terminal with a cloud-edge-device collaborative system, the instability and error issues of multimodal screening in handheld mobile scenarios are solved, achieving efficient and reliable diabetic foot screening, which is suitable for primary care clinics and home follow-ups.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-26
AI Technical Summary
Existing multimodal diabetic foot screening devices suffer from problems such as unstable imaging, large quantitative errors, insufficient coverage, limited edge computing power, and lack of closed-loop optimization in handheld mobile scenarios, making it difficult to meet the high-frequency screening needs of primary clinics, bedside and home follow-up.
A handheld integrated multimodal diabetic foot screening terminal was designed, which combines near-infrared multispectral imaging, infrared thermal imaging and ToF depth perception. Through image stabilization gating, multi-frame anti-shake, three-dimensional geometric normalization and lightweight edge inference, it realizes synchronous triggering and unified time reference labeling of multimodal data, and optimizes the model through cloud-edge-device collaborative mechanism.
It significantly improved the consistency and availability of handheld data collection, reduced quantitative errors, enabled rapid offline screening, and enhanced robustness in long-tail scenarios through cloud-based iterative optimization.
Smart Images

Figure CN122074899A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of biomedical engineering, optoelectronic detection technology and artificial intelligence medical image analysis technology, and relates to a handheld integrated multimodal diabetic foot screening terminal and a cloud-edge-device collaborative screening system. Background Technology
[0002] Diabetic foot ulcer (DFU) is a common and serious chronic complication of diabetes, characterized by high disability and recurrence rates. Before ulceration is clinically visible, abnormalities in foot microcirculation perfusion, altered local tissue oxygenation, and abnormal temperature distribution often already exist. Therefore, non-invasive, rapid, and repeatable risk screening is necessary in the early stages.
[0003] Existing commonly used risk assessment methods include monofilament tactile testing, ankle-brachial index measurement, and transcutaneous oxygen partial pressure testing, but these methods have problems such as strong subjectivity, long time consumption, limited testing range, or insensitivity to early functional abnormalities, making it difficult to meet the high-frequency screening needs of primary care clinics, bedside and home follow-up.
[0004] Near-infrared optical imaging can utilize the absorption characteristics of hemoglobin in the visible-near-infrared band to retrieve hemodynamic parameters; infrared thermal imaging can be used to obtain the temperature distribution of the sole surface and identify inflammatory "hot spots" or ischemic "cold spots". Multimodal fusion can provide richer physiological information, but existing multimodal devices are mostly desktop or cart-type, relying on fixed supports or controlled environments, making it difficult to deploy them in mobile scenarios.
[0005] As multimodal screening evolves towards handheld operation, new key technical challenges will be introduced: (1) motion blur and cross-modal registration errors caused by hand shake; (2) changes in light intensity and quantitative deviations in reflectivity caused by distance and angle fluctuations (distance square attenuation and Lambert cosine effect); (3) edge darkening, field of view loss and artifacts caused by foot surface curvature and occlusion; (4) omission of high-risk areas due to incomplete coverage of single-shot images; (5) difficulty in deploying complex models due to limited computing power and power consumption on the device side; (6) lack of a continuous optimization mechanism driven by difficult samples after the device leaves the factory.
[0006] Therefore, there is an urgent need for an integrated multimodal screening terminal and method for handheld mobile scenarios, which can achieve "image stabilization gating triggering + multi-frame anti-shake + three-dimensional geometric normalization" at the acquisition end, and combine lightweight inference on the edge side and uncertainty-driven cloud-edge-device collaborative mechanism to continuously improve model performance while ensuring privacy and availability. Summary of the Invention
[0007] In view of this, the purpose of this invention is to provide a handheld integrated multimodal diabetic foot screening terminal and a cloud-edge-device collaborative screening system, which can overcome the problems of unstable imaging, large quantitative error, insufficient coverage, limited computing power on the edge, and lack of closed-loop optimization in existing multimodal diabetic foot screening technologies in handheld mobile scenarios.
[0008] To achieve the above objectives, the present invention provides the following technical solution: A handheld integrated multimodal diabetic foot screening terminal includes: a handheld shell (110) with an imaging window (111) at its front end; a near-infrared multispectral imaging unit (120), an infrared thermal imaging unit (130), and a ToF depth sensing unit (140) whose fields of view at least partially overlap; an IMU (145) for outputting terminal attitude-related angular velocity and acceleration information; a ring-shaped multi-band LED illumination module (150) for providing controllable pulse illumination; an embedded processing and storage unit (160) including a processor and a neural network acceleration unit for performing stability determination, synchronous triggering, continuous shooting frame selection / fusion, three-dimensional modeling, surface normalization, feature extraction, end-side inference, and uncertainty assessment; a human-computer interaction module (170) for displaying shooting guidance and screening results and outputting audio-visual or vibration feedback; a wireless communication module (180) and a power management module (190) for secure communication with the cloud and powering each module; the embedded processing and storage unit (160) 0) Includes a synchronization trigger and timestamp management module, used to uniformly label multispectral, thermal imaging and depth data with a time reference, and trigger quality control failure when a time deviation exceeds a preset threshold; the ring multi-band LED lighting module (150) is equipped with a diffuse reflection uniform light structure and / or a polarization suppression structure, and the embedded processing and storage unit (160) adaptively adjusts the LED driving current and exposure / integration time according to the real-time distance measured by ToF to suppress near-distance overexposure and far-distance underexposure; the human-computer interaction module (170) displays a foot coverage heat map or coverage indicator bar in scanning mode, and prohibits triggering acquisition and outputs prompts when incomplete coverage, angle deviation or jitter exceeds the limit; the wireless communication module (180) uploads difficult sample data packets only when the preset network conditions are met, and the difficult sample data packets are at least de-identified and encrypted before being uploaded; the encryption process can be selected from one or more of end-to-end symmetric encryption, asymmetric key negotiation or homomorphic encryption.
[0009] Furthermore, the terminal performs the following steps: S1: Terminal initialization and multimodal self-test: Perform power-on self-test on the near-infrared multispectral imaging unit, infrared thermal imaging unit, ToF depth sensing unit, IMU and ring multiband LED lighting module, and load the terminal inference model and acquisition control parameters; S2: Multidimensional Imaging Guidance and Stability Trigger Judgment: In preview mode, the ToF depth sensing unit acquires the distance field and local normal vector distribution of the sole surface in real time, while the IMU acquires the terminal angular velocity and acceleration information simultaneously; guidance information such as distance suitability, angle perpendicularity, and jitter status is displayed on the human-computer interaction interface; a joint stability criterion including distance deviation, incident angle deviation, and motion amplitude is constructed, and an acquisition trigger signal is generated if and only if the target area of the sole is within a preset distance range, the angle between the imaging optical axis and the local normal of the sole is less than a preset angle threshold, and the motion amplitude detected by the IMU is continuously lower than the stability threshold within a preset time window; S3: Multimodal short-time burst shooting and intelligent frame selection: In response to the acquisition trigger signal, the ring multi-band LED illumination module is controlled to flash illumination according to the preset spectral timing, and each imaging unit is synchronously triggered to perform short-time multi-frame burst shooting; the quality control of the burst shooting frame sequence is performed on the end side, and frames with motion blur, occlusion, specular reflection overexposure, saturation or field of view artifacts are removed, and the best frame with the highest quality score is selected or lightweight multi-frame fusion is performed to generate a multimodal raw dataset, and the acquisition environment parameters and timestamps are recorded; S4: 3D Modeling, Texture Mapping and Surface Normalization: Reconstruct a 3D mesh model of the foot based on the ToF depth map; use calibration parameters to map multispectral and thermal images as textures onto the 3D mesh surface; unfold and map the textured 3D model to a standard 2D plane coordinate system, and perform distance squared compensation and incident angle compensation based on geometric information to generate a normalized reflectance spectrum and temperature distribution map. At the same time, generate low-confidence masks for areas with excessive incident angles or low depth confidence, and use them for weight reduction or interpolation repair in subsequent analysis. S5: Anatomical zoning and feature extraction: Key anatomical regions such as heel, metatarsal heads, arch, and toes are automatically marked on the standard two-dimensional plane, and near-infrared multispectral tissue hemodynamic features and infrared thermographic features are extracted in each region. S6: End-side reasoning and uncertainty assessment: Integrating multispectral features, thermal features, and ToF geometric features, inputting a lightweight artificial intelligence model deployed on the terminal neural network acceleration unit, and outputting the risk level of diabetic foot, the location of suspected lesions, and uncertainty indices representing the confidence level of prediction.
[0010] Furthermore, in step S2, the joint stability criterion employs a multi-level AND logic gating mechanism, including at least: Distance gating: This refers to the distance statistics of the target area measured by Time-of-Flight (ToF). For the preset target distance; when The time-distance gating is true; Angle gating: ,in The unit vector of the imaging optical axis. Let be the unit vector of the local normal vector of the foot obtained by fitting the ToF point cloud; when The time angle gating is true; Jitter gating: within a preset time window Internally, the angular velocity vectors of the IMU are respectively... With acceleration vector Calculate the time-averaged amplitude or root mean square amplitude when and The jitter gate is true; Only if the above gating conditions are all true and the duration exceeds The acquisition trigger signal is generated in time, and when any gate is false, the operator is prompted to adjust the distance, angle or remain still through display, sound and light or vibration feedback.
[0011] Furthermore, in step S3, short burst shooting and intelligent frame selection include: N frames of multispectral images, thermal images, and depth maps were continuously acquired at a preset frame rate. ; Calculate a sharpness score for each frame. Inter-frame displacement ,in Obtained based on Laplace variance or gradient magnitude Obtained based on optical flow estimation or IMU time synchronization integration; Maximum inter-frame displacement ,when When the signal-to-noise ratio is not greater than a preset micro-motion threshold, registration and weighted fusion are performed on multiple frames to improve the signal-to-noise ratio; when Select when the value is greater than the preset micro-motion threshold The frame with the highest saturation ratio and the lowest saturation ratio is selected as the best frame.
[0012] Furthermore, in step S3, the quality control includes at least determining and triggering a resampling prompt for any of the following: motion blur, missing foot area, occlusion, specular highlights, overexposure saturation, low depth confidence, or multimodal time synchronization anomaly.
[0013] Furthermore, in step S4, surface normalization specifically includes: Constructing 3D points to two-dimensional parametric plane coordinates mapping And map the multimodal texture onto a two-dimensional plane; For multispectral band pixel values Obtained by distance square compensation and Lambert cosine compensation : , ,in This is the actual imaging distance. To calibrate the distance, For local incident angle, To prevent the lower limit of the denominator from being too small; when When the preset effective threshold is exceeded or the ToF depth confidence is lower than the threshold, the corresponding pixel is marked as a low-confidence region and its weight is reduced in subsequent analysis or repaired by interpolation based on neighborhood information.
[0014] Furthermore, in step S6, the uncertainty index is obtained based on at least one uncertainty estimation strategy, including Monte Carlo Dropout, deep integration, or confidence estimation based on output entropy; and the terminal dynamically selects a low-power fast model or a high-precision enhancement strategy based on the stability score and the uncertainty index.
[0015] This invention also provides a multimodal diabetic foot cloud-edge-device collaborative screening system. This system employs the screening terminal and execution steps described above, and further includes a cloud service module and an application terminal module. The cloud service module receives selectively uploaded difficult sample data packets and follow-up tags from the terminal, performs model training, evaluation, and version management, and generates a lightweight model update package adapted for terminal deployment. The application terminal module allows medical personnel or subjects to view screening results, historical trends, and early warning information. The cloud service module sends the model update package to the handheld multimodal intelligent screening terminal, which performs integrity verification and supports rollback to the previous stable version. The system execution steps also include: S7: Uncertainty-driven cloud-edge collaborative closed loop: Based on the uncertainty index, quality control results, and acquisition environment parameters, hierarchical screening is performed. When the uncertainty exceeds a preset threshold or a preset quality control rule is triggered, the desensitized and encrypted difficult sample data packets are selectively uploaded to the cloud. After retraining and model compression, the cloud sends updated model parameters to the terminal, which completes version switching during idle periods and supports rollback.
[0016] Furthermore, the cloud service module is also used to analyze the quality control failure logs uploaded by the terminal, and update the stability threshold, time window parameters or guidance strategy on the terminal side accordingly, forming a remote optimization of the acquisition control logic.
[0017] Furthermore, the lightweight model update package is obtained through teacher-student knowledge distillation and is further pruned and quantized to adapt to the neural network acceleration unit architecture of the terminal.
[0018] The beneficial effects of this invention are as follows: Compared with the prior art, the present invention has at least the following beneficial effects: (i) By using stability gating triggering and Burst multi-frame anti-shake, the consistency and availability of handheld acquisition are significantly improved; (ii) By using 3D modeling, unfolding and geometric compensation, the quantitative error introduced by distance / angle fluctuation and foot surface is reduced, and the comparability of cross-follow-up is improved; (iii) Offline rapid screening is achieved through a lightweight edge model, while uncertainty-driven selective cloud migration reduces privacy and bandwidth pressure; (iv) The continuous iteration of the model and threshold is achieved through cloud-edge-edge closed loop, improving the robustness of long-tail scenarios.
[0019] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 This is a schematic diagram of the handheld device's appearance / overall structure. Figure 2 This is a hardware block diagram / hardware connection diagram for a handheld terminal. Figure 3 Flowchart for handheld shooting; Figure 4 A schematic diagram illustrating the division of key anatomical regions of the foot and multimodal registration; Figure 5 This is a schematic diagram of the process for geometric correction of the plantar surface and calculation of normalized reflectivity based on ToF depth information; Figure 6 This is a schematic diagram of the cloud-edge-device collaborative system for screening diabetic foot. Detailed Implementation
[0021] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.
[0022] Figure 1 The figure shows a schematic diagram of the appearance / overall structure of the handheld device. The present invention provides a handheld multimodal intelligent screening terminal (100) for diabetic foot based on the fusion of near-infrared multispectral imaging, infrared thermal imaging and ToF depth information. The terminal (100) adopts an integrated and modular design, integrating the functions of data acquisition, lighting, processing, interaction, communication and power supply into a single compact shell, which is convenient for one-handed holding and use in bedside / outpatient and primary care mobile scenarios.
[0023] Specifically, the terminal (100) includes a handheld housing (110) and an imaging window (111) disposed at the front end of the housing. The handheld housing (110) may be designed as a strip-shaped or "gun-shaped" grip structure, and the outer surface may be provided with anti-slip texture or soft covering to improve grip stability and operational safety; without limiting the present invention, the handheld housing (110) may further adopt a dustproof and splashproof structure to adapt to clinical and bedside use environments.
[0024] Specifically, the imaging window (111) is adjacent to a near-infrared multispectral imaging unit (120), an infrared thermal imaging unit (130), and a ToF depth sensing unit (140). The fields of view of the imaging units overlap at least partially within the target area on the sole of the foot to reduce cross-modal parallax and improve subsequent registration and fusion accuracy. Preferably, the near-infrared multispectral imaging unit (120), the infrared thermal imaging unit (130), and the ToF depth sensing unit (140) are arranged in a compact or approximately shared field of view arrangement to ensure that multimodal corresponding information can be obtained for the same sole area at the same acquisition time.
[0025] Specifically, a ring-shaped multi-band LED illumination module (150) is provided around the imaging window (111) to provide controllable illumination to the area to be tested on the sole of the foot. The LED illumination module (150) may include multiple near-infrared narrowband light sources and optionally include a visible light auxiliary illumination source, supporting group illumination according to a preset time sequence, pulse illumination, or brightness adjustment, thereby reducing motion artifacts and improving illumination uniformity in handheld shooting scenarios. Preferably, the LED illumination module (150) may be configured with a diffuser or anti-glare structure to reduce the impact of strong reflections from the sole surface on image quality.
[0026] Specifically, to achieve handheld image stabilization and attitude gating, an inertial measurement unit (IMU) (145) is fixed inside the terminal (100) to output angular velocity and acceleration information to characterize handheld jitter and attitude changes. Preferably, the IMU (145) is located in a relatively stable position inside the housing (110) (e.g., near the device's center of mass or near the main control circuit board) to improve jitter estimation consistency and facilitate fusion with ToF depth information.
[0027] Specifically, the terminal (100) further includes an embedded processing and storage unit (160), which is used for unified control and data management of multimodal acquisition, and performs functions such as synchronization triggering, timestamp annotation, caching, preprocessing, 3D modeling and edge inference. Preferably, the embedded processing and storage unit (160) may include a synchronization triggering and timestamp management function module, which timestamps the data from the near-infrared multispectral imaging unit (120), the infrared thermal imaging unit (130) and the ToF depth sensing unit (140) using a unified time base to ensure the temporal consistency of Burst continuous shooting, cross-modal texture mapping and 3D modeling.
[0028] Specifically, the terminal (100) includes a human-computer interaction module (170), which includes at least a display screen (171) and a prompting component (sound, light, and vibration). The display screen (171) is used to display shooting guidance information (such as distance / angle / shake indication), acquisition status, multimodal preview, and screening results; the prompting component is used to provide immediate feedback to the operator in states such as gated failure, quality control abnormality, or acquisition completion, so as to reduce the learning cost of operation and improve acquisition consistency.
[0029] Specifically, the terminal (100) further includes a wireless communication module (180) and a power management module (190). The wireless communication module (180) is used to communicate securely with external devices or cloud service modules to achieve data uploading, model / threshold policy distribution, and remote maintenance; the power management module (190) is used for battery management, charge and discharge protection, and multi-channel voltage regulation to provide stable power support for the imaging unit, lighting module, processing unit, and interactive communication module.
[0030] In this embodiment, the embedded processing and storage unit (160) may be a system-on-a-chip or an embedded motherboard with CPU and neural network acceleration capabilities; the IMU (145) may be a 6-axis or 9-axis MEMS device; the ToF depth sensing unit (140) may output depth maps and depth confidence information; and the infrared thermal imaging unit (130) may be an uncooled focal plane detector module.
[0031] Please see Figure 2 , Figure 2 This illustration shows the hardware connection relationships and data / control interaction architecture between functional modules of the handheld multimodal intelligent screening terminal (100) described in this invention. The terminal (100) forms a closed-loop hardware system of "sensor acquisition - synchronous control - edge computing - interactive feedback - secure communication - power management" within a single compact housing to adapt to the real-time acquisition and instant prompting requirements in handheld mobile scenarios.
[0032] Specifically, the terminal (100) includes an embedded processing and storage unit (160), which serves as the main control and computing core. It establishes data connections with the near-infrared multispectral imaging unit (120), the infrared thermal imaging unit (130), the ToF depth sensing unit (140), and the inertial measurement unit (IMU) (145) to receive multimodal raw data, perform caching and preprocessing, and generate intermediate results required for shooting stability assessment, 3D modeling, and end-side inference. Preferably, the near-infrared multispectral imaging unit (120), the infrared thermal imaging unit (130), and the ToF depth sensing unit (140) are connected to the embedded processing and storage unit (160) through high-speed image / data interfaces to meet the low-latency transmission of multi-channel data. The IMU (145) can be connected to the embedded processing and storage unit (160) through a low-speed serial bus to periodically output angular velocity and acceleration data for jitter estimation and attitude gating.
[0033] Specifically, the embedded processing and storage unit (160) establishes a control connection with the ring-shaped multi-band LED lighting module (150) to realize pulse lighting, brightness adjustment, and timing synchronization by band / group, thereby uniformly scheduling the lighting and reducing motion artifacts in Burst mode. Preferably, the terminal (100) can be equipped with a lighting driving circuit or a constant current driving module to support the control of the duty cycle, pulse width, or current amplitude of LEDs of different bands, thereby realizing distance-adaptive exposure and adaptive adjustment of lighting energy.
[0034] Specifically, the embedded processing and storage unit (160) establishes a bidirectional connection with the human-machine interaction module (170), which includes a display screen (171) and a prompting component. The embedded processing and storage unit (160) is used to output distance / angle / jitter guidance information, acquisition status, multimodal preview and screening results to the display screen (171), and can control the prompting component to output prompts in states such as gate control failure, quality control abnormality or acquisition completion; the human-machine interaction module (170) can transmit operator input back to the embedded processing and storage unit (160) to realize the switching between point-and-shoot mode and scanning mode and the acquisition process control.
[0035] Specifically, the embedded processing and storage unit (160) establishes a bidirectional connection with the wireless communication module (180) for secure communication with the cloud service module when the network is available, enabling data packet uploading, model / threshold policy distribution, and remote maintenance. Preferably, the terminal (100) can perform data desensitization, compression, and encryption on the embedded processing and storage unit (160), and selectively send uploaded content based on uncertainty thresholds, quality control rules, and authorization policies to reduce bandwidth consumption and improve privacy protection.
[0036] Specifically, the terminal (100) includes a power management module (190), which is electrically connected to the near-infrared multispectral imaging unit (120), the infrared thermal imaging unit (130), the ToF depth sensing unit (140), the IMU (145), the lighting module (150), the embedded processing and storage unit (160), the human-computer interaction module (170), and the wireless communication module (180), respectively, for providing charge and discharge management, overcurrent / overvoltage protection, and multi-channel regulated power supply, thereby ensuring the terminal's battery life and operational stability in handheld mobile scenarios. Preferably, the power management module (190) can perform hierarchical management of the power consumption of the lighting module (150) and the embedded processing and storage unit (160), and dynamically schedule the power supply and power consumption modes during standby, preview, and Burst acquisition and inference stages to balance response speed and battery life.
[0037] Preferably, the embedded processing and storage unit (160) includes a synchronization triggering and timestamp management module, which performs unified triggering control on the acquisition process from the near-infrared multispectral imaging unit (120), the infrared thermal imaging unit (130) and the ToF depth sensing unit (140), and timestamps the multimodal data based on a unified time base, so as to provide time consistency guarantee for Burst continuous shooting frame selection, cross-modal texture mapping and 3D modeling.
[0038] Without limiting the scope of the invention, the embedded processing and storage unit (160) may be a system-on-a-chip or an embedded motherboard with CPU and neural network acceleration capabilities; the wireless communication module (180) may support one or more of Wi-Fi, Bluetooth or cellular mobile communication; the power management module (190) may include power metering and temperature monitoring circuits for battery safety management; the sensing and acquisition unit may further include ambient light, temperature and humidity or distance auxiliary sensors to improve imaging adaptive control capabilities.
[0039] Please see Figure 3 , Figure 3 This paper illustrates the multimodal acquisition and screening method of the present invention in a handheld mobile scenario. The process follows the main steps of "stability gating trigger - Burst burst acquisition - quality control and frame selection / fusion - 3D modeling and surface normalization - lightweight inference on the device side - uncertainty-driven collaboration" to reduce imaging errors caused by hand shake, distance / angle fluctuations, and the curvature of the foot, and to achieve real-time screening under the condition of limited device side resources.
[0040] Specifically, step S1 is terminal initialization and self-test. The terminal completes power and sensor status detection, and loads the current version of the threshold strategy, acquisition mode parameters, and edge inference model; preferably, the terminal performs basic calibration or black level correction on each imaging unit before entering the acquisition preview mode to improve subsequent data consistency.
[0041] Specifically, step S2 involves handheld shooting guidance and stability assessment. During the preview phase, the terminal simultaneously acquires ToF depth information and IMU inertial information, and calculates the distance statistics to the target area. and combined with the preset target distance and permitted scope Distance gating is formed; local normal vectors and unit vectors are obtained based on ToF point cloud fitting. , and the unit vector of the imaging optical axis Calculate the included angle Angle gating is achieved; simultaneously, the angular velocity vector is output to the IMU. With acceleration vector In the time window The terminal calculates the average amplitude or root mean square amplitude to form a jitter gating system. Preferably, the terminal outputs the gating result as a distance bar, level / attitude indicator, jitter bar, or color prompt, and prompts the operator to adjust the distance, pitch / roll angle, or remain stationary when the gating is not satisfied.
[0042] Specifically, step S3 involves combined gating triggering and Burst burst shooting. When distance gating, angle gating, and shake gating simultaneously meet preset thresholds and the duration exceeds [a certain value], [the process is complete]. At that time, the terminal generates a data acquisition trigger signal and controls the ring-shaped multi-band LED lighting module to light up according to the preset band timing pulse, simultaneously triggering the near-infrared multispectral imaging unit, infrared thermal imaging unit, and ToF depth sensing unit to acquire data within a short time window. Frame multimodal data (of which Gated triggering and Burst acquisition can prevent jitter introduced by the operator when pressing buttons and reduce the number of unstable frames entering subsequent 3D processing and inference. Preferably, the terminal uses a unified time base for timestamping multimodal frames to ensure the temporal consistency of cross-modal data.
[0043] Specifically, step S4 involves end-side quality control and frame selection / lightweight fusion. The terminal calculates a sharpness score for each frame in the Burst sequence. And estimate the inter-frame displacement. (This can be obtained based on optical flow estimation or IMU time synchronization integration), and the maximum displacement is taken. Simultaneously, artifacts such as occlusion, specular highlights, overexposure / saturation, field-of-view loss, and low depth confidence are detected, forming quality control markers. Preferably, when When the signal-to-noise ratio is not greater than the micro-motion threshold and artifact detection passes, the terminal performs registration and weighted fusion on multiple frames to improve the signal-to-noise ratio and stability; when When the value exceeds the micro-motion threshold or artifacts are severe, the terminal selects... The frame with the highest score and the smallest artifacts is selected as the best frame and proceeds to the next stage of processing.
[0044] Specifically, step S5 involves 3D modeling, surface geometry correction, and standard plane unfolding. The terminal back-projects pixel coordinates from the ToF depth map to obtain a point cloud and reconstructs a 3D mesh model of the foot. Near-infrared multispectral and infrared thermal imaging textures are mapped onto the 3D mesh surface using calibration parameters, achieving cross-modal alignment. The terminal further constructs 3D points... to two-dimensional coordinates mapping The foot surface is unfolded to a standard two-dimensional plane using conformal mapping or a minimum deformation unfolding algorithm to facilitate subsequent ROI statistics and cross-period follow-up alignment. Preferably, the terminal calculates the actual imaging distance for each pixel / area. With the angle of incidence And compensate according to the square of the distance and the Lambert cosine; when If the confidence level of the ToF depth is too small or below the threshold, the corresponding region will be marked as a low-confidence region and reduced in weight, interpolated and repaired, or prompted for reshoot in subsequent analysis to avoid compensation divergence caused by edge incident angle.
[0045] Specifically, step S6 involves lightweight inference and result output on the terminal side. The terminal extracts multimodal features in a standard two-dimensional plane or three-dimensional grid coordinate system and calls a lightweight artificial intelligence model to output risk classification results, suspected lesion locations, or risk heatmaps. Preferably, the terminal dynamically selects a model or inference strategy based on stability score, quality score, and scene complexity to improve robustness in complex scenarios while ensuring real-time performance.
[0046] Specifically, step S7 involves uncertainty assessment and cloud-edge-device collaborative closed-loop. During inference, the terminal estimates prediction uncertainty and classifies samples according to uncertainty thresholds, quality control rules, and authorization strategies. For high-uncertainty or quality-controlled problematic samples, data packets are generated and desensitized, compressed, and encrypted. When the network is available and compliance requirements are met, these packets are selectively uploaded to the cloud service module. The cloud performs human-in-the-loop annotation and retraining based on the problematic samples, and generates update packages through distillation, pruning, and quantization, which are then distributed to the terminal. The terminal performs integrity verification on the update packages and supports rollback to the previous stable version to achieve continuous optimization of the acquisition control logic and the edge-side model.
[0047] Please see Figure 4 , Figure 4 This illustration demonstrates the division of key anatomical regions of the foot in a standard two-dimensional plane coordinate system and the spatial registration relationship of multimodal data according to the present invention. Figure 4 This is used to illustrate how, after acquiring near-infrared multispectral images, infrared thermal images, and ToF depth information through handheld acquisition, different modalities can be established in the same plantar region to provide a basis for subsequent ROI statistics, risk assessment, and cross-follow-up alignment.
[0048] Specifically, the terminal first acquires multimodal data of the target area on the sole of the foot, including near-infrared multispectral images, infrared thermal images, and ToF depth maps / point cloud information, and establishes spatial mapping relationships between each imaging unit based on calibration parameters. Preferably, the calibration parameters include intrinsic parameters of each imaging unit and extrinsic parameters between them, used to map pixel coordinates under different modalities to a unified sole coordinate system or a unified standard two-dimensional plane coordinate system, so as to reduce the impact of cross-modal parallax on the regional correspondence.
[0049] Specifically, such as Figure 4 As shown, the terminal anatomically divides the foot region on a standard two-dimensional plane or an unfolded foot plane. These anatomical divisions include at least the heel region, arch region, metatarsal head region, and toe region. The metatarsal head region can be further subdivided into the first to fifth metatarsal head regions, or other subdivision methods can be set according to clinical screening needs. Preferably, the divisions can be automatically located based on key points of the foot contour, geometric proportions, or preset templates, and the boundaries can be corrected using depth information to improve the stability of the divisions for different individual foot types.
[0050] Specifically, the terminal aligns near-infrared multispectral images, infrared thermal images, and depth information to the same partition frame through the spatial mapping relationship, enabling a one-to-one or computable correspondence between multimodal pixels / areas within any anatomical partition. This allows for the simultaneous statistical analysis of multispectral features, thermal features, and depth / geometric features within the same ROI. Preferably, the terminal can... Figure 4 The partitioning results are overlaid with a multimodal alignment effect for use in quality checks or for operator visualization.
[0051] Please see Figure 5 , Figure 5 This illustrates the process of surface geometry correction and normalized reflectivity calculation based on ToF depth information according to the present invention. Figure 5 This is used to illustrate how, under conditions where there are fluctuations in handheld shooting distance and incident angle, and the sole surface of the foot is a three-dimensional curved surface, the terminal can use depth information to estimate the imaging geometry, and perform geometric compensation and low-confidence region processing on the near-infrared multispectral imaging results, thereby improving the comparability of multimodal data under different shooting postures and the stability of subsequent screening results.
[0052] Specifically, the terminal acquires a ToF depth map and performs preprocessing, which may include noise suppression, hole filling, and confidence level filtering to obtain effective depth data for 3D reconstruction. Preferably, the terminal simultaneously acquires depth confidence information and marks pixels with confidence levels below a threshold as areas to be removed or repaired to reduce the impact of depth anomalies on geometric correction.
[0053] Specifically, the terminal backprojects pixel coordinates into a point cloud based on the depth map and reconstructs a 3D mesh model of the foot. Based on this, the terminal calculates the actual imaging distance $d$ and the incident angle $$\theta$$ for each pixel / patch, where the incident angle characterizes the angle between the imaging optical axis and the local normal of the foot. The terminal performs geometric compensation on the near-infrared multispectral image; the squared distance compensation can be expressed as: in, This is the original strength value. This is the intensity value after distance compensation. For reference distance.
[0054] Specifically, the terminal further performs Lambert cosine compensation to reduce the impact of irradiance differences caused by changes in the incident angle, and introduces a lower limit. To avoid compensation divergence at large angles, the compensation can be expressed as: in, This is the normalized intensity value.
[0055] Specifically, when When the depth confidence is too small or below the threshold, the terminal marks the corresponding pixel / area as a low-confidence region and performs weight reduction, interpolation repair, or prompts for reshooting in the subsequent analysis to avoid the propagation of compensation errors caused by abnormal edge incidence angles or depths; the criterion can be expressed as: Specifically, the terminal can map the geometrically compensated multispectral texture back to the standard two-dimensional plane coordinate system for subsequent anatomical partitioning, ROI statistics and cross-follow-up alignment; and can output low-confidence regions in the form of masks for subsequent feature extraction and risk inference stages to ignore or reduce weight.
[0056] Without limiting the scope of this invention, the surface geometry correction process can be used for near-infrared multispectral modes, and can also be extended to other imaging modes affected by distance and incident angle; the reference distance Angle threshold Confidence threshold and lower limit It can be preset or adaptively adjusted according to different terminal structures, imaging optical parameters and clinical acquisition environments to obtain more stable normalized results in different scenarios.
[0057] Please see Figure 6 , Figure 6 This illustrates the cloud-edge-device collaborative system structure and the uncertainty-driven closed-loop optimization process of the present invention. Figure 6 This is used to illustrate how, under the conditions of limited resources on handheld terminals and the need to balance privacy and bandwidth, the terminal can perform real-time screening and uncertainty assessment on the device side, and selectively upload difficult samples to the cloud when authorization and network conditions are met, thereby enabling continuous iterative updates of the model and threshold strategy.
[0058] Specifically, such as Figure 6 As shown, the system includes at least a handheld multimodal intelligent screening terminal (100), a cloud service module, and an application terminal module. The handheld multimodal intelligent screening terminal (100) is used to perform multimodal data acquisition, stability gating and Burst burst shooting, quality control and frame selection / fusion, 3D modeling and surface normalization, and edge-side lightweight inference, and outputs screening results and uncertainty indicators. The application terminal module can be a doctor's terminal or a patient's terminal, used to display screening results, risk classification, trend follow-up information, and alarm prompts, and can perform follow-up annotation or result confirmation within authorized scope.
[0059] Specifically, the handheld terminal (100) estimates the prediction uncertainty during the on-device inference process and classifies the samples according to the uncertainty threshold, quality control rules, and acquisition environment parameters. When a sample meets the high uncertainty or quality control triggering conditions, the terminal generates a difficult sample data packet and performs de-identification, desensitization, compression, and encryption processing. When the network is available and compliance requirements are met, the packet is selectively uploaded to the cloud service module to reduce bandwidth consumption and privacy risks. Preferably, the data packet may include multimodal raw data or feature data, acquisition geometric information, quality control markers, and necessary metadata; and only local area data related to the difficult determination may be uploaded to further reduce the data volume.
[0060] Specifically, after receiving the problematic sample data packets, the cloud service module can aggregate them into a sample pool and perform human-in-the-loop annotation and retraining. The cloud can optimize the model structure or threshold strategy based on the statistics of the problematic samples, and generate a lightweight model update package adapted to the terminal through distillation, pruning, and quantization. The cloud service module sends the model update package or threshold strategy update to the handheld terminal (100) through a secure channel to achieve continuous iteration of the terminal-side model and acquisition control logic.
[0061] Specifically, after receiving the update package from the cloud, the handheld terminal (100) performs integrity verification and version management, and supports online upgrades and rollbacks to the previous stable version to reduce the impact of update failures on clinical use. Preferably, the terminal can record operational indicators such as quality control failure statistics and data collection gate pass rates, and report the statistical information to the cloud when authorized, for adaptive optimization of thresholds and guidance strategies, thereby achieving more consistent handheld data collection results across different institutions and operators.
[0062] Preferably, the edge-side lightweight artificial intelligence model adopts a lightweight backbone network and multi-branch output structure adapted to mobile devices: the backbone network is used to extract shared representations of multimodal inputs, the fusion module is used to align and fuse near-infrared multispectral features, thermal imaging features and geometric / mass features, and the output end includes at least a risk grading output head and a lesion localization / thermal map output head, thereby improving the sensitivity to local abnormal areas while maintaining edge-side real-time performance.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications should be covered within the scope of the claims of the present invention.
Claims
1. A handheld integrated multimodal diabetic foot screening terminal, characterized in that, The terminal includes: a handheld housing (110) with an imaging window (111) at its front end; a near-infrared multispectral imaging unit (120), an infrared thermal imaging unit (130), and a ToF depth sensing unit (140), whose fields of view at least partially overlap; an IMU (145) for outputting terminal attitude-related angular velocity and acceleration information; a ring-shaped multi-band LED illumination module (150) for providing controllable pulse illumination; an embedded processing and storage unit (160) including a processor and a neural network acceleration unit for performing stability determination, synchronous triggering, continuous shooting frame selection / fusion, 3D modeling, surface normalization, feature extraction, end-side inference, and uncertainty assessment; a human-computer interaction module (170) for displaying shooting guidance and screening results and outputting audio-visual or vibration feedback; a wireless communication module (180) and a power management module (190) for secure communication with the cloud and power supply to each module; the embedded processing and storage unit (160) includes a synchronous triggering and a... The timestamp management module is used to uniformly label multispectral, thermal imaging, and depth data with a time reference, and triggers quality control failure when a time deviation exceeds a preset threshold. The ring-shaped multi-band LED lighting module (150) is equipped with a diffuse reflection uniform light structure and / or a polarization suppression structure, and the embedded processing and storage unit (160) adaptively adjusts the LED driving current and exposure / integration time according to the real-time distance measured by ToF to suppress near-distance overexposure and far-distance underexposure. The human-computer interaction module (170) displays a foot coverage heatmap or coverage indicator bar in scanning mode, and prohibits triggering acquisition and outputs prompts when incomplete coverage, angle deviation, or excessive jitter is detected. The wireless communication module (180) uploads difficult sample data packets only when preset network conditions are met. The difficult sample data packets undergo at least de-identification, desensitization, and encryption processing before uploading. The encryption processing can be selected from one or more of end-to-end symmetric encryption, asymmetric key negotiation, or homomorphic encryption.
2. The handheld integrated multimodal diabetic foot screening terminal according to claim 1, characterized in that, The terminal performs the following steps: S1: Terminal initialization and multimodal self-test: Perform power-on self-test on the near-infrared multispectral imaging unit, infrared thermal imaging unit, ToF depth sensing unit, IMU and ring multiband LED lighting module, and load the terminal inference model and acquisition control parameters; S2: Multidimensional Imaging Guidance and Stability Trigger Judgment: In preview mode, the ToF depth sensing unit acquires the distance field and local normal vector distribution of the sole surface in real time, while the IMU acquires the terminal angular velocity and acceleration information simultaneously; guidance information such as distance suitability, angle perpendicularity, and jitter status is displayed on the human-computer interaction interface; a joint stability criterion including distance deviation, incident angle deviation, and motion amplitude is constructed, and an acquisition trigger signal is generated if and only if the target area of the sole is within a preset distance range, the angle between the imaging optical axis and the local normal of the sole is less than a preset angle threshold, and the motion amplitude detected by the IMU is continuously lower than the stability threshold within a preset time window; S3: Multimodal short-time burst shooting and intelligent frame selection: In response to the acquisition trigger signal, the ring multi-band LED illumination module is controlled to flash illumination according to the preset spectral timing, and each imaging unit is synchronously triggered to perform short-time multi-frame burst shooting; the quality control of the burst shooting frame sequence is performed on the end side, and frames with motion blur, occlusion, specular reflection overexposure, saturation or field of view artifacts are removed, and the best frame with the highest quality score is selected or lightweight multi-frame fusion is performed to generate a multimodal raw dataset, and the acquisition environment parameters and timestamps are recorded; S4: 3D Modeling, Texture Mapping and Surface Normalization: Reconstruct a 3D mesh model of the foot based on the ToF depth map; use calibration parameters to map multispectral and thermal images as textures onto the 3D mesh surface; unfold and map the textured 3D model to a standard 2D plane coordinate system, and perform distance squared compensation and incident angle compensation based on geometric information to generate a normalized reflectance spectrum and temperature distribution map. At the same time, generate low-confidence masks for areas with excessive incident angles or low depth confidence, and use them for weight reduction or interpolation repair in subsequent analysis. S5: Anatomical zoning and feature extraction: Key anatomical regions such as heel, metatarsal heads, arch, and toes are automatically marked on the standard two-dimensional plane, and near-infrared multispectral tissue hemodynamic features and infrared thermographic features are extracted in each region. S6: End-side reasoning and uncertainty assessment: Integrating multispectral features, thermal features, and ToF geometric features, inputting a lightweight artificial intelligence model deployed on the terminal neural network acceleration unit, and outputting the risk level of diabetic foot, the location of suspected lesions, and uncertainty indices representing the confidence level of prediction.
3. A handheld integrated multimodal diabetic foot screening terminal according to claim 2, characterized in that, In step S2, the joint stability criterion employs a multi-level AND logic gating mechanism, including at least: Distance gating: This refers to the distance statistics of the target area measured by Time-of-Flight (ToF). For the preset target distance; when The time-distance gating is true; Angle gating: ,in The unit vector of the imaging optical axis. Let be the unit vector of the local normal vector of the foot obtained by fitting the ToF point cloud; when The time angle gating is true; Jitter gating: within a preset time window Internally, the angular velocity vectors of the IMU are respectively... With acceleration vector Calculate the time-averaged amplitude or root mean square amplitude when and The jitter gate is true; Only if the above gating conditions are all true and the duration exceeds The acquisition trigger signal is generated in time, and when any gate is false, the operator is prompted to adjust the distance, angle or remain still through display, sound and light or vibration feedback.
4. A handheld integrated multimodal diabetic foot screening terminal according to claim 3, characterized in that, In step S3, short burst shooting and intelligent frame selection include: N frames of multispectral images, thermal images, and depth maps were continuously acquired at a preset frame rate. ; Calculate a sharpness score for each frame. Inter-frame displacement ,in Obtained based on Laplace variance or gradient magnitude Obtained based on optical flow estimation or IMU time synchronization integration; Maximum inter-frame displacement ,when When the signal-to-noise ratio is not greater than a preset micro-motion threshold, registration and weighted fusion are performed on multiple frames to improve the signal-to-noise ratio; when Select when the value is greater than the preset micro-motion threshold The frame with the highest saturation ratio and the lowest saturation ratio is selected as the best frame.
5. A handheld integrated multimodal diabetic foot screening terminal according to claim 4, characterized in that, In step S3, the quality control includes at least determining and triggering a resampling prompt for any of the following: motion blur, missing foot area, occlusion, specular highlights, overexposure saturation, low depth confidence, or multimodal time synchronization anomaly.
6. A handheld integrated multimodal diabetic foot screening terminal according to claim 5, characterized in that, In step S4, surface normalization specifically includes: Constructing 3D points to two-dimensional parametric plane coordinates mapping And map the multimodal texture onto a two-dimensional plane; For multispectral band pixel values Obtained by distance square compensation and Lambert cosine compensation : , ,in This is the actual imaging distance. To calibrate the distance, For local incident angle, To prevent the lower limit of the denominator from being too small; when When the preset effective threshold is exceeded or the ToF depth confidence is lower than the threshold, the corresponding pixel is marked as a low-confidence region and its weight is reduced in subsequent analysis or repaired by interpolation based on neighborhood information.
7. A handheld integrated multimodal diabetic foot screening terminal according to claim 6, characterized in that, In step S6, the uncertainty index is obtained based on at least one uncertainty estimation strategy, including Monte Carlo Dropout, deep integration, or confidence estimation based on output entropy; and the terminal dynamically selects a low-power fast model or a high-precision enhancement strategy based on the stability score and the uncertainty index.
8. A multimodal cloud-edge collaborative screening system for diabetic foot, characterized in that: The system employs the screening terminal and execution steps described in any one of claims 1 to 7, and further includes a cloud service module and an application terminal module. The cloud service module is used to receive data packets of difficult samples and follow-up tags selectively uploaded by the terminal, perform model training, evaluation, and version management, and generate a lightweight model update package adapted to the terminal deployment. The application terminal module is used to allow medical personnel or subjects to view screening results, historical trends, and early warning information. The cloud service module sends the model update package to the handheld multimodal intelligent screening terminal, and the terminal performs integrity verification and supports rollback to the previous stable version. The system execution steps also include: S7: Uncertainty-driven cloud-edge collaborative closed loop: Based on the uncertainty index, quality control results and collection environment parameters, a graded screening is performed. When the uncertainty is higher than the preset threshold or a preset quality control rule is triggered, the difficult sample data packets that have been desensitized and encrypted are selectively uploaded to the cloud. After retraining and model compression are completed in the cloud, updated model parameters are sent to the terminal. The terminal can switch versions during idle periods and supports rollback.
9. A multimodal diabetic foot cloud-edge collaborative screening system according to claim 8, characterized in that: The cloud service module is also used to analyze the quality control failure logs uploaded by the terminal, and update the stability threshold, time window parameters or guidance strategy on the terminal side accordingly, forming a remote optimization of the acquisition control logic.
10. A multimodal diabetic foot cloud-edge collaborative screening system according to claim 8, characterized in that: The lightweight model update package is obtained through teacher-student knowledge distillation and further pruned and quantized to adapt to the neural network acceleration unit architecture of the terminal.