Non-contact blood glucose monitoring device and method

By using a non-contact blood glucose monitoring device with a near-infrared camera and lightweight modeling algorithms, high-precision and low-power blood glucose monitoring is achieved, solving the problems of large size and unstable accuracy of existing devices. It is suitable for home and telemedicine.

CN121040903APending Publication Date: 2025-12-02SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511252967.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing contact-type near-infrared blood glucose monitoring devices are bulky, expensive, and have unstable accuracy in remote and continuous monitoring, making it difficult to meet the needs of medical-grade applications.

Method used

A non-contact blood glucose monitoring device is used to capture facial video through a near-infrared camera. A lightweight modeling algorithm is used to extract key facial feature points, perform dynamic monitoring area selection and global modeling, and combine the physiological guidance time module to estimate blood glucose, so as to achieve end-to-end real-time monitoring.

Benefits of technology

It improves the accuracy and stability of blood glucose monitoring, is suitable for high-frequency, long-term continuous monitoring, and reduces the size and power consumption of the device, making it suitable for home and telemedicine applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121040903A_ABST
    Figure CN121040903A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of non-contact blood glucose monitoring, and provides a non-contact blood glucose monitoring device and method.The device comprises the steps that infrared video frames of the face of a user are obtained; compressing the original infrared video frame to obtain a compressed video frame; identifying an original infrared video frame, extracting key facial feature points, and dynamically selecting an optimal monitoring area based on facial physiological feature stability; continuous face information is extracted from the selected optimal monitoring area, remodeling is carried out, parallel large kernel and small kernel convolution paths are utilized, multi-scale context representation is generated, global modeling is carried out to capture a time dependency relationship, features are integrated, and the current blood glucose concentration of the user is obtained through estimation; face detection is carried out based on the compressed video frame, corresponding user information and blood glucose concentration historical records are called according to a detection result, and blood glucose fluctuation is determined in combination with the current blood glucose concentration. The accuracy of non-contact blood glucose monitoring is improved, and the detection precision is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of non-contact blood glucose monitoring technology, specifically relating to a non-contact blood glucose monitoring device and method. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Monitoring blood glucose levels is crucial for disease prevention, diagnosis, and treatment. Currently, most clinically used blood glucose monitoring methods are invasive, such as finger-prick blood sampling and venous blood sampling. While these methods offer a degree of accuracy, the procedure requires skin puncture, leading to pain, bleeding, and infection risks, resulting in poor patient compliance and hindering the requirements for high-frequency, long-term continuous blood glucose monitoring. Therefore, developing convenient, comfortable, and non-invasive blood glucose monitoring methods has become a research hotspot.

[0004] Among numerous non-contact detection technologies, near-infrared (NIR) methods have been widely studied for the quantitative analysis of glucose concentration in tissues due to their excellent penetration and specific recognition capabilities for tissue components. This method involves irradiating the skin surface with near-infrared light within a specific wavelength range, and then analyzing the spectral signals absorbed or scattered by the tissue to reflect blood glucose levels to a certain extent. Compared to invasive blood glucose detection methods, NIR offers advantages such as being non-invasive, rapid, and real-time, and has broad application prospects.

[0005] However, current near-infrared blood glucose monitoring devices are mostly contact-based, requiring the probe to be placed close to the skin or fixedly positioned on the site of detection, limiting their application in remote, continuous monitoring scenarios. Furthermore, some existing methods rely on bulky spectrometers or high-power lasers, resulting in large systems, high costs, and high power consumption, making them unsuitable for wearable, portable, and home-based deployments. Moreover, spectral signals are easily affected by ambient light interference, differences in skin condition, and hand movements, leading to unstable detection accuracy and failing to meet the requirements of medical-grade applications. Summary of the Invention

[0006] To address the aforementioned problems, this invention proposes a non-contact blood glucose monitoring device and method. This invention simultaneously acquires the user's facial near-infrared video stream, identifies key facial feature points by analyzing the near-infrared video stream, dynamically selects the optimal monitoring region based on the stability of facial physiological features, extracts continuous facial information from the selected ROI region, and performs deep fusion analysis using a lightweight modeling algorithm. Ultimately, this achieves continuous monitoring of the user's current blood glucose concentration, improving the accuracy of non-contact blood glucose monitoring and ensuring detection precision.

[0007] According to some embodiments, the present invention adopts the following technical solution: A non-contact blood glucose monitoring device, comprising: The near-infrared supplementary lighting module includes a near-infrared light source array and an image sensor. The near-infrared light source array is used to provide multi-band near-infrared light to illuminate the user's face, and the image sensor is used to acquire infrared video frames of the user's face. The video processing module is used to compress the original infrared video frames to obtain compressed video frames; The video feature extraction unit is used to identify the original infrared video frames, extract key facial feature points, and dynamically select the optimal monitoring area based on the stability of facial physiological features. The blood glucose estimation unit is used to extract continuous facial information from the selected optimal monitoring area, reshape it, generate multi-scale contextual representation using parallel large and small kernel convolutional paths, perform global modeling to capture time dependencies, integrate features, and estimate the user's current blood glucose concentration. The interaction unit is used to perform face detection based on compressed video frames, and retrieve the corresponding user information and blood glucose concentration history based on the detection results, and determine blood glucose fluctuations by combining the current blood glucose concentration.

[0008] As an alternative implementation, the device further includes a visible light cutoff filter component and a steady-state drive circuit.

[0009] As an alternative implementation, the blood glucose estimation unit utilizes a lightweight blood glucose monitoring network to achieve blood glucose estimation, including a multi-scale convolution module, a physiological guidance time module, and a fusion and prediction module. The multi-scale convolution module is used to extract continuous facial information from the selected optimal monitoring region, reconstruct it, and generate a multi-scale context representation using parallel large-kernel and small-kernel convolution paths. The physiological guidance time module is used to perform global modeling to capture temporal dependencies. The fusion and prediction module is used to integrate features and estimate the user's current blood glucose concentration.

[0010] As a further implementation, the input of the multi-scale convolutional module is the video data of the selected optimal monitoring area, which is reshaped to process the spatial features of all time frames simultaneously. Parameters are shared between time frames. Parallel large-kernel and small-kernel convolutional paths are used. The multi-scale large-kernel convolutional path is used to perceive contextual relationships, while the small-kernel convolutional path aggregates features from relevant local contexts, and then the multi-scale features are fused.

[0011] As a further implementation, the large-kernel convolutional path employs large-kernel depthwise convolution to capture extensive contextual relationships. This is achieved by applying a single filter to each input channel for the token. The operation is as follows: ; in, Indicated by Token The size of the center is The neighborhood, For large kernel depthwise convolution, Point convolution is used to model spatial relationships and outputs... To generate context-adaptive weights.

[0012] As a further implementation, the small kernel convolution path uses small kernel dynamic convolution to aggregate features in highly relevant local contexts. The input channels are divided into multiple groups, each group sharing aggregation weights. For each channel... Belongs to group Token The polymerization process is as follows: ; in, These are reshaped adaptive weights generated from the large kernel convolution path. This indicates a convolution operation.

[0013] As an alternative implementation, the outputs of the two parallel paths of the multi-scale convolutional module are fused through point convolution to integrate multi-scale features, and then a squeeze-excitation layer is added to enhance feature representation through a channel attention mechanism.

[0014] As an optional implementation, the physiological guidance time module includes a physiological embedding layer and an adaptive time grouping attention layer. The physiological embedding layer injects physiological prior information into the input tensor, generates sine / cosine position codes for each time point, captures the sequence order, calculates the local trend of the input sequence through first-order difference, maps it to a set dimension through a fully connected layer, and finally embeds it as the sum of the three. The trend embedding tensor is expanded through point convolution and fused with the input tensor. The adaptive temporal grouping attention layer is used to perform global average pooling on the temporal dimension of the input tensor, calculate the fluctuation intensity of the time series, generate group number weights through a fully connected layer based on the fluctuation intensity, dynamically adjust the number of groups, group the channels according to the number of groups, and apply multi-head self-attention to each group.

[0015] As an alternative implementation, the fusion and prediction module flattens the output of the physiological guidance time module in the spatial dimension and uses a prediction head to predict blood glucose levels. The prediction head contains two linear layers with a ReLU activation function and a dropout rate used in between to prevent overfitting. The final output is the predicted blood glucose level value.

[0016] A non-contact blood glucose monitoring method includes the following steps: Provides multi-band near-infrared light to illuminate the user's face and acquires infrared video frames of the user's face; The original infrared video frames are compressed to obtain compressed video frames. The original infrared video frames are identified, key facial feature points are extracted, and the optimal monitoring area is dynamically selected based on the stability of facial physiological features. Continuous facial information is extracted from the selected optimal monitoring area, reconstructed, and multi-scale contextual representations are generated using parallel large-kernel and small-kernel convolutional paths. Global modeling based on the physiological guidance time module is then performed to capture time dependencies, integrate features, and estimate the user's current blood glucose concentration. Face detection is performed based on compressed video frames. The corresponding user information and blood glucose concentration history are retrieved based on the detection results, and the blood glucose fluctuation is determined in combination with the current blood glucose concentration.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention uses a near-infrared camera to capture facial video for blood glucose monitoring, eliminating the need for skin contact, thus removing the pain and infection risks of traditional invasive methods. It significantly improves patient comfort and monitoring compliance, and is particularly suitable for high-frequency, long-term continuous monitoring needs.

[0018] This invention integrates a processing model that achieves end-to-end acceleration from signal acquisition to result output, helping to control inference latency, ensuring real-time blood glucose monitoring, and is suitable for dynamic health management scenarios. Through ingenious model design, it maintains low computational complexity while ensuring powerful feature extraction and global time modeling capabilities, which helps improve detection accuracy and guarantee the accuracy of results.

[0019] This invention employs a compact near-infrared camera and an embedded platform, which is small in size, low in power consumption, and low in cost, making it suitable for home monitoring and telemedicine applications and promoting the popularization of the technology.

[0020] This invention effectively mitigates interference from ambient light, skin condition differences, and motion artifacts by using optimized signal processing algorithms and physiologically guided modeling, thereby improving the stability and accuracy of non-contact blood glucose monitoring and meeting the needs of medical-grade applications.

[0021] This invention enhances the practicality and scalability of non-contact blood glucose monitoring, providing an efficient and convenient solution for chronic disease management, telemedicine, and daily health monitoring. Its non-contact design, combined with high real-time performance, portability, and high accuracy, significantly improves the universality and user experience of blood glucose monitoring, offering diabetic patients a more flexible and comfortable way to manage their health, while also laying a foundation for the future development of health monitoring technologies.

[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0023] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0024] Figure 1 This is a schematic diagram of the lens module of a non-contact blood glucose monitoring device according to one embodiment, wherein (a) is a schematic diagram of the distribution of the lens and the supplementary light module; and (b) is a schematic diagram of the distribution of the modules inside the lens module. Figure 2 This is a monitoring flowchart of one embodiment of the device; Figure 3 This is a schematic diagram of a non-contact blood glucose monitoring model according to one embodiment; Figure 4 This is a schematic diagram of a multi-scale convolutional module in a non-contact blood glucose monitoring model according to one embodiment; Figure 5 This is a schematic diagram of the adaptive time-grouping attention module in a non-contact blood glucose monitoring model according to one embodiment; Figure 6 This is a schematic diagram illustrating the hardware and software co-optimization of a non-contact blood glucose monitoring device according to one embodiment. Detailed Implementation

[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0026] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0027] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0028] Where there is no conflict, the embodiments and features described in this application may be combined with each other.

[0029] Example 1 A non-contact blood glucose monitoring device, such as Figure 2As shown, it includes: The near-infrared supplementary lighting module includes a near-infrared light source array and an image sensor. The near-infrared light source array is used to provide multi-band near-infrared light to illuminate the user's face, and the image sensor is used to acquire infrared video frames of the user's face. The video processing module is used to compress the original infrared video frames to obtain compressed video frames; The video feature extraction unit is used to identify the original infrared video frames, extract key facial feature points, and dynamically select the optimal monitoring area based on the stability of facial physiological features. The blood glucose estimation unit is used to extract continuous facial information from the selected optimal monitoring area, reshape it, generate multi-scale contextual representation using parallel large and small kernel convolutional paths, perform global modeling to capture time dependencies, integrate features, and estimate the user's current blood glucose concentration. The interaction unit is used to perform face detection based on compressed video frames, and retrieve the corresponding user information and blood glucose concentration history based on the detection results, and determine blood glucose fluctuations by combining the current blood glucose concentration.

[0030] In this embodiment, as Figure 1 As shown in (a) and (b), the near-infrared supplementary lighting module (hereinafter referred to as the supplementary lighting module) of the device includes a near-infrared light source array, a CMOS image sensor, a visible light cutoff filter assembly, a lens group, and a steady-state driving circuit. The near-infrared light source array is configured with a multi-band near-infrared LED array (typical center wavelength range: 850–940nm in this embodiment), and brightness control and modulation frequency synchronization are achieved through PWM to improve signal modulation depth and dynamic response capability. A visible light cutoff filter assembly is placed in front of the lens group to selectively block visible light, allowing only near-infrared or other specific wavelengths to pass through, thereby reducing ambient visible light interference. Simultaneously, to avoid diffuse reflection interference and ambient stray light pollution, the near-infrared light source array is equipped with a focusing lens and a light shield with a set angle to ensure uniform light intensity and good directionality of the light irradiated to the face.

[0031] The near-infrared supplementary lighting module is controlled by the main control board.

[0032] This embodiment employs a CMOS image sensor designed for near-infrared enhancement, which significantly outperforms ordinary visible light sensors in the 850–940 nm wavelength band. It supports high frame rate (≥60fps) sampling and features a global shutter function, effectively avoiding motion artifacts from interfering with the quality of near-infrared image acquisition. The image resolution of the CMOS image sensor can be adaptively adjusted according to scene requirements, with typical configurations being 1280×720 or 1920×1080.

[0033] In addition, the lens assembly uses a fixed-focus near-infrared lens (typical focal length: 6mm~8mm), which has wide-angle characteristics to ensure clear imaging at different face distances. The lens assembly is connected to the main control board via MIPI CSI-2 to ensure real-time, low-latency image data upload.

[0034] The steady-state drive circuit and temperature control circuit of the supplementary lighting module are key components to ensure the imaging quality and system stability of near-infrared light.

[0035] In this embodiment, the steady-state drive circuit mainly adopts a constant voltage and constant current source control strategy, integrating a high-precision linear regulator and a current limiting module to ensure the brightness consistency and spectral stability of the near-infrared light source during long-term operation, effectively suppressing signal drift and image noise caused by power supply fluctuations. Simultaneously, the temperature control circuit integrates a thermistor and a high-precision digital-to-analog converter sampling system to monitor the lens module's operating temperature in real time, and uses a proportional-integral-differential algorithm to drive the stepless fan for active temperature control.

[0036] In this example, the system utilizes a deep learning-based MediaPipe Face Mesh algorithm to perform high-precision facial geometry analysis on the original video frames, obtaining the locations of key facial landmarks. The system leverages the geometric constraints of these landmarks and facial physiological feature stability assessments, combined with a dynamic optimization algorithm based on signal quality metrics (peak signal-to-noise ratio, light intensity variance), to select the optimal monitoring region in real time. Within the candidate regions, the system first calculates the peak signal-to-noise ratio, measuring the ratio of blood flow pulse signal to background noise; simultaneously, it calculates the light intensity variance to assess local illumination changes and stability. Then, head pose constraints are used to eliminate side-face regions, further filtering unstable areas. Finally, the system performs temporal filtering on the selected ROI locations in the time domain and employs a weighted fusion strategy to smooth ROI coordinate changes between consecutive frames, achieving steady-state tracking of region locations and enhancing video quality.

[0037] In this embodiment, as Figure 3As shown, the blood glucose estimation unit utilizes a lightweight blood glucose monitoring network to perform blood glucose estimation. This network consists of three main parts: (1) a multi-scale large-small convolutional block (MSLSConv) for extracting spatial features from the input tensor; (2) a lightweight physiologically-guided time transformer block (LPGTT) for global modeling to capture temporal dependencies; and (3) a fusion and prediction block (FPB) for integrating features and generating blood glucose predictions. This architecture is designed to maintain low computational complexity while ensuring powerful feature extraction and global temporal modeling capabilities to support real-time inference on edge devices.

[0038] Among them, such as Figure 4 As shown, the multi-scale convolutional module MSLSConv takes video data of shape [B, T, C, H, W] as input and reshapes it to [B×T, C, H, W] to simultaneously process spatial features across all time frames, with parameters shared between frames. Furthermore, to achieve multi-scale feature extraction, the module employs parallel large-kernel and small-kernel convolutional paths, combining a "large-kernel perception, small-kernel aggregation" approach with a dynamic convolution mechanism to simulate the broad perception and precise focusing capabilities of the human visual system.

[0039] Large kernel perception stage: Employs large kernel depthwise convolution (e.g., kernel size...). This allows for the capture of broad contextual relationships. Depthwise convolutions significantly reduce computational complexity while achieving a large receptive field by applying a single filter to each input channel. For tokens... The specific operation can be described as follows: (1) in, Indicated by Token The size of the center is The neighborhood, For large kernel depthwise convolution, This is a point convolution used to model spatial relationships. Output Context-adaptive weights are generated for subsequent aggregation steps. This design ensures that the module can efficiently capture global spatial information.

[0040] Small kernel aggregation stage: Employs small kernel dynamic convolution (e.g., kernel size...). Features are aggregated within highly relevant local contexts. To reduce memory overhead, the input channels are divided into... Groups, each sharing the aggregate weight. For channels... Belongs to group Token The polymerization process is as follows: (2) in, These are reshaped adaptive weights generated from LKP. This represents a convolution operation. This dynamic aggregation ensures accurate fusion of local features, capturing subtle spatial patterns associated with blood glucose monitoring.

[0041] Multi-scale integration: To achieve multi-scale feature extraction, the MSLSConv module applies large convolutional kernels of different kernel sizes through parallel paths (e.g., and The MSLSConv module generates multi-scale contextual representations. The outputs of the two paths are fused through pointwise convolutions to integrate multi-scale features. Subsequently, a squeeze-activation layer is added to enhance feature representation through a channel attention mechanism. The MSLSConv module maintains a lightweight design through depthwise convolutions and dynamic weights, making it suitable for real-time applications.

[0042] The physiological guidance time module is described below: The physiological guidance time module consists of two core components: (1) Physiological Embedding Layer (PE Layer); and (2) Adaptive Temporal Grouping Attention (ATGA).

[0043] Physiological Embedding Layer: Blood glucose time series typically exhibit periodic fluctuations and local trends. To model these physiological characteristics, the input tensor [B, T, C', H', W'] is first injected with prior physiological information through a physiological embedding layer. First, for each time point... Generate sine / cosine position codes and capture sequence order using the following formula: (3) (4) in, For time point indexing, For embedding dimension indexes, For the embedded dimension (set as) ).

[0044] The local trend of the input sequence is calculated using first-order differencing (e.g., The input tensor is mapped to c' / 4 dimensions via a fully connected layer. The final embedding is the sum of the three, with a shape of [B, T, C' / 4, H', W']. This is then expanded to a trend embedding tensor TE of [B, T, C', H', W'] via point convolutions and fused with the input tensor. (5) Here, X is the input tensor, and PE and TE are the position and trend embeddings, respectively. This embedding mechanism enhances the module's sensitivity to the physiological characteristics of blood glucose sequences.

[0045] Adaptive temporal grouping attention: Traditional grouping attention divides channels equally, ignoring the dynamic characteristics of time series. Adaptive temporal grouping attention, such as... Figure 5 As shown, the number of groups is dynamically allocated based on the fluctuation intensity of the blood glucose sequence, enhancing the modeling capability for key time points. Therefore, global average pooling (GAP) is performed on the time dimension of the input tensor [B, T, C', H', W'] to obtain [B, T, C']. The fluctuation intensity of the time series is then calculated: (6) in, The fluctuation intensity at time point t. Based on The number of group weights is generated through the fully connected layer. (Range [4, C' / 4], e.g., groups 4 to c' / 4), dynamically adjust the number of groups: (7) according to Divide the channel into Groups, each group has a dimension of Multi-head self-attention (MHSA) is applied to each group: (8) in, , Dynamic grouping allows modules to allocate more groups in highly volatile regions to capture fine-grained information, and reduce the number of groups in stable regions to reduce complexity.

[0046] Fusion and Predictive Head The output of the FPB module is flattened in the spatial dimension and used for glucose level prediction via a lightweight feedforward network. The prediction head contains two linear layers with a ReLU activation function and a dropout rate (p=0.2) in between to prevent overfitting. The final output is a scalar glucose level prediction for each batch of samples.

[0047] This example uses video compression technology to compress the input video, reducing data volume and improving transmission and computation efficiency. Then, it utilizes a face detection and recognition algorithm based on MediaPipe Face Detection to quickly locate and identify the user from the compressed video frames. By combining the identified user information with their historical blood glucose concentration, the currently detected blood glucose concentration is compared and analyzed to assess blood glucose fluctuation trends in real time.

[0048] In addition, in this embodiment, as Figure 6 As shown, hardware and software co-optimization is employed.

[0049] The core objective is to capture and analyze a patient's blood glucose status in real time through facial video-based dynamic physiological signal sensing technology, thereby ensuring efficient and accurate blood glucose monitoring and control.

[0050] To achieve high-precision monitoring in a short time, the workflow in this embodiment is divided into three key stages: video acquisition and preprocessing, blood glucose calculation, and GUI display. Considering the mixed parallel execution of deep learning (DL) tasks (such as blood glucose calculation) and non-deep learning tasks (such as video acquisition and GUI display) in the system, and that DL tasks typically consume a lot of hardware resources and time, it is particularly important to rationally plan the system workflow.

[0051] Therefore, in the specific implementation process, this embodiment uses the NVIDIA Jetson embedded module as the system hardware platform and combines heterogeneous computing technology to reasonably schedule different computing units (such as CPU, GPU and DLA) to make full use of system hardware resources and improve computing efficiency.

[0052] For the computationally intensive and time-consuming deep learning (DL) tasks, this embodiment utilizes the Deep Learning Accelerator (DLA) in the NVIDIA Jetson embedded module to offload the workload of the DL tasks. Furthermore, this embodiment uses the two DLAs equipped in the NVIDIA Jetson embedded module to construct a two-stage pipeline, running physiological signal perception and intelligent decision-making inference tasks respectively, further reducing system latency and ensuring real-time decision-making.

[0053] For the computationally intensive video acquisition and preprocessing parts of the non-DL tasks, this embodiment deploys them on the GPU of the embedded system and utilizes CUDA acceleration and zero-memory copy technology to further improve the computational speed of video acquisition and preprocessing. Since the computational requirements of the control output part in the non-DL tasks are relatively small, this embodiment deploys it on the CPU. By rationally allocating computing resources according to different task requirements, this embodiment constructs a physiological signal-guided perception-computation-display closed-loop blood glucose monitoring system, providing a flexible and standardized interface for future continuous iterative updates of the system.

[0054] Furthermore, this embodiment optimizes the deep learning model of the blood glucose monitoring system to reduce model size and inference time, adapting to the resource limitations of the NVIDIA Jetson embedded module. Lightweight measures such as network pruning and inference optimization are implemented. First, the importance of each network block in the model is evaluated, and the performance after pruning is tested using a small number of blood glucose samples to measure performance recovery and the degree of inference time reduction. Combining these two aspects, an importance score is calculated for each network block, prioritizing the pruning of blocks with minimal impact on performance and significant acceleration effects until the model size meets the target. Subsequently, the pruned model is fine-tuned using a small number of samples, and a lightweight adapter is added to restore accuracy.

[0055] During deployment, the pruned model is run in a two-stage pipeline on two DLAs of the NVIDIA Jetson embedded module, further accelerating inference time. Through heterogeneous computing and pruning optimization, the accuracy is restored to over 95% of the original while reducing overall system latency, making it suitable for real-time blood glucose monitoring.

[0056] Finally, the lightweight blood glucose monitoring network proposed in this embodiment further accelerates and optimizes the model through TensorRT. By utilizing its graph fusion, accuracy calibration and layer fusion mechanisms, the network parameters are effectively compressed and the inference speed is improved. Through techniques such as few-sample pruning and heterogeneous computation optimization, the model accuracy and inference speed are balanced, providing an efficient and lightweight solution for wearable blood glucose monitoring devices, while also supporting future system iterations.

[0057] Example 2 A non-contact blood glucose monitoring method based on the device provided in the embodiments includes the following steps: Provides multi-band near-infrared light to illuminate the user's face and acquires infrared video frames of the user's face; The original infrared video frames are compressed to obtain compressed video frames. The original infrared video frames are identified, key facial feature points are extracted, and the optimal monitoring area is dynamically selected based on the stability of facial physiological features. Continuous facial information is extracted from the selected optimal monitoring area, reconstructed, and multi-scale contextual representation is generated using parallel large-kernel and small-kernel convolutional paths. Global modeling is then performed to capture temporal dependencies, features are integrated, and the user's current blood glucose concentration is estimated. Face detection is performed based on compressed video frames. The corresponding user information and blood glucose concentration history are retrieved based on the detection results, and the blood glucose fluctuation is determined in combination with the current blood glucose concentration.

[0058] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of one or more computer-usable storage media (including, but not limited to, disk storage, etc.) containing computer-usable program code. CD - ROM It takes the form of a computer program product implemented on (such as optical memory, etc.).

[0059] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0060] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0061] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0062] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made by those skilled in the art without creative effort within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A non-contact blood glucose monitoring device, characterized in that, include: The near-infrared supplementary lighting module includes a near-infrared light source array and an image sensor. The near-infrared light source array is used to provide multi-band near-infrared light to illuminate the user's face, and the image sensor is used to acquire infrared video frames of the user's face. The video processing module is used to compress the original infrared video frames to obtain compressed video frames; The video feature extraction unit is used to identify the original infrared video frames, extract key facial feature points, and dynamically select the optimal monitoring area based on the stability of facial physiological features. The blood glucose estimation unit is used to extract continuous facial information from the selected optimal monitoring area, reshape it, generate multi-scale contextual representation using parallel large and small kernel convolutional paths, perform global modeling to capture time dependencies, integrate features, and estimate the user's current blood glucose concentration. The interaction unit is used to perform face detection based on compressed video frames, and retrieve the corresponding user information and blood glucose concentration history based on the detection results, and determine blood glucose fluctuations by combining the current blood glucose concentration.

2. The non-contact blood glucose monitoring device as described in claim 1, characterized in that, The device also includes a visible light cutoff filter assembly and a steady-state drive circuit.

3. The non-contact blood glucose monitoring device as described in claim 1, characterized in that, The blood glucose estimation unit utilizes a lightweight blood glucose monitoring network to estimate blood glucose levels. It includes a multi-scale convolution module, a physiological guidance time module, and a fusion and prediction module. The multi-scale convolution module extracts continuous facial information from the selected optimal monitoring region, reconstructs it, and generates a multi-scale context representation using parallel large and small kernel convolution paths. The physiological guidance time module performs global modeling to capture temporal dependencies. The fusion and prediction module integrates features to estimate the user's current blood glucose concentration.

4. A non-contact blood glucose monitoring device as described in claim 3, characterized in that, The input of the multi-scale convolutional module is the video data of the selected optimal monitoring area, which is reshaped to process the spatial features of all time frames simultaneously. Parameters are shared between time frames. Parallel large-kernel and small-kernel convolutional paths are used. The multi-scale large-kernel convolutional path is used to perceive the context relationship, while the small-kernel convolutional path aggregates features from the relevant local context, and then the multi-scale features are fused.

5. A non-contact blood glucose monitoring device as described in claim 4, characterized in that, The large kernel convolution path uses large kernel depthwise convolution to capture a wide range of contextual relationships, for Token The operation is as follows: ; in, Indicated by Token The size of the center is The neighborhood, For large kernel depthwise convolution, Point convolution is used to model spatial relationships and outputs... To generate context-adaptive weights.

6. A non-contact blood glucose monitoring device as described in claim 4, characterized in that, The kernel-small convolution path uses dynamic kernel-small convolution to aggregate features in highly relevant local contexts. The input channels are divided into multiple groups, each sharing aggregation weights. For each channel... Belongs to group Token The polymerization process is as follows: ; in, These are reshaped adaptive weights generated from the large kernel convolution path. This indicates a convolution operation.

7. A non-contact blood glucose monitoring device as described in claim 1, characterized in that, The outputs of the two parallel paths of the multi-scale convolution module are fused through point convolution to integrate multi-scale features, and then a squeeze-excitation layer is added to enhance feature representation through a channel attention mechanism.

8. A non-contact blood glucose monitoring device as described in claim 1, characterized in that, The physiological guidance timing module includes a physiological embedding layer and an adaptive time grouping attention layer. The physiological embedding layer injects physiological prior information into the input tensor, generates sine / cosine position codes for each time point, captures the sequence order, calculates the local trend of the input sequence through first-order difference, maps it to a set dimension through a fully connected layer, and finally embeds it as the sum of the three. The trend embedding tensor is expanded through point convolution and fused with the input tensor. The adaptive temporal grouping attention layer is used to perform global average pooling on the temporal dimension of the input tensor, calculate the fluctuation intensity of the time series, generate group number weights through a fully connected layer based on the fluctuation intensity, dynamically adjust the number of groups, group the channels according to the number of groups, and apply multi-head self-attention to each group.

9. A non-contact blood glucose monitoring device as described in claim 1, characterized in that, The fusion and prediction module flattens the output of the physiological guidance time module in the spatial dimension and uses a prediction head to predict blood glucose levels. The prediction head contains two linear layers with a ReLU activation function and a dropout rate in between to prevent overfitting. The final output is the predicted blood glucose level.

10. A non-contact blood glucose monitoring method, characterized in that, Includes the following steps: Provides multi-band near-infrared light to illuminate the user's face and acquires infrared video frames of the user's face; The original infrared video frames are compressed to obtain compressed video frames. The original infrared video frames are identified, key facial feature points are extracted, and the optimal monitoring area is dynamically selected based on the stability of facial physiological features. Continuous facial information is extracted from the selected optimal monitoring area, reconstructed, and multi-scale contextual representation is generated using parallel large-kernel and small-kernel convolutional paths. Global modeling is then performed to capture temporal dependencies, features are integrated, and the user's current blood glucose concentration is estimated. Face detection is performed based on compressed video frames. The corresponding user information and blood glucose concentration history are retrieved based on the detection results, and the blood glucose fluctuation is determined in combination with the current blood glucose concentration.