An intelligent stethoscope based on multi-modal data collaborative diagnosis
Patent Information
- Application Number
- CN202521022278.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2035-05-23
AI Technical Summary
[0002]现有的智能听诊器设备仅聚焦心音或呼吸音分析,缺乏与血氧、心电图等生理参数的协同诊断,因为光吸收偏差导致血氧饱和度检测误差高,未解决运动伪影干扰问题,导致疾病筛查准确率受限,精度不高
[0012] (1) This utility model connects the acoustic sensor component and the physiological parameter sensor component to the stethoscope, processes the signal analysis and multimodal fusion unit and calls the existing pre-trained AI model to realize the fusion of acoustic and physiological features, which solves the problem that traditional stethoscopes lack physiological parameter references, rely on a single basis and have low accuracy.
Smart Images

Figure CN224792354U_ABST
Abstract
Description
Technical Field
[0001] This utility model relates to an intelligent stethoscope based on multimodal data collaborative diagnosis, belonging to the field of intelligent devices and medical electronic equipment technology. Background Technology
[0002] Existing intelligent stethoscope devices only focus on analyzing heart sounds or breath sounds, lacking collaborative diagnosis with physiological parameters such as blood oxygenation and electrocardiograms. High errors in blood oxygen saturation detection are caused by light absorption bias, and motion artifact interference remains unresolved, limiting the accuracy and precision of disease screening. Furthermore, most devices rely on cloud computing, making real-time processing impossible in offline environments, and computational latency further impacts diagnostic efficiency. Utility Model Content
[0003] The purpose of this invention is to overcome the above-mentioned shortcomings and provide an intelligent stethoscope based on multimodal data collaborative diagnosis.
[0004] The technical solution adopted by this utility model is as follows:
[0005] An intelligent stethoscope based on multimodal data collaborative diagnosis includes an acoustic sensor assembly, a physiological parameter sensor assembly, a signal analysis and multimodal fusion unit, and a human-computer interaction and communication unit. The signal analysis and multimodal fusion unit includes a signal preprocessing unit, a heterogeneous computing core unit, and an edge computing unit connected in sequence. The outputs of the acoustic sensor assembly and the physiological parameter sensor assembly are connected to the input of the signal preprocessing unit of the signal analysis and multimodal fusion unit. The heterogeneous computing core unit includes ARM and FPGA modules, which are connected to the edge computing unit via PCIe interface and SPI bus, respectively. The edge computing unit adopts the NVIDIA Jetson embedded system, which directly calls the existing pre-trained AI model to achieve the fusion of acoustic and physiological features. The AI model directly outputs the diagnostic results. The output of the edge computing unit is connected to the input of the human-computer interaction and communication unit.
[0006] The acoustic sensor assembly includes a disposable food-grade silicone mouthpiece, which is connected to a directional acoustic waveguide via a one-way airflow valve and a magnetic interface. A microphone array is fixed to the outlet end of the directional acoustic waveguide to collect cough sounds using beamforming technology. A piezoelectric sensor is located on the outer wall of the directional acoustic waveguide to collect vibration signals from the waveguide. The directional acoustic waveguide is made of stainless steel bellows with spiral guide grooves on its inner wall to optimize the transmission efficiency of low-frequency cough sounds (100-500Hz) and reduce signal attenuation. The microphone array consists of four microphones arranged in a diamond shape and fixed to the outlet end of the acoustic waveguide. The piezoelectric sensor is attached to the outer wall of the acoustic waveguide and fixed with epoxy resin adhesive.
[0007] The physiological parameter sensor assembly is rectangular in shape, with a data acquisition inlet on the front, a dual-wavelength transmission sensor on one side, and a microlens array on the other side. The microlens array employs a gradient refractive index design, consisting of lenses with micron-level apertures and relief depths, effectively suppressing motion artifacts. The dual-wavelength transmission sensor, with its 45° tilted dual-light source layout and narrow-band filtering, enhances blood oxygen detection accuracy. The internal contact surface of the data acquisition inlet uses a flexible silicone substrate with a nano-hydrophobic coating to conform to the patient's finger. Adaptive deformation enhances skin adhesion stability and suppresses signal attenuation caused by sweat, making it particularly suitable for patients with darker skin tones. The patient acquires parameters by inserting their finger into the data acquisition inlet, ensuring it conforms to the internal contact surface. The dual-wavelength transmission sensor first transmits red and infrared light through the patient's fingertip. The microlens array uses a gradient refractive index design to compensate for optical path shifts caused by finger movement, suppressing motion artifacts. The transmitted light is focused by the filter in the sensor and received by a photodiode. The received signal is then transmitted to the signal preprocessing unit. The physiological parameter sensor assembly enhances signal coupling through a flexible bonding structure, improves blood oxygenation accuracy through dual-wavelength optical detection, and suppresses motion interference through gradient refractive index microlenses, achieving high-precision and highly interference-resistant blood oxygen saturation detection that meets the medical-grade requirements for dynamic monitoring scenarios.
[0008] The signal preprocessing unit includes a low-noise amplifier, a bandpass filter, a dynamic baseline correction circuit, and an analog-to-digital converter (ADC) to perform preliminary noise reduction, characteristic frequency band extraction, and high-precision digitization of acoustic and physiological signals. After passing through the amplifier and bandpass filter in the preprocessing module, the acoustic signal undergoes characteristic frequency band extraction and high-precision digitization via the ADC. The physiological signal passes through the dynamic baseline correction circuit using an analog high-pass filter to eliminate low-frequency motion artifacts and respiratory baseline drift, and is then quantized via the ADC.
[0009] The FPGA module integrates an MFCC IP core, and the NVIDIA Jetson edge computing unit features a TensorCore hardware accelerator. The ARM and FPGA are co-designed to achieve hardware acceleration and reduce latency. The MFCC IP core directly calls existing MFCC firmware from the Xilinx Vitis library, integrated into the FPGA for hardware acceleration. The TensorCore hardware accelerator is deployed in the NVIDIA Jetson edge computing unit, directly calling existing pre-trained AI models via the ONNX format to achieve acoustic and physiological feature fusion, improving diagnostic specificity. The computing power of the NVIDIA Jetson supports real-time inference of complex AI models. The FPGA and ARM share the computational load, reducing end-to-end latency and meeting clinical real-time requirements. Hardware-level fusion of acoustic and physiological data improves disease detection accuracy.
[0010] The human-computer interaction and communication unit includes a detachable AI terminal and a cloud-based collaborative architecture. The AI terminal connects to the main body of the signal analysis and multimodal fusion unit via a magnetic interface. The AI terminal includes a display screen, warning indicator lights, and Bluetooth connectivity. The 1.44-inch TFT screen is installed at a 30° angle, displaying heart sound waveforms, SpO2 values, and diagnostic results. The cloud-based collaborative architecture uploads data to the cloud via a standard Bluetooth / LoRa communication module and interfaces with the hospital's HIS system using the HL7 protocol, supporting remote review of high-risk cases by doctors.
[0011] The beneficial effects of this utility model are:
[0012] (1) This utility model connects the acoustic sensor component and the physiological parameter sensor component to the stethoscope, processes the signal analysis and multimodal fusion unit and calls the existing pre-trained AI model to realize the fusion of acoustic and physiological features, which solves the problem that traditional stethoscopes lack physiological parameter references, rely on a single basis and have low accuracy.
[0013] (2) The FPGA module integrates the MFCC IP core, and the NVIDIA Jetson layout of the edge computing unit has TensorCore hardware accelerators. Through the collaboration and modular interaction of ARM, FPGA and Jetson (detachable terminal and standardized protocol), the end-to-end latency is reduced.
[0014] (3) The acoustic sensor assembly uses a stainless steel bellows with a spiral guide groove on the inner wall to reduce signal attenuation. The anti-interference design of the flexible substrate contact surface and the microlens array of the physiological parameter sensor assembly improves the data acquisition accuracy of the device.
[0015] The stethoscope of this invention enables intelligent diagnosis of cardiopulmonary diseases with multimodal, high precision, and low latency, and is especially suitable for primary healthcare and environments without network access. Attached Figure Description
[0016] Figure 1 This is a structural block diagram of the present utility model;
[0017] Figure 2 This is a structural diagram of the acoustic sensor array;
[0018] Figure 3 This is a structural diagram of the physiological parameter sensor group;
[0019] Figure 4 Here is a flowchart of the preprocessing unit;
[0020] Among them, 1. disposable food-grade silicone mouthpiece, 2. directional acoustic waveguide, 3. piezoelectric sensor, 4. microphone array, 5. magnetic interface, 6. one-way airflow valve, 7. acquisition inlet, 8. dual-wavelength transmission sensor, and 9. microlens array. Detailed Implementation
[0021] The present application will be further described below with reference to specific embodiments.
[0022] Example 1: An intelligent stethoscope based on multimodal data collaborative diagnosis includes an acoustic sensor assembly, a physiological parameter sensor assembly, a signal analysis and multimodal fusion unit, and a human-computer interaction and communication unit. The signal analysis and multimodal fusion unit includes a signal preprocessing unit, a heterogeneous computing core unit, and an edge computing unit connected in sequence. The output terminals of the acoustic sensor assembly and the physiological parameter sensor assembly are connected to the input terminal of the signal preprocessing unit of the signal analysis and multimodal fusion unit. The heterogeneous computing core unit includes ARM and FPGA modules. The ARM and FPGA modules are connected to the edge computing unit through PCIe interface and SPI bus, respectively. The edge computing unit adopts the NVIDIA Jetson embedded system. NVIDIA Jetson directly calls the existing pre-trained AI model through ONNX format to realize the fusion of acoustic and physiological features. The AI model directly outputs the diagnostic results. The output terminal of the edge computing unit is connected to the input terminal of the human-computer interaction and communication unit.
[0023] The acoustic sensor assembly includes a disposable food-grade silicone mouthpiece 1, which is connected to a directional acoustic waveguide 2 via a one-way airflow valve 6 and a magnetic interface 5. A microphone array 4 is fixed to the outlet end of the directional acoustic waveguide 2 to collect cough sounds using beamforming technology. A piezoelectric sensor 3 is installed on the outer wall of the directional acoustic waveguide 2 to collect vibration signals from the waveguide. The disposable food-grade silicone mouthpiece uses a one-way airflow valve to prevent saliva backflow and contamination of the equipment, reducing the risk of cross-infection. The magnetic interface ensures aseptic operation. The directional acoustic waveguide uses a stainless steel corrugated tube design with spiral guide grooves (groove depth 0.8±0.3mm, pitch 3.5±1.5mm) inside the wall to optimize the transmission efficiency of low-frequency cough sounds (100-500Hz) and reduce signal attenuation. The microphone array employs four microphones arranged in a diamond topology, fixed at the outlet of the directional acoustic waveguide. Combined with beamforming algorithms, it directionally acquires cough sounds. The patient coughs through a disposable mouthpiece, and the sound waves are transmitted to the microphone array via the directional acoustic waveguide's flow channel. The ambient noise suppression ratio reaches 20dB. An auxiliary piezoelectric sensor is attached to the outer wall of the directional acoustic waveguide and fixed with epoxy resin, synchronously acquiring vibration harmonic characteristics.
[0024] The physiological parameter sensor assembly is cuboid in shape, with a data acquisition inlet 7 on the front, a dual-wavelength transmissive sensor 8 on one side, and a microlens array 9 on the other side. The internal contact surface of the data acquisition inlet uses a flexible silicone substrate with a nano-hydrophobic coating to conform to the patient's finger. This adaptive deformation enhances skin adhesion stability and suppresses signal attenuation caused by sweat, reducing blood oxygen detection errors. The microlens array employs a gradient refractive index design to effectively suppress motion artifacts. The patient acquires parameters by inserting their finger into the data acquisition inlet, ensuring contact with the internal surface. The dual-wavelength transmissive sensor first transmits red and infrared light through the patient's fingertip. The microlens array uses a gradient refractive index design to compensate for optical path shifts caused by finger movement, suppressing motion artifacts. The transmitted light is focused by a filter in the sensor and received by a photodiode. The received signal is then transmitted to a signal preprocessing unit. The physiological parameter sensor assembly, through its flexible fit structure enhancing signal coupling, dual-wavelength optical detection improving blood oxygen accuracy, and gradient refractive index lenses suppressing motion interference, achieves high-precision, highly interference-resistant blood oxygen saturation detection, meeting the medical-grade requirements for dynamic monitoring scenarios.
[0025] The signal preprocessing unit includes a low-noise amplifier, a bandpass filter, a dynamic baseline correction circuit, and an analog-to-digital converter (ADC) to perform preliminary noise reduction, characteristic frequency band extraction, and high-precision digitization of acoustic and physiological signals. After passing through the amplifier and bandpass filter in the preprocessing module, the acoustic signal undergoes characteristic frequency band extraction and high-precision digitization via the ADC. The physiological signal passes through the dynamic baseline correction circuit using an analog high-pass filter to eliminate low-frequency motion artifacts and respiratory baseline drift, and is then quantized via the ADC.
[0026] The heterogeneous computing core unit includes ARM and FPGA modules, with ARM+FPGA co-designed. The FPGA module connects via I... 2 The S-interface receives pre-processed acoustic signals, and the ARM module receives pre-processed physiological signals via the I3C interface. The ARM and FPGA modules are connected to the edge computing unit via PCIe and SPI bus, respectively.
[0027] The FPGA module integrates an MFCC IP core, and the NVIDIA Jetson edge computing unit features a Tensor Core hardware accelerator. The MFCC IP core directly calls existing MFCC firmware based on the XilinxVitis library, integrated into the FPGA for hardware acceleration. The Tensor Core hardware accelerator is deployed in the NVIDIA Jetson edge computing unit, calling pre-trained AI models through a standardized ONNX interface. It utilizes the Jetson's built-in Tensor Core hardware accelerator to perform multimodal feature fusion operations, improving diagnostic specificity. The pre-trained AI model can employ existing multimodal fusion models to achieve acoustic and physiological feature fusion. For example, the joint cross-attention mechanism in Reference 1 dynamically captures the complementary relationship between acoustic and physiological signals, calculates attention weights based on correlation, reduces intermodal heterogeneity, and provides a pre-trained model in an open-source code library that can be directly called for data fusion. The OpenBioMed open-source platform developed by the Tsinghua University AIR team provides a cross-modal pre-training toolkit, supporting the aligned fusion of acoustic features and physiological signals, which can be directly called through the toolkit. Reference 1:
[0028] [1]Rajasekar GP,De Melo WC,Ullah N,et al.A Joint Cross-AttentionModel for Audio-Visual Fusion in Dimensional Emotion Recognition[J].2022.DOI:10.48550 / arXiv.2203.14779.
[0029] Cough acoustic signals are acquired through acoustic sensor components, and after passing through the preprocessing unit's amplifier, bandpass filter, and ADC, they are then transmitted via I... 2The S-interface (industry standard protocol) directly connects to the FPGA. The FPGA integrates MFCC firmware based on the Xilinx Vitis library to complete acoustic feature extraction. The extracted MFCC feature vector is transmitted to the NVIDIA Jetson via the PCIe 3.0 standard interface. Physiological signals are collected by physiological parameter sensor components, processed by dynamic baseline correction circuit and ADC, and then transmitted to the ARM module via I3C interface (standard communication protocol) to calculate characteristic parameters such as blood oxygen saturation. Subsequently, the data is sent to the NVIDIA Jetson via the SPI bus (Universal Serial Peripheral Interface). The NVIDIA Jetson directly calls the existing pre-trained AI model, which is stored in the Jetson's eMMC memory in ONNX standard format. It is directly loaded by the TensorRT inference framework, and the attention weight calculation is accelerated using hardware Tensor Cores to complete the fusion of multimodal data. Finally, the deployed AI model directly outputs the diagnostic results, which are transmitted to the display screen of the human-computer interaction module.
[0030] The human-computer interaction and communication unit includes a detachable AI terminal and a cloud-based collaborative architecture. The AI terminal connects to the main body of the signal analysis and multimodal fusion unit via a magnetic interface. The AI terminal includes a 1.44-inch TFT display (tilted at 30°), a warning indicator light (triggered by a threshold comparison circuit), and a standard Bluetooth 5.0 module. The display shows heart sound waveforms, SpO2 values, and diagnostic results via existing GUI firmware. After the patient's cough and SpO2 data are collected, the diagnostic results are transmitted to the display via an edge computing module. The warning indicator light is linked to a preset threshold via a hardware comparator, automatically illuminating when data exceeds the limit. The cloud-based collaborative architecture uploads data to the cloud via a standard Bluetooth / LoRa communication module and interfaces with the hospital's HIS system using the HL7 protocol, supporting doctors to remotely review high-risk cases.
[0031] The above is a further description of the present utility model in conjunction with the embodiments, and the protection scope of the present utility model is not limited thereto.
Claims
1. An intelligent stethoscope based on multimodal data collaborative diagnosis, characterized in that, The system includes acoustic sensor components, physiological parameter sensor components, a signal analysis and multimodal fusion unit, and a human-computer interaction and communication unit. The signal analysis and multimodal fusion unit includes a signal preprocessing unit, a heterogeneous computing core unit, and an edge computing unit connected in sequence. The outputs of the acoustic sensor components and physiological parameter sensor components are connected to the input of the signal preprocessing unit of the signal analysis and multimodal fusion unit. The heterogeneous computing core unit includes ARM and FPGA modules, which are connected to the edge computing unit via PCIe interface and SPI bus, respectively. The edge computing unit uses the NVIDIA Jetson embedded system, which directly calls the existing pre-trained AI model to achieve acoustic and physiological feature fusion. The AI model directly outputs diagnostic results. The output of the edge computing unit is connected to the input of the human-computer interaction and communication unit.
2. The intelligent stethoscope based on multimodal data collaborative diagnosis according to claim 1, characterized in that, The acoustic sensor assembly includes a disposable food-grade silicone mouthpiece, which is connected to a directional acoustic waveguide via a one-way airflow valve and a magnetic interface. A microphone array is fixed at the outlet end of the directional acoustic waveguide to collect cough sounds in a directional manner using beamforming technology. A piezoelectric sensor is provided on the outer wall of the directional acoustic waveguide to collect vibration signals of the waveguide.
3. The intelligent stethoscope based on multimodal data collaborative diagnosis according to claim 2, characterized in that, The directional acoustic waveguide is made of stainless steel corrugated pipe with spiral guide grooves on the inner wall.
4. The intelligent stethoscope based on multimodal data collaborative diagnosis according to claim 1, characterized in that, The physiological parameter sensor assembly is rectangular in shape, with a data acquisition inlet on the front, a dual-wavelength transmission sensor on one side, and a microlens array on the other side.
5. The intelligent stethoscope based on multimodal data collaborative diagnosis according to claim 4, characterized in that, The microlens array features a gradient refractive index design and consists of lenses with micron-level aperture and relief depth, effectively suppressing motion artifacts. The internal contact surface of the acquisition inlet uses a flexible silicone substrate with a composite nano-hydrophobic coating to fit the patient's finger.
6. The intelligent stethoscope based on multimodal data collaborative diagnosis according to claim 1, characterized in that, The signal preprocessing unit includes a low-noise amplifier, a bandpass filter, a dynamic baseline correction circuit, and an analog-to-digital converter (ADC) to perform preliminary noise reduction, feature frequency extraction, and high-precision digitization of acoustic and physiological signals. After the acoustic signal passes through the amplifier and bandpass filter of the preprocessing module, feature frequency extraction is achieved, and high-precision digitization is completed by the ADC. The physiological signal uses a dynamic baseline correction circuit with an analog high-pass filter to eliminate low-frequency motion artifacts and respiratory baseline drift in the physiological signal, and the physiological signal is quantized by the ADC.
7. The intelligent stethoscope based on multimodal data collaborative diagnosis according to claim 1, characterized in that, The FPGA module integrates an MFCC IP core, and the NVIDIA Jetson layout of the edge computing unit features a Tensor Core hardware accelerator.
8. The intelligent stethoscope based on multimodal data collaborative diagnosis according to claim 1, characterized in that, The human-computer interaction and communication unit includes a detachable AI terminal and a cloud-based collaborative architecture. The AI terminal is connected to the main body of the signal analysis and multimodal fusion unit via a magnetic interface. The AI terminal includes a display screen, a warning indicator light, and Bluetooth.