Method and apparatus for configuring a heart sound processing model based on multi-modal knowledge distillation

By training a lightweight student model using multimodal knowledge distillation technology, the problem of insufficient generalization ability of existing heart sound processing models under different devices and environments is solved, enabling efficient diagnosis on portable devices and improving diagnostic accuracy and stability.

CN121483574BActive Publication Date: 2026-04-21TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
Filing Date
2026-01-08
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing heart sound processing models rely on single-modal data, resulting in insufficient generalization ability under different acquisition environments and devices. The models are large in scale and computationally complex, making it difficult to achieve real-time inference on portable devices or mobile terminals.

Method used

Employing multimodal knowledge distillation technology, a lightweight student model is trained using multimodal data such as heart sounds, electrocardiograms, electrocardiograms, and electronic medical records through a knowledge distillation mechanism between teacher and student models, thus achieving integrated multimodal training and lightweight deployment.

Benefits of technology

This improves the diagnostic accuracy, stability, and engineering usability of the heart sound processing model on portable devices, meeting the real-time diagnostic needs of mobile and wearable devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121483574B_ABST
    Figure CN121483574B_ABST
Patent Text Reader

Abstract

This application provides a configuration method and apparatus for a heart sound processing model based on multimodal knowledge distillation. It introduces knowledge distillation technology on the basis of deep learning-based heart sound data processing and builds a novel model configuration scheme in a deep and adaptive manner. This realizes the integration of multimodal training and lightweight deployment of knowledge distillation, so as to comprehensively improve the accuracy, stability and engineering usability of diagnosis, and has good application prospects in clinical work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical technology, specifically to a configuration method and apparatus for a heart sound processing model based on multimodal knowledge distillation. Background Technology

[0002] Phonocardiogram (PCG) is an important physiological signal reflecting the opening and closing of heart valves and hemodynamics. Due to its non-invasive, low-cost, and portable collection characteristics, it is widely used in the auxiliary diagnosis of cardiovascular diseases.

[0003] Currently, various automated analysis methods have been proposed. Early research was mostly based on signal processing and traditional machine learning methods. With the development of deep learning, more and more methods are using convolutional neural networks, recurrent neural networks, and Transformer structures to perform end-to-end modeling of heart sound signals, achieving progress in heart murmur recognition, valvular disease classification, and heart sound event detection. Although these methods have achieved high accuracy on public datasets, they usually require large model sizes and computational resources, making them difficult to apply directly to mobile devices or portable terminals.

[0004] However, the inventors of this application have found that the prior art relies too much on single modality data and fails to fully combine multiple clinical information, resulting in insufficient generalization ability under different acquisition environments and devices. At the same time, the model is large in scale and computationally complex, making it difficult to achieve real-time inference on portable devices or mobile terminals. Summary of the Invention

[0005] This application provides a configuration method and apparatus for a heart sound processing model based on multimodal knowledge distillation. It introduces knowledge distillation technology on the basis of deep learning-based heart sound data processing and builds a novel model configuration scheme in a deep and adaptive manner. This realizes the integration of multimodal training and lightweight deployment of knowledge distillation, so as to comprehensively improve the accuracy, stability and engineering usability of diagnosis, and has good application prospects in clinical work.

[0006] Firstly, this application provides a configuration method for a heart sound processing model based on multimodal knowledge distillation, the method comprising:

[0007] Configure a first heart sound processing model as the teacher model and a second heart sound processing model as the student model;

[0008] During the model training phase, the first heart sound processing model is used to guide the second heart sound processing model to learn based on the knowledge distillation mechanism. The multimodal inputs of the first heart sound processing model include heart sound modality, electrocardiogram modality, echocardiogram modality and electronic medical record modality, while the single-modal inputs of the second heart sound processing model include heart sound modality.

[0009] The second heart sound processing model, which learns through a knowledge distillation mechanism, is deployed on a target device used to perform the model's actual application tasks. The target device includes at least one of a portable stethoscope, a mobile terminal, and a wearable device.

[0010] Secondly, this application provides a configuration device for a heart sound processing model based on multimodal knowledge distillation, the device comprising:

[0011] A configuration unit is used to configure the first heart sound processing model as a teacher model and the second heart sound processing model as a student model.

[0012] The training unit is used to guide the second heart sound processing model to learn using the first heart sound processing model based on the knowledge distillation mechanism during the model training phase. The multimodal inputs of the first heart sound processing model include heart sound modality, electrocardiogram modality, echocardiogram modality and electronic medical record modality, and the single-modal inputs of the second heart sound processing model include heart sound modality.

[0013] The deployment unit is used to deploy a second heart sound processing model, which has been learned through a knowledge distillation mechanism, on a target device used to perform the model's actual application tasks. The target device includes at least one of a portable stethoscope, a mobile terminal, and a wearable device.

[0014] Thirdly, this application provides a processing device, including a processor and a memory, wherein a computer program is stored in the memory, and when the processor invokes the computer program in the memory, it executes the method provided by the first aspect of this application or any possible implementation of the first aspect of this application.

[0015] Fourthly, this application provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to perform the method provided in the first aspect of this application or any possible implementation thereof.

[0016] From the above, it can be concluded that this application has the following beneficial effects:

[0017] Targeting the goal of heart sound processing, this application introduces knowledge distillation technology on the basis of deep learning-based heart sound data processing, and builds a novel model configuration scheme with deep adaptation, realizing the integration of multimodal training and lightweight deployment of knowledge distillation, so as to comprehensively improve the accuracy, stability and engineering usability of diagnosis, and has good application prospects in clinical work. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic flowchart of a configuration method for a heart sound processing model based on multimodal knowledge distillation, as described in this application.

[0020] Figure 2 This is a schematic diagram of the configuration device for the heart sound processing model based on multimodal knowledge distillation in this application;

[0021] Figure 3 This is a schematic diagram of one type of processing equipment used in this application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. The naming or numbering of steps appearing in this application does not imply that the steps in the method flow must be performed in the chronological / logical order indicated by the naming or numbering. The execution order of named or numbered process steps can be changed according to the desired technical purpose, as long as the same or similar technical effect is achieved.

[0024] The module division described in this application is a logical division. In practical applications, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection between modules shown or discussed may be through some interfaces, and the indirect coupling or communication connection between modules may be electrical or other similar forms, none of which are limited in this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed in multiple circuit modules. Some or all of the modules may be selected to achieve the purpose of the solution in this application according to actual needs.

[0025] Before introducing the configuration method of the heart sound processing model based on multimodal knowledge distillation provided in this application, we will first introduce the background content involved in this application.

[0026] The configuration method, apparatus, and computer-readable storage medium for a heart sound processing model based on multimodal knowledge distillation provided in this application can be applied to processing devices. It introduces knowledge distillation technology on the basis of deep learning-based heart sound data processing and builds a novel model configuration scheme in a deeply adapted manner, realizing the integration of multimodal training and lightweight deployment of knowledge distillation, so as to comprehensively improve the accuracy, stability and engineering usability of diagnosis, and has good application prospects in clinical work.

[0027] The configuration method for the heart sound processing model based on multimodal knowledge distillation mentioned in this application can be implemented by a configuration device for the heart sound processing model based on multimodal knowledge distillation, or by different types of processing devices such as servers, physical hosts, or user equipment (UE) that integrate the configuration device for the heart sound processing model based on multimodal knowledge distillation. The configuration device for the heart sound processing model based on multimodal knowledge distillation can be implemented in hardware or software. The UE can specifically be a terminal device such as a smartphone, tablet, laptop, desktop computer, or personal digital assistant (PDA). The processing devices can be configured in a device cluster manner.

[0028] It is understandable that the solution in this application is usually based on existing data or data that has already been collected. Therefore, the processing device that executes the configuration method of the heart sound processing model based on multimodal knowledge distillation of this application or the application service corresponding to the configuration method of the heart sound processing model based on multimodal knowledge distillation of this application usually only needs to meet the required data processing capabilities, and its specific device type and device deployment form are quite flexible.

[0029] If the direct acquisition of existing data mentioned above is involved, then further hardware and software adaptations are needed for the processing equipment to enable it to acquire data. For example, if real-time acquisition of heart sounds and other data is required, the corresponding data acquisition equipment can be incorporated into the processing equipment cluster, or the processing equipment itself can be the control unit of the data acquisition equipment. Alternatively, a third-party call can be used to trigger external data acquisition equipment to perform real-time data acquisition operations.

[0030] Furthermore, if there is a need to display the processing progress (including processing results, especially the heart sound diagnosis results given by the heart sound processing model), the processing device itself can be configured with the necessary display screen (including a touch screen) to display the specific content. Of course, the processing device can also display the specific content through an external display device or other devices with a display screen. Clinical work often involves warning outputs based on alarm output components such as buzzers and indicator lights, i.e., warning outputs combining heart sound diagnosis results and corresponding threshold / feature conditions, which are similar to the content display settings here.

[0031] The following section introduces the configuration method of the heart sound processing model based on multimodal knowledge distillation provided in this application.

[0032] First, refer to Figure 1 , Figure 1 The diagram illustrates a configuration method for the heart sound processing model based on multimodal knowledge distillation according to this application. The configuration method for the heart sound processing model based on multimodal knowledge distillation provided by this application may specifically include the following steps S101 to S103:

[0033] Step S101: Configure the first heart sound processing model as the teacher model and the second heart sound processing model as the student model;

[0034] Understandably, this application specifically introduced knowledge distillation technology in the early model configuration work for the heart sound processing model to be applied in clinical work. Based on this, combined with the model configuration scheme specially designed in this application, a high-performance second heart sound processing model is obtained as the final heart sound processing model that can be used in practice.

[0035] Knowledge distillation (KD) is a technique that uses a high-performance "teacher model" to transfer knowledge to a lightweight "student model" during the training phase. It can significantly compress the model size while maintaining diagnostic accuracy as much as possible. Unlike transfer learning, the mode adopted in this application allows the teacher model to use multimodal, large-scale, or privileged information-containing data during the training phase, while the student model only retains lightweight input based on heart sound signals for deployment. This achieves the characteristic of "multimodal support during training and single-modal support during application," which will be specifically demonstrated in the subsequent scheme introduction.

[0036] Understandably, the configuration of the first and second heart sound processing models here is part of the preliminary model configuration work. For example, the corresponding initial model may be automatically extracted from local or remote locations, or the initial model may be built under the manual operation of relevant users. In addition, the application of knowledge distillation technology may also involve related model configuration work.

[0037] Step S102: In the model training phase, the first heart sound processing model is used to guide the second heart sound processing model to learn based on the knowledge distillation mechanism. The multimodal input of the first heart sound processing model includes heart sound mode, electrocardiogram mode, echocardiogram mode and electronic medical record mode, and the single-modal input of the second heart sound processing model includes heart sound mode.

[0038] It is easy to see that the teacher model, i.e. the first heart sound processing model, is a multimodal model, specifically involving multimodal input data of heart sounds (Phonocardiogram, PCG), electrocardiogram (ECG), echocardiography (Echo) and electronic health record (EHR), while the student model, i.e. the second heart sound processing model, is a unimodal model, specifically involving unimodal input of heart sounds.

[0039] At this point, we can see that the model configuration of this application, which relies on multimodal models during the training phase and only requires single-modal models during the application phase, helps to significantly improve the generalization and cross-scenario applicability of the model while ensuring lightweight edge devices.

[0040] In the training or learning phase of a model based on a knowledge distillation mechanism, the system boundary allows access to multimodal data, including heart sounds as the core input, as well as optional electrocardiograms, structured echocardiogram indicators, and electronic health record text summaries. After processing and alignment, the multimodal data is input into the teacher model, i.e., the first heart sound processing model. Used to generate soft tags and time-frequency mask prior Then, through a knowledge distillation mechanism, it is passed on to the student model, namely the second heart sound processing model. .

[0041] Specifically, the data collection and processing of training samples involves the following steps:

[0042] 1.1) Data Acquisition

[0043] Heart sounds are acquired using an electronic stethoscope or multichannel microphone, with a sampling rate typically between 2 and 4 kHz. Locus of sound, body position, and device parameters are simultaneously recorded in metadata. During the training period, electrocardiograms (sampling rate ≥ 250 Hz) can be acquired for cardiac cycle alignment and temporal prior generation; echocardiograms and electronic medical records are converted into weakly labeled or auxiliary features in a structured manner.

[0044] Data time synchronization is achieved through hardware clocks (NTP / PTP / GPS) or software alignment (cross-correlation and dynamic time warping) to ensure that the R peak of the electrocardiogram is aligned with the S1 of the heart sound, and to normalize each cardiac cycle to a uniform length. Event labels (such as S1, S2, S3, S4, OS, EC, MC) and task labels such as murmur level and valvular disease type are mapped to the same time axis to form a complete multi-task training sample.

[0045] Among them, S1, S2, S3, S4, OS, EC, and MC are existing parameters in clinical heart sound diagnosis, representing the first heart sound, second heart sound, third heart sound, fourth heart sound, mitral valve opening snap, ejection click, and mitral valve closing click, respectively.

[0046] 1.2) Data Quality

[0047] To ensure data quality, this application can also automatically classify all heart sound segments.

[0048] Based on metrics such as spectral energy ratio, artifact detection, and effective period count, signals are classified into three categories: high quality (Q2), usable (Q1), and low quality (Q0). Q0 samples are removed during training and trigger secondary acquisition or delay judgment during deployment. Data that has passed quality screening undergoes further bandpass filtering (25-400Hz), wavelet or spectral subtraction denoising, amplitude normalization, and resampling normalization to ensure input consistency.

[0049] 1.3) Data Augmentation

[0050] Data augmentation employs strategies such as time scaling, random noise mixing, SpecAugment, and time-frequency domain mixing to improve the robustness of the model across multiple scenarios and devices.

[0051] 1.4) Data Processing

[0052] To address the frequency response differences among different acquisition devices, this application can also introduce amplitude-frequency response correction and add device embedding vectors to the model. Used for cross-device adaptation.

[0053] All data is managed through a unified data dictionary, and fields and value ranges are divided into training, validation and test sets by patient or center to avoid information leakage.

[0054] Original and derived data are version-controlled and accompanied by hash identifiers to ensure the traceability of data governance.

[0055] Regarding the model input, this application may also involve configuration work for input representation and multimodal fusion, as detailed below:

[0056] 2.1) Heart sound waveform Two-dimensional time-frequency representation is obtained by performing a short-time Fourier transform (STFT) or a log-Melbourne spectrum transform. It can use a 25-millisecond window, a 10-millisecond step size, and 64-80 frequency bands.

[0057] To mitigate the impact of individual differences and heart rate variations, this application can also superimpose a periodic normalization channel into the input, enabling the model to remain stable across patients and scenarios.

[0058] 2.2) For the student model, i.e., the second heart sound processing model In terms of model design, it can be based on a lightweight CNN-TCN architecture or a lightweight Conformer architecture, using depthwise separable convolutions and dilated convolutions to jointly extract time-frequency and temporal features, and introducing device embedding vectors through conditional normalization layers. This will further enhance cross-device adaptability.

[0059] By setting a multi-task head at the output end, it is possible to specifically realize heart murmur detection, valvular disease classification, heart sound event recognition, signal quality assessment, and cardiac function level prediction, thus completing multi-task diagnosis within a single model framework.

[0060] 2.3) For the teacher model, i.e., the first heart sound processing model In terms of model design, a higher-capacity heart sound branch can be added to the student model, and modal branches such as electrocardiogram, electrocardiogram and electronic health record can be introduced.

[0061] Each modal branch is aligned and fused through a cross-modal attention mechanism to generate a fine-grained time-frequency mask. Multi-tasking soft tags (i.e., output probability distribution) and intermediate representations It is then passed to the student model during the distillation stage for logit distillation, feature distillation, and mask distillation.

[0062] 2.4) Meanwhile, during the characterization and fusion stage, this application can also add structured constraints to ensure logical consistency between tasks, such as the valvular disease type needing to match the murmur phase, the S1 / S2 boundary needing to be continuous with the mask region, and the mask confidence of low-quality signals needing to be lowered to avoid noise propagation.

[0063] Next, let's examine the specific distillation path designed under the knowledge distillation mechanism in this application, or the complementary paths covered by the distillation process.

[0064] Understandably, the core of this application lies in using a knowledge distillation mechanism to transform the multimodal teacher model during the training period. The diagnostic capabilities can be transferred to a lightweight student model that relies solely on heart sounds. This approach balances performance and real-time capabilities during deployment. The distillation process encompasses multiple complementary paths, including output layer logit distillation, intermediate layer feature distillation, time-frequency mask distillation in the time-frequency domain, comparative distillation of samples, and uncertainty adjustment of confidence levels. This ensures that the student model, while structurally compact, reproduces the discriminative ability of the teacher model as closely as possible. Specifically, these include:

[0065] 3.1) Logit (output probability distribution) distillation

[0066] Logit distillation is performed at the output layer. The temperature-based probability distribution generated by the teacher model for the input samples is represented as follows:

[0067] ,

[0068] in, Let the teacher's logit vector be... The distillation temperature. , This represents the Softmax function.

[0069] The output prediction distribution of the student model is represented as follows:

[0070] ,

[0071] in, This is the student's logit vector.

[0072] and The difference between the two, constrained by the Kullback-Leibler divergence (or KL divergence, relative entropy), can be expressed as:

[0073] ,

[0074] This loss guides the student model to learn the teacher model's "hidden knowledge" at class boundaries, thereby further improving generalization ability.

[0075] 3.2) Characteristic distillation

[0076] Feature distillation is performed in the intermediate layer, and the intermediate representation of the teacher model is denoted as... The student model corresponding to the layer representation is denoted as The two are connected through a linear mapping function. After aligning the dimensions, the Euclidean distance is minimized, which is expressed as:

[0077] ,

[0078] This ensures that the student model can capture the same time-frequency representation structure as the teacher model.

[0079] 3.3) Time-frequency mask distillation

[0080] At the time-frequency domain level, the teacher model outputs a attention mask. This indicates its discriminative region on the heart sound time-frequency map; correspondingly, the student model simultaneously predicts the mask. Both are jointly constrained by Dice loss and cross-entropy:

[0081] ,

[0082] This design enables students to maintain sensitivity to critical phases and noise areas even in noisy environments.

[0083] 3.4) Comparative distillation

[0084] To enhance cross-sample consistency, this application also introduces comparative distillation, whereby students embed [the following information] for different enhanced samples or different sampling points of the same subject. embedded with teachers Samples should be kept similar, while samples from different diseases should be kept distinct, and this should be achieved through InfoNCE format, represented as follows:

[0085]

[0086] in, Indicates temperature parameter, This represents the inner product similarity.

[0087] 3.5) Adjusting for uncertainty in confidence levels

[0088] Furthermore, considering the noise and data uncertainty present in clinical settings, the prediction confidence of the teacher model can also be used to adjust the student loss weights. Therefore, this application employs an uncertainty regularization term. Suppress overconfident predictions and improve robustness.

[0089] In summary, the training loss of the student model is a weighted combination of multiple terms:

[0090] ,

[0091] in, These are dynamically adjusted weighting coefficients.

[0092] This design ensures that the student model can still maintain multi-task diagnostic accuracy, temporal consistency, and cross-domain robustness under limited parameter conditions.

[0093] Thus, the specific configuration of the distillation path described above enables the student model to not only learn the soft knowledge of the teacher model at the classification level, but also inherit its intermediate feature structure and time-frequency attention region. Through comparative learning, it enhances consistency across individuals and across devices, breaking through the limitations of the traditional single distillation method and ensuring the comprehensiveness and robustness of heart sound diagnosis under a limited parameter scale.

[0094] Regarding the specific model training process, this application comprises three stages: teacher model training, student model pre-training, and joint distillation training. Through phased optimization and dynamic weight control, it ensures that the student model maintains diagnostic capabilities close to those of the teacher model even with limited computational resources. Specifically:

[0095] 4.1) Loss function during the teacher model training phase Specifically, it can be expressed as:

[0096] ,

[0097] in, Represents the standard cross-entropy loss function. Indicates the main task label. Indicates the weight of auxiliary tasks. This indicates auxiliary labels that include event segmentation and signal quality.

[0098] The output of this stage includes soft targets. Intermediate representation Time and frequency masks This serves as a source of knowledge for subsequent distillation.

[0099] 4.2) Loss function during the pre-training phase of the student model Specifically, it can be expressed as:

[0100] ,

[0101] in, This represents the predicted distribution output by the student model. , This represents the Softmax function. Represents the student's logit vector. The distillation temperature. ;

[0102] As can be seen, the student model The input is only a single-mode waveform of heart sounds, and it is trained using standard cross-entropy loss. This stage enables the model to have preliminary classification and event recognition capabilities, ensuring that the model converges stably when entering the distillation stage.

[0103] 4.3) Loss function during the joint distillation training phase Specifically, it can be expressed as:

[0104] ,

[0105] in, For dynamically adjusted different weighting coefficients, This represents the logit distillation loss. Indicates characteristic distillation loss, This indicates the time-frequency mask distillation loss. This indicates the difference in distillation losses. This indicates a regularization term indicating uncertainty.

[0106] As can be seen, during the joint distillation training phase, the student model is simultaneously supervised by real labels and constrained by teacher knowledge.

[0107] To avoid excessive noise in the early distillation stages, this application also employs phased weighted scheduling, which can be dynamically adjusted at different specific stages during the joint distillation training phase. To carry out specific scheduling.

[0108] As an example, to avoid excessive distillation noise in the early stages of training, only supervised loss and logit distillation can be enabled. In the medium term, characteristic distillation and mask distillation will be gradually introduced. Later, comparative distillation and uncertainty adjustment were added to further improve cross-domain robustness and generalization performance.

[0109] Continuing with the loss function during the joint distillation training phase here The specific details of each item involved in the different distillation routes will be explained.

[0110] , ,

[0111] in, This represents the temperature probability distribution output by the teacher model. and Kullback-Leibler divergence, Represents the teacher's logit vector.

[0112] ,

[0113] in, Represents the square of the L2 norm. This represents the intermediate representation of the teacher model. Represents a linear mapping function. This represents the layer representation of the student model.

[0114] ,

[0115] in, This represents the mask predicted by the teacher model. This represents the mask predicted by the student model. for and Dice loss, express and Standard cross-entropy loss,

[0116]

[0117] in, Indicates student embedding, Indicates teacher embedding, Indicates inner product similarity. This represents the temperature parameter.

[0118] Meanwhile, regarding model training, in addition to the aforementioned loss configuration, this application also involves another optimization configuration. Specifically, the optimization strategy configured during training may include the following processing:

[0119] 5.1) The Adamw optimizer is used;

[0120] 5.2) Learning rate Using cosine annealing scheduling can be specifically represented as:

[0121] ,

[0122] in, The initial learning rate, For the annealing cycle;

[0123] 5.3) Dynamic weight adjustment adopts the GradNorm method based on gradient norm to keep different loss components balanced during training;

[0124] 5.4) Introduce two regularization methods: Dropout and SpecAugment;

[0125] Various regularization measures such as Dropout and SpecAugment are used to prevent overfitting and improve generalization ability. Specifically, Dropout and SpecAugment can be used to prevent student models from memorizing specific spectral patterns.

[0126] 5.5) An early stopping strategy is introduced to monitor the multi-task performance on the validation set to prevent overfitting;

[0127] 5.6) Introduce mixed-precision training involving FP16 or INT8 low-precision formats to reduce memory usage and accelerate convergence.

[0128] 5.7) Quantize and prune the student model during training to ensure resource constraints during deployment.

[0129] Step S103 involves deploying the second heart sound processing model, which has been learned through a knowledge distillation mechanism, onto a target device used to perform the model's actual application tasks. The target device includes at least one of a portable stethoscope, a mobile terminal, and a wearable device.

[0130] Understandably, once the training / learning based on the knowledge distillation mechanism is completed, the lightweight student model to be used, namely the second heart sound processing model, can then be deployed to portable stethoscopes (or handheld digital stethoscopes), mobile terminals, and wearable devices that may be involved in clinical work.

[0131] Specifically, mobile terminals can include user-side terminal devices such as personal digital assistants (PDAs), smartphones, tablets, laptops, and smart bracelets. The deployment can be carried out by using a compressed model to deploy the device's processing unit for easy operation.

[0132] Thus, by fully utilizing the advantages of multimodal and large-scale models in the training phase to obtain the second heart sound processing model, it can maintain efficient diagnostic performance in a lightweight manner during the application phase, thus addressing the shortcomings of existing technologies in terms of robustness, real-time performance, and task comprehensiveness.

[0133] Correspondingly, the method in this application may also specifically involve the processing of the corresponding model application stage:

[0134] Heart sound analysis and processing are performed based on the second heart sound processing model.

[0135] During this deployment phase, it is understandable that this application also includes further optimization settings, specifically:

[0136] 1.1) Sliding Window Inference Mechanism

[0137] The deployment strategy for the second heart sound processing model when deployed on the target device may include the following processing content:

[0138] To ensure real-time performance, a sliding window inference mechanism can be used to segment the continuous heart sound signal waveform into segments of length [missing information]. The windows, spaced apart by steps The sliding operation, with each window independently performing time-frequency transformation and model forward inference, has an approximate complexity of:

[0139] ,

[0140] in, The output represents an approximate complexity. Indicates the total length of the input sequence. This indicates the computational complexity of a single-window model.

[0141] Thus, through reasonable selection and For example, choosing a solution ( This helps to significantly reduce overall latency while ensuring diagnostic accuracy.

[0142] 1.2) Model Quantization and Pruning

[0143] The deployment strategy for the second heart sound processing model when deployed on the target device also includes the following processing:

[0144] Quantization and pruning of the second heart sound processing model;

[0145] The quantization strategy employs symmetric INT8 fixed-point quantization to ensure that convolutional and fully connected layers operate in the integer domain;

[0146] The pruning strategy structurally reduces redundant channels, and the corresponding constraints are:

[0147] ,

[0148] in, This indicates the accuracy of the model after pruning. This represents the model's accuracy before pruning. This indicates the allowable accuracy loss threshold, which can be set to 2%.

[0149] As is understandable, model quantization and pruning have been briefly mentioned in the previous section on optimization configuration during model training. This operation is performed before the model is used, either during model training or during deployment.

[0150] For model quantization, it helps ensure that convolutional and fully connected layers operate in the integer domain, thereby reducing inference latency to 30-40% of the original floating-point model; for model pruning, it helps to structurally reduce redundant channels.

[0151] Thus, after quantization and pruning, the model's parameter size is controlled within the range of 1-2 million, and the single-window inference latency is less than 25 milliseconds, enabling it to run smoothly on mainstream mobile SoCs or low-power processors.

[0152] Furthermore, this application also includes corresponding optimized configurations for actual use after deployment, i.e., the model's inference work. Specifically, these include:

[0153] 2) Signal quality self-checking mechanism

[0154] The second heart sound processing model completes its inference process on the target device, and the introduced signal quality self-checking mechanism includes the following processing:

[0155] The input fragment is first processed by a quality head to predict its quality level. , , The signal quality levels are represented from low to high;

[0156] when When this happens, the system triggers a secondary acquisition or a delayed judgment to prevent noise fragments from directly entering the diagnostic process;

[0157] when At this time, the output results are marked with a low confidence level, indicating a need for clinical review;

[0158] when At that time, the result directly enters the task layer.

[0159] It is easy to see that the above-mentioned signal quality self-test operation architecture can effectively reduce the risk of false alarms and misdiagnosis in practical applications, bringing a better user experience.

[0160] Furthermore, regarding the student model, i.e. the second heart sound processing model, as briefly mentioned in the previous introduction, it can involve multiple tasks based on a single modal input. This can be further illustrated by the task space configuration below.

[0161] 3) Task Space

[0162] 3.1) The second heart sound processing model will divide the task space Multi-task output heads are uniformly mapped to the same backbone network to achieve resource sharing and information complementarity.

[0163] It is understandable that when the student model, i.e. the second heart sound model, is deployed on the edge, it can simultaneously complete the various clinically relevant tasks described above, avoiding the redundant architecture of "single task - multiple models" in traditional methods.

[0164] Correspondingly, the time-frequency representation of the model input The shared representation vector is obtained through the backbone network. and with the mission space By mapping multiple task output heads to the same backbone network, resource sharing and complementary use of cross-task information are achieved through a unified shared backbone and multiple task output heads. Structured consistency constraints are introduced to ensure that predictions for different tasks logically conform to clinical common sense, and the parallel predictions mentioned above are achieved through multiple task heads.

[0165] 3.2) The second heart sound processing model performs five prediction tasks under single-modal input conditions, specifically involving heart murmur detection, valvular disease classification, heart sound event detection, signal quality assessment, and heart function level prediction.

[0166] Specifically, heart murmur detection can involve grade determination, valvular disease classification can involve site identification, heart sound event detection can involve phase boundary marking, and cardiac function grade prediction can involve specific grading methods such as NYHA classification and Killip classification.

[0167] 3.2.1) Heart murmur detection

[0168] For heart murmur detection, a Softmax classifier is used to output the presence or absence of murmur and the murmur level. , The predicted probability is expressed as:

[0169] ,

[0170] in, Indicates a shared representation vector. , This represents the time-frequency representation of heart sounds (or the time-frequency diagram of heart sounds) of the model input.

[0171] 3.1.2) Classification of Valvular Diseases

[0172] In valvular disease classification, the task header outputs the disease type. , Information on the location and body part.

[0173] in, These respectively represent aortic stenosis, aortic regurgitation, mitral stenosis, and mitral regurgitation.

[0174] This output can be jointly modeled with the murmur level, introducing consistency constraints into the training loss. For example, if aortic stenosis is predicted, the corresponding murmur will be mainly concentrated during systole.

[0175] 3.2.3) Heart sound event detection

[0176] In the heart sound event detection task, the model needs to provide the boundary sequence of events such as S1 and S2.

[0177] In practice, this application employs a time series binary classification method, predicting at each time step whether the event belongs to the event boundary, with the output represented as follows:

[0178] ,

[0179] in, This indicates the results of heart sound event detection. This represents the Sigmoid function. Indicates the first Frame features.

[0180] Prediction results and masks generated by the teacher model Used together for time-frequency interpretation.

[0181] 3.2.4) Signal quality assessment

[0182] Signal quality evaluation task output , , The signal quality levels indicated range from low to high, and are used for self-checking and verification mechanisms during deployment.

[0183] 3.2.5) Prediction of cardiac function level

[0184] The cardiac function grade prediction task can use a graded classifier (NYHA I-IV) or regression output (predicting ejection fraction intervals), the output of which also comes from a shared representation vector. .

[0185] This application can also package the results of each task in a structured format for the model output, including task labels, prediction confidence and time-frequency masks. The final output not only provides diagnostic conclusions, but also intuitively presents the model’s focus area and event phases, which is convenient for doctors to review and interpret clinically.

[0186] 3.3) Safety and compliance requirements

[0187] In the second heart sound processing model, all data is stored locally in a single-center application scenario, and the original heart sounds are not uploaded. In a multi-center collaborative scenario, a federated distillation mode is adopted.

[0188] In the federated distillation model, each center runs a teacher model locally to generate soft labels or intermediate representations, which are then uploaded to a central server for aggregation. The training parameters of the student model are updated synchronously in each center.

[0189] The above mechanism ensures that "data does not leave the domain," thus protecting privacy and enabling cross-center knowledge sharing, which has better application value in practical situations.

[0190] In this way, together with several other aspects, a diagnostic assurance mechanism is formed that combines interpretability masking, signal quality grading, security confidence thresholds, and federated distillation.

[0191] 3.4) Explainability and security design

[0192] 3.4.1) Explainability of presentation

[0193] In inference operations, the prediction mask is displayed in the visualization interface that outputs the model's prediction results. Superimposed on the corresponding time-frequency plot of the heart sound signal input to the model Above, the event boundaries of the heart sound event detection are displayed corresponding to the mask.

[0194] The former helps to intuitively label the energy and temporal regions that the model focuses on, while the latter makes it easier for doctors to verify the consistency of "conclusion-basis-temporal" (for example, in the case of aortic stenosis, the mask should mainly cover the higher frequency band of systole), thus improving clinical acceptability.

[0195] 3.4.2) Safety Thresholds and Quality Defense Lines

[0196] The system uses a combination of confidence and uncertainty gating:

[0197] when or In such cases, instead of providing a direct conclusion, the process is transferred to "manual review / supplementary sampling";

[0198] in, Indicates the confidence level. Indicates uncertainty, threshold and It can be configured according to the scenario (pre-hospital emergency care / outpatient clinic).

[0199] Signal quality head output As the first line of defense:

[0200] Q0 rejects the diagnosis and prompts for resampling; Q1 is marked with a low confidence level and proceeds for review; Q2 is released normally, thus reducing misjudgments caused by noise and artifacts at the source.

[0201] 3.4.3) Privacy Compliance and Human-Machine Collaboration

[0202] In multi-center collaboration, federated distillation is used to exchange only teacher soft tags / intermediate projections and not transmit original heart sounds, thus achieving privacy compliance by ensuring that "data does not leave the domain".

[0203] The terminal interface simultaneously displays the diagnostic conclusions, Visualization and confidence levels support doctors to confirm / revise with one click. All manual revisions are fed back into incremental learning and recalibration to continuously improve the model's safety and clinical acceptability.

[0204] In summary, regarding the above solutions, this application, based on deep learning for heart sound data processing, introduces knowledge distillation technology and deeply adapts and builds a novel model configuration scheme to achieve integrated multimodal training and lightweight deployment of knowledge distillation, thereby comprehensively improving the accuracy, stability, and engineering usability of diagnosis, and has good application prospects in clinical work.

[0205] In terms of details, the beneficial effects that this application can achieve are specifically as follows:

[0206] 1) Combining multimodal training with single-modal deployment improves diagnostic performance and generalization ability.

[0207] Existing methods often rely on a single heart sound signal, resulting in insufficient robustness across different devices and scenarios. In contrast, this application introduces multimodal data such as heart sounds, electrocardiograms, electrocardiograms, and electronic health records as input to the classroom model during the training phase. Through a knowledge distillation mechanism, its diagnostic capabilities are transferred to a lightweight student model that relies solely on heart sounds. This eliminates the need for additional multimodal data collection during deployment, while still maintaining high accuracy and cross-device stability, thus improving the versatility of practical applications.

[0208] 2) Lightweight design and distillation optimization meet the real-time needs of mobile devices and wearables.

[0209] Existing deep learning methods have large model sizes, making it difficult to run in real time on edge devices. In contrast, this application compresses the model size through multi-level distillation (logit distillation, feature distillation, time-frequency mask distillation, etc.) and combines quantization and pruning techniques to keep the student model within the millions of parameters. The single-window inference latency is less than 25 milliseconds, enabling it to run efficiently on mobile and wearable devices and meet the requirements of real-time assisted diagnosis.

[0210] 3) A unified multi-task diagnostic framework covering core clinical needs.

[0211] Traditional methods are typically designed for single tasks such as heart murmur detection or valvular disease classification, making it difficult to complete a comprehensive diagnosis within a single system. In contrast, this application constructs a unified multi-task output space that can simultaneously perform heart murmur detection, valvular disease classification, heart sound event recognition, signal quality evaluation, and cardiac function level prediction. This avoids redundant computation caused by parallel multi-model operation, improves diagnostic efficiency and task coverage, and better meets actual clinical needs.

[0212] 4) Explainability and safety mechanisms enhance clinical credibility and acceptability of application.

[0213] This application outputs time-frequency masks and event boundaries during the inference process, which can intuitively display the model's areas of interest and diagnostic criteria. At the same time, it establishes a multi-level security gating mechanism by combining confidence level, uncertainty estimation and signal quality classification. When the input is uncertain or the signal quality is poor, it triggers manual review or secondary acquisition, thereby effectively reducing the risk of misdiagnosis and improving the safety and reliability in clinical use.

[0214] To better understand the above content, we can further illustrate it with the two application examples shown below.

[0215] Application Example 1: Multi-task Intelligent Heart Sound Diagnosis System for Multimodal Teachers and Unimodal Students

[0216] 1.1) Scope of application and system composition

[0217] Scope of application: Adult heart sound screening and auxiliary diagnosis for outpatient clinics, wards, pre-hospital emergency care and home follow-up.

[0218] Hardware components: Electronic stethoscope (4kHz / 16-bit, supporting 4 auscultation points: aorta, pulmonary artery, tricuspid valve, mitral valve), optional single-lead / tri-lead ECG acquisition module (≥250Hz), Android or iOS mobile terminal (SoC with 8GB RAM or above), Bluetooth 5.0.

[0219] Software components: mobile app (for data collection / display / local inference), training platform (for teacher / student training and distillation), federated distillation server (optional, aggregating soft tags and model parameters).

[0220] 1.2) Data Acquisition and Preprocessing

[0221] 1.2.1) Data Acquisition Process and Key Points

[0222] Subjects were asked to breathe calmly, either supine or sitting; auscultation was taken at the four auscultation points in sequence, for 15 seconds at each point; if necessary, auscultation was also taken in both supine and sitting positions.

[0223] The PCG sampling rate is 4kHz; if an ECG is connected, the sampling rate is set to 500Hz to facilitate R-peak detection.

[0224] Metadata also records: gender / age group (de-sensitized), body position, auscultation point, device ID, and environmental noise level (A-weighted, optional).

[0225] 1.2.2) Quality Grading and Noise Reduction

[0226] The fragments were pre-screened using a quality head: the broadband SNR proxy, artifact ratio, and effective cardiac cycle count were calculated to obtain the grade. , .

[0227] Q0 is removed from the training set; Q0 triggers a "re-sampling / delay determination" during deployment.

[0228] Subsequently, bandpass filtering (25-400Hz, optional 50 / 60Hz notch filtering), wavelet or spectral subtraction denoising, and amplitude normalization were performed on the PCG.

[0229] ECG was performed with a bandpass of 5-45 Hz and the R-peak was labeled with a Pan-Tompkins or convolutional R-peak detector.

[0230] 1.2.3) Alignment and Window Splitting

[0231] Soft alignment is performed using R-peak-S1 phase estimation;

[0232] Dynamic time warping (DTW) is used to normalize each cardiac cycle to a uniform length.

[0233] PCG converted to time-frequency plot: STFT window length 25ms, step size 10ms, 80-dimensional logarithmic Mel spectrum, superimposed with "period normalized channel".

[0234] Training / Inference Window: Window Length s, step size s.

[0235] 1.2.4) Data Augmentation

[0236] Random time scaling (±10%), perturbation noise (target SNR 6-20dB), SpecAugment (temporal masking ≤10 frames × 2, frequency masking ≤8 bands × 2), random time shift (±80ms), and device response simulation (convolution of amplitude-frequency response library by device ID) are used to improve cross-device robustness.

[0237] 1.3 Model Structure and Parameters

[0238] 1.3.1) Student Model (End-side deployment)

[0239] Lightweight Conformer / CNN-TCN hybrid backbone:

[0240] Four coding blocks, each containing depthwise separable convolutions (kernel=3, dilation={1,2,4,8} alternating), lightweight self-attention (4 heads, 64 channels), residuals, and layer normalization; channel base width 64, Dropout 0.2, total parameters ≈1.3M. Embedded using a conditional normalization / FiLM injection device. .

[0241] Multi-task head:

[0242] Noise detection / level (0–6) Softmax;

[0243] Valvular heart disease classification (AS / AR / MS / MR / Other);

[0244] Frame-by-frame Sigmoid for event boundary sequences (S1 / S2 / noise intervals);

[0245] Signal quality (Q0 / Q1 / Q2) Softmax;

[0246] Heart function class (NYHA I-IV) Softmax or segmental regression.

[0247] 1.3.2) Teacher Model (For use during training)

[0248] The system uses a high-capacity PCG branch (base width 256, 8 encoding blocks) + an ECG branch (1D-CNN + BiGRU×2) + structured Echo / EHR features (32-dimensional gated vectors). Cross-modal attention is used to complete alignment and fusion. Output: soft labels. Intermediate representation With time-frequency mask .

[0249] 1.4) Training and Distillation Process

[0250] 1.4.1) Teacher Training

[0251] Input: (PCG, optional ECG, Echo / EHR features). Loss:

[0252] ,

[0253] .

[0254] Optimization: AdamW (lr=3e-4, wd=1e-4), Cosine annealing ( ) epochs, batch size 48, early stop tolerance 8 epochs.

[0255] Export .

[0256] 1.4.2) Student Pre-training

[0257] PCG input only , AdamW(lr=2e-3), 30epochs.

[0258] 1.4.3) Combined distillation

[0259] Total student losses:

[0260] ,

[0261] Distillation temperature ; Dice+CE; InfoNCE (queue size 512, ).

[0262] Weighted scheduling:

[0263] 0-10 epochs: ;

[0264] 11-30 epochs: ;

[0265] 31-50 epochs: .

[0266] GradNorm is used to balance the gradients of multiple task heads; SpecAugment and focal-CE are used for unbalanced classes.

[0267] The validation set is stratified by "patient + center".

[0268] 1.5) Compressed Deployment and End-Side Inference

[0269] 1.5.1) Pruning and Quantification

[0270] Structured channel pruning ratio 30%, constraint Symmetric INT8 quantization (calibration set 2,000 window) was used, with edge-side ARM-NEON / NNAPI inference. The final model parameters are approximately 0.95-1.3M, and the file size is approximately 3.8-5.2MB.

[0271] 1.5.2) Sliding Window and Complexity

[0272] Continuous signal press s、 Sliding window; Total computational cost:

[0273] ,

[0274] Single-window latency is 22-25ms on mid-range mobile SoCs, with peak memory usage of <20MB.

[0275] 1.5.3) Safety Gating and Output

[0276] Gating threshold: or Proceed to "Manual Review / Re-collection"; Quality Q0 results in direct rejection. Output includes: multi-task labels and confidence levels, event boundaries, Visualize (overlaid on the time-frequency graph) and retain audit logs (timestamps, device ID hashes).

[0277] 1.6) Installation, Operation and Usage Steps

[0278] Installation: Install the app on your mobile device, pair it with the stethoscope via Bluetooth for the first time, and complete the device response self-calibration (play the calibration tone or load the response curve corresponding to the device ID); (optional) bind the ECG module.

[0279] operate:

[0280] Select participant information and agree to the privacy policy;

[0281] Enter the "Collection Wizard" and place the samples at the four auscultation points in sequence as prompted on the screen, collecting data for 15 seconds at each point;

[0282] The quality bar is displayed in real time during the data acquisition process, and if Q0 is reached, a re-acquisition prompt will be automatically displayed.

[0283] Data collection is complete; the endpoint automatically infers and provides multi-task results. ;

[0284] Doctors can "confirm / revise" on the interface, and the revisions will be written to the local database for subsequent incremental updates. Uses: for rapid screening and triage during outpatient / physical examinations; for initial assessment and transport recommendations in pre-hospital emergency care; for long-term monitoring and synchronization with the HIS / follow-up system during home follow-ups.

[0285] 1.7) Implementation Results

[0286] On three-center data (5 stethoscopes, stratified by patient and center), the baseline of single-modal PCG and fine-tuning by transfer learning were compared:

[0287] Overall performance (macro average): AUROC: baseline 0.89 → migration 0.90 → this application 0.93; Macro-F1: baseline 0.79 → migration 0.81 → this application 0.84.

[0288] Low SNR (5-8dB): FPR@95%TPR: Baseline 18.5% → Migration 16.2% → This application 12.1%; Event boundary error (ms): Baseline 46.7 → Migration 42.5 → This application 35.9.

[0289] Cross-device generalization (no domain seen): Macro-F1: Baseline 0.76 → Migration 0.78 → This application 0.83.

[0290] Edge-side efficiency: Multi-model parallelism ≈ 6.5M parameters → This application's unified multi-task efficiency is 1.2M; INT8 inference 22-25ms / window, peak memory <20MB.

[0291] Security Gating: Enabled After quality grading, the false positive rate decreased by about 30%, and the retest trigger rate was approximately 8-10%, which is in line with clinical procedures.

[0292] Example 2: Multicenter Heart Sound Screening and Diagnostic System Based on Federated Distillation

[0293] 2.1) Application Scenarios and System Composition

[0294] This embodiment focuses on regional community hospitals and primary healthcare institutions to illustrate how this application can be used in multi-center, low-data-volume, and privacy-restricted environments.

[0295] Application scenarios: Multiple medical institutions have limited heart sound data, diverse equipment models, complex noise environments, and cannot directly share raw data.

[0296] Hardware configuration: The branch center is equipped with a portable electronic stethoscope (sampling rate 2-4kHz, 16-bit) and optional single-lead ECG (≥250Hz); the local terminal is a regular PC or embedded ARM SoC. The central node deploys a secure aggregation server (GPU / TPU optional).

[0297] Software architecture: Each branch center trains a local teacher model. Soft labels and intermediate representations are generated; these are then encrypted and uploaded to the central node, where they are aggregated and distilled into the student model. Then it is sent back to each branch center for deployment.

[0298] 2.2) Data Acquisition and Governance

[0299] Each branch center has approximately 200-300 patients' data, with each heart sound sample lasting 10-15 seconds. Some locations also have ECG data.

[0300] Quality classification: Based on the spectral energy ratio, artifact detection, and effective cycle count, the signal is classified into Q2 (high quality), Q1 (usable), and Q0 (low quality). Q0 is directly eliminated during training and inference.

[0301] Timing alignment: If ECG exists, the S1 of PCG is aligned with the R peak; if there is no ECG, PCG adaptive peak detection combined with dynamic time warping (DTW) is used to ensure periodic consistency across devices and individuals.

[0302] Device adaptation: Frequency response differences between different stethoscopes are compensated through amplitude-frequency response correction; device IDs are converted into embedded vectors. It is introduced into the conditional normalization layer of the student model for cross-device adaptation.

[0303] Data augmentation: Temporal scaling, noise mixing (street noise, breathing sounds), SpecAugment, and hybrid augmentation are employed to improve the model's robustness to complex environments.

[0304] 2.3) Model Training and Federated Distillation Process

[0305] 2.3.1) Local Teacher Model Training

[0306] Each center trains a local teacher model. Input PCG±ECG, output soft label Feature representation and mask The optimization objective is:

[0307] ,

[0308] 2.3.2 Encrypted Transmission and Aggregation

[0309] Each center only uploads , and The statistical summary is provided, but the original PCG is not uploaded. Homomorphic encryption and differential privacy are used to ensure that "data does not leave the domain".

[0310] 2.3.3) Central distillation and student model update

[0311] The central node unifies and aggregates teacher knowledge to train student models. The loss function is:

[0312] ,

[0313] Among them, the comparative distillation term ensures consistency of characteristics across centers and across devices.

[0314] 2.3.4) Model Distribution and Update Mechanism

[0315] The central node periodically (e.g., quarterly) sends the distilled and updated student model parameters back to each branch center to achieve continuous iteration.

[0316] 2.4) Deployment and Inference Process

[0317] The student model parameter size is controlled at around 1.1M, and after quantization and compression, it is <5MB, with a single-window inference latency of <25ms on the edge. Inference flow:

[0318] Procedure: Doctor collects 10-15 seconds of heart sounds → Local terminal runs → Outputs heart murmur detection, valvular disease classification, event boundaries, quality grade, and cardiac function prediction.

[0319] Quality self-check: If the output is Q0, prompt for resampling; Q1, add a low confidence flag and proceed to review; Q2, normal output.

[0320] Security gating: When or At that time, manual review is triggered.

[0321] 2.5) Implementation Results

[0322] In a joint experiment involving three community hospitals and one county-level hospital:

[0323] Overall accuracy: The traditional single-center trained student model has a Macro-F1 score of ≈0.77, while the federated distillation scheme in this application improves it to 0.83.

[0324] Cross-device generalization: On unseen devices, the student model in this application achieves a Macro-F1 improvement of approximately 6%, outperforming direct transfer learning methods.

[0325] The low-quality data tolerance mechanism reduces the false positive rate by 25% under Q1 samples, and the Q0 rejection mechanism reduces misjudgments.

[0326] Real-time performance: Inference latency of <20ms / window for ordinary PCs and <30ms / window for ARM SoCs, meeting the needs of pre-hospital ambulance and outpatient applications.

[0327] The above is an introduction to the configuration method of the heart sound processing model based on multimodal knowledge distillation provided in this application. In order to facilitate better implementation of the configuration method of the heart sound processing model based on multimodal knowledge distillation provided in this application, this application also provides a configuration device of the heart sound processing model based on multimodal knowledge distillation from the perspective of functional modules.

[0328] See Figure 2 , Figure 2 This is a schematic diagram of a configuration device for the heart sound processing model based on multimodal knowledge distillation according to this application. In this application, the configuration device 200 for the heart sound processing model based on multimodal knowledge distillation may specifically include the following structure:

[0329] Configuration unit 201 is used to configure a first heart sound processing model as a teacher model and a second heart sound processing model as a student model.

[0330] Training unit 202 is used to guide the second heart sound processing model to learn using the first heart sound processing model based on the knowledge distillation mechanism during the model training phase. The multimodal inputs of the first heart sound processing model include heart sound modality, electrocardiogram modality, echocardiogram modality and electronic medical record modality, and the single-modal inputs of the second heart sound processing model include heart sound modality.

[0331] Deployment unit 203 is used to deploy a second heart sound processing model, which has been learned through a knowledge distillation mechanism, on a target device used to perform the model's actual application tasks. The target device includes at least one of a portable stethoscope, a mobile terminal, and a wearable device.

[0332] As an exemplary embodiment, the distillation path involved in the knowledge distillation mechanism specifically includes logit distillation of the output layer, feature distillation of the intermediate layer, time-frequency mask distillation in the time-frequency domain, comparative distillation of samples, and uncertainty adjustment of confidence.

[0333] As another exemplary embodiment, the entire model training phase includes a teacher model training phase, a student model pre-training phase, and a joint distillation training phase;

[0334] Loss function during the teacher model training phase Represented as:

[0335] ,

[0336] in, Represents the standard cross-entropy loss function. Indicates the main task label. Indicates the weight of auxiliary tasks. This indicates auxiliary labels including event segmentation and signal quality;

[0337] Loss function during the pre-training phase of the student model Represented as:

[0338] ,

[0339] in, This represents the predicted distribution output by the student model. , This represents the Softmax function. Represents the student's logit vector. The distillation temperature. ;

[0340] Loss function during the joint distillation training phase Represented as:

[0341] ,

[0342] in, For dynamically adjusted different weighting coefficients, This represents the logit distillation loss. Indicates characteristic distillation loss, This indicates the time-frequency mask distillation loss. This indicates the difference in distillation losses. Represents an uncertain regularization term.

[0343] , ,

[0344] in, This represents the temperature probability distribution output by the teacher model. and Kullback-Leibler divergence, Represents the teacher's logit vector.

[0345] ,

[0346] in, Represents the square of the L2 norm. This represents the intermediate representation of the teacher model. Represents a linear mapping function. This represents the layer representation of the student model.

[0347] ,

[0348] in, This represents the mask predicted by the teacher model. This represents the mask predicted by the student model. for and Dice loss, express and Standard cross-entropy loss,

[0349]

[0350] in, Indicates student embedding, Indicates teacher embedding, Indicates inner product similarity. This represents the temperature parameter.

[0351] As another exemplary embodiment, the optimization strategy configured during training includes the following processing:

[0352] Use the Adamw optimizer;

[0353] Learning rate Use cosine annealing for scheduling;

[0354] Dynamic weight adjustment employs the GradNorm method based on gradient norm;

[0355] Two regularization methods, Dropout and SpecAugment, are introduced.

[0356] Introduce an early stop strategy to monitor the multi-task performance of the validation set;

[0357] Introduce mixed-precision training involving FP16 or INT8 low-precision formats.

[0358] As another exemplary embodiment, the deployment strategy involved when the second heart sound processing model is deployed on the target device includes the following processing:

[0359] Using a sliding window inference mechanism, the continuous heart sound signal waveform is segmented into segments with a length of [missing information]. The windows, spaced apart by steps The sliding operation, with each window independently performing time-frequency transformation and model forward inference, has an approximate complexity of:

[0360] ,

[0361] in, The output represents an approximate complexity. Indicates the total length of the input sequence. This indicates the computational complexity of a single-window model.

[0362] The deployment strategy for the second heart sound processing model when deployed on the target device also includes the following processing:

[0363] Quantization and pruning of the second heart sound processing model;

[0364] The quantization strategy employs symmetric INT8 fixed-point quantization to ensure that convolutional and fully connected layers operate in the integer domain;

[0365] The pruning strategy structurally reduces redundant channels, and the corresponding constraints are:

[0366] ,

[0367] in, This indicates the accuracy of the model after pruning. This represents the model's accuracy before pruning. This indicates the allowable accuracy loss threshold.

[0368] As another exemplary embodiment, the second heart sound processing model, after deployment on the target device, introduces a signal quality self-checking mechanism, including the following processing:

[0369] The input fragment is first processed by a quality head to predict its quality level. , , The signal quality levels are represented from low to high;

[0370] when When this happens, the system triggers a secondary acquisition or a delayed judgment to prevent noise fragments from directly entering the diagnostic process;

[0371] when At this time, the output results are marked with a low confidence level, indicating a need for clinical review;

[0372] when At that time, the result directly enters the task layer.

[0373] As yet another exemplary embodiment, the second heart sound processing model will use the task space Multi-task output heads are uniformly mapped to the same backbone network to achieve resource sharing and information complementarity;

[0374] The second heart sound processing model performs five prediction tasks under single-modal input conditions, specifically involving heart murmur detection, valvular disease classification, heart sound event detection, signal quality assessment, and heart function level prediction.

[0375] In the second heart sound processing model, all data is stored locally in a single-center application scenario, and the original heart sounds are not uploaded. In a multi-center collaborative scenario, a federated distillation mode is adopted.

[0376] In inference operations, the prediction mask is displayed in the visualization interface that outputs the model's prediction results. The event boundaries of the heart sound event detection are superimposed on the corresponding time-frequency map of the model input heart sound signal, and the event boundaries of the output heart sound event detection are displayed corresponding to the mask.

[0377] As another exemplary embodiment, the device further includes an application unit 204 for:

[0378] Heart sound analysis and processing are performed based on the second heart sound processing model.

[0379] This application also provides a processing device from a hardware architecture perspective. As mentioned earlier, in practice, a processing device may exist as a device cluster. In this case, each device in the device cluster can also be referred to as a processing device. See [reference needed]. Figure 3 , Figure 3This diagram illustrates a structural schematic of the processing device of this application. Specifically, the processing device may include a processor 301, a memory 302, and an input / output device 303. The processor 301 executes the computer program stored in the memory 302 to implement, for example... Figure 1 The steps of the configuration method for the heart sound processing model based on multimodal knowledge distillation in the corresponding embodiment; or, when the processor 301 executes the computer program stored in the memory 302, it implements as follows: Figure 2 Corresponding to the functions of each unit in the embodiment, the memory 302 is used to store the functions executed by the processor 301 as described above. Figure 1 The computer program required for configuring the heart sound processing model based on multimodal knowledge distillation in the corresponding embodiment.

[0380] For example, a computer program may be divided into one or more modules / units, one or more of which are stored in memory 302 and executed by processor 301 to complete this application. One or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in a computer device.

[0381] The processing device may include, but is not limited to, processor 301, memory 302, and input / output device 303. Those skilled in the art will understand that the illustrations are merely examples of the processing device and do not constitute a limitation on the processing device. It may include more or fewer components than illustrated, or combine certain components, or different components. For example, the processing device may also include network access devices, buses, etc., and processor 301, memory 302, input / output device 303, etc., are connected via a bus.

[0382] Processor 301 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the processing device, connecting various parts of the device through various interfaces and lines.

[0383] The memory 302 can be used to store computer programs and / or modules. The processor 301 implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory 302 and by calling data stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the processing device, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0384] When processor 301 executes a computer program stored in memory 302, it can specifically perform the following functions:

[0385] Configure a first heart sound processing model as the teacher model and a second heart sound processing model as the student model;

[0386] During the model training phase, the first heart sound processing model is used to guide the second heart sound processing model to learn based on the knowledge distillation mechanism. The multimodal inputs of the first heart sound processing model include heart sound modality, electrocardiogram modality, echocardiogram modality and electronic medical record modality, while the single-modal inputs of the second heart sound processing model include heart sound modality.

[0387] The second heart sound processing model, which learns through a knowledge distillation mechanism, is deployed on a target device used to perform the model's actual application tasks. The target device includes at least one of a portable stethoscope, a mobile terminal, and a wearable device.

[0388] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the configuration device, processing equipment, and corresponding units of the heart sound processing model based on multimodal knowledge distillation described above can be found in the following reference: Figure 1 The configuration method of the heart sound processing model based on multimodal knowledge distillation in the corresponding embodiment will not be described in detail here.

[0389] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0390] Therefore, this application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the present application. Figure 1 The steps of the configuration method for the heart sound processing model based on multimodal knowledge distillation in the corresponding embodiment can be found in the following example. Figure 1 The configuration method of the heart sound processing model based on multimodal knowledge distillation in the corresponding embodiment will not be repeated here.

[0391] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0392] Because of the instructions stored in the computer-readable storage medium, the present application can be executed as described above. Figure 1 The steps of the configuration method for the heart sound processing model based on multimodal knowledge distillation in the corresponding embodiment can therefore achieve the results of this application. Figure 1 The beneficial effects that the configuration method of the heart sound processing model based on multimodal knowledge distillation can achieve in the corresponding embodiment are detailed in the preceding description and will not be repeated here.

[0393] The configuration method, apparatus, processing device, and computer-readable storage medium of the heart sound processing model based on multimodal knowledge distillation provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A configuration method for a heart sound processing model based on multimodal knowledge distillation, characterized in that, The method includes: Configure a first heart sound processing model as the teacher model and a second heart sound processing model as the student model; During the model training phase, the first heart sound processing model is used to guide the second heart sound processing model to learn based on a knowledge distillation mechanism. The multimodal input of the first heart sound processing model includes heart sound modality, electrocardiogram modality, echocardiogram modality and electronic medical record modality, and the single-modal input of the second heart sound processing model includes the heart sound modality. The second heart sound processing model, which has been learned through the knowledge distillation mechanism, is deployed on a target device for performing actual application tasks of the model. The target device includes at least one of a portable stethoscope, a mobile terminal, and a wearable device. The distillation paths involved in the knowledge distillation mechanism specifically include logit distillation of the output layer, feature distillation of the intermediate layer, time-frequency mask distillation in the time-frequency domain, comparative distillation of samples, and uncertainty adjustment of confidence. The entire model training phase includes a teacher model training phase, a student model pre-training phase, and a joint distillation training phase. The loss function during the teacher model training phase Represented as: , in, Represents the standard cross-entropy loss function. Indicates the main task label. This indicates the main task label prediction result. Indicates the weight of auxiliary tasks. This indicates auxiliary labels including event segmentation and signal quality. This indicates the prediction result of the auxiliary label; The loss function during the pre-training phase of the student model Represented as: , in, This represents the predicted distribution output by the student model. , This represents the Softmax function. Represents the student's logit vector. The distillation temperature. ; The loss function during the joint distillation training phase Represented as: , in, For dynamically adjusted different weighting coefficients, This represents the logit distillation loss. Indicates characteristic distillation loss, This indicates the time-frequency mask distillation loss. This indicates the difference in distillation losses. Represents an uncertain regularization term. , , in, This represents the temperature probability distribution output by the teacher model. and Kullback-Leibler divergence, Represents the teacher's logit vector. , in, Represents the square of the L2 norm. This represents the intermediate representation of the teacher model. Represents a linear mapping function. This represents the layer representation of the student model. , in, This represents the mask predicted by the teacher model. This represents the mask predicted by the student model. for and Dice loss, express and Standard cross-entropy loss, in, Indicates student embedding, Indicates teacher embedding, Indicates inner product similarity. This represents the temperature parameter.

2. The method according to claim 1, characterized in that, The optimization strategies configured during training include the following processing: Use the Adamw optimizer; Learning rate Use cosine annealing for scheduling; Dynamic weight adjustment employs the GradNorm method based on gradient norm; Two regularization methods, Dropout and SpecAugment, are introduced. Introduce an early stop strategy to monitor the multi-task performance of the validation set; Introduce mixed-precision training involving FP16 or INT8 low-precision formats.

3. The method according to claim 1, characterized in that, The deployment strategy for the second heart sound processing model when deployed on the target device includes the following processing: Using a sliding window inference mechanism, the continuous heart sound signal waveform is segmented into segments with a length of [missing information]. The windows, the windows being spaced in steps The sliding process involves each window independently undergoing time-frequency transformation and model forward inference, with a corresponding complexity approximately as follows: , in, The output represents an approximate complexity. Indicates the total length of the input sequence. This indicates the computational complexity of a single-window model. The deployment strategy for the second heart sound processing model when deployed on the target device also includes the following processing: The second heart sound processing model is quantized and pruned; The quantization strategy employs symmetric INT8 fixed-point quantization to ensure that convolutional and fully connected layers operate in the integer domain; The pruning strategy structurally reduces redundant channels, and the corresponding constraints are: , in, This indicates the accuracy of the model after pruning. This represents the model's accuracy before pruning. This indicates the allowable accuracy loss threshold.

4. The method according to claim 1, characterized in that, The second heart sound processing model completes its inference process after deployment on the target device, and introduces a signal quality self-checking mechanism, including the following processing: The input fragment is first processed by a quality head to predict its quality level. , , The signal quality levels are represented from low to high; when When this happens, the system triggers a secondary acquisition or a delayed judgment to prevent noise fragments from directly entering the diagnostic process; when At this time, the output results are marked with a low confidence level, indicating a need for clinical review; when At that time, the result directly enters the task layer.

5. The method according to claim 4, characterized in that, The second heart sound processing model will have a task space. Multi-task output heads are uniformly mapped to the same backbone network to achieve resource sharing and information complementarity; The second heart sound processing model performs five prediction tasks under the single-modal input operating conditions, specifically involving heart murmur detection, valvular disease classification, heart sound event detection, signal quality assessment, and heart function level prediction. In the second heart sound processing model, all data is stored locally in a single-center application scenario, and the original heart sounds are not uploaded. In a multi-center collaborative scenario, a federated distillation mode is adopted. In inference operations, the prediction mask is displayed in the visualization interface that outputs the model's prediction results. The event boundaries of the detected heart sound events are superimposed on the corresponding time-frequency map of the model input heart sound signal, and the event boundaries of the heart sound event are displayed corresponding to the mask.

6. A configuration device for a heart sound processing model based on multimodal knowledge distillation, characterized in that, The device includes: A configuration unit is used to configure the first heart sound processing model as a teacher model and the second heart sound processing model as a student model. The training unit is used to guide the second heart sound processing model to learn using the first heart sound processing model based on a knowledge distillation mechanism during the model training phase. The multimodal inputs of the first heart sound processing model include heart sound modality, electrocardiogram modality, echocardiogram modality, and electronic medical record modality, and the single-modal inputs of the second heart sound processing model include the heart sound modality. The deployment unit is used to deploy the second heart sound processing model, which has been learned through the knowledge distillation mechanism, on a target device for performing the actual application task of the model, wherein the target device includes at least one of a portable stethoscope, a mobile terminal, and a wearable device. The distillation paths involved in the knowledge distillation mechanism specifically include logit distillation of the output layer, feature distillation of the intermediate layer, time-frequency mask distillation in the time-frequency domain, comparative distillation of samples, and uncertainty adjustment of confidence. The entire model training phase includes a teacher model training phase, a student model pre-training phase, and a joint distillation training phase. The loss function during the teacher model training phase Represented as: , in, Represents the standard cross-entropy loss function. Indicates the main task label. This indicates the main task label prediction result. Indicates the weight of auxiliary tasks. This indicates auxiliary labels including event segmentation and signal quality. This indicates the prediction result of the auxiliary label; The loss function during the pre-training phase of the student model Represented as: , in, This represents the predicted distribution output by the student model. , This represents the Softmax function. Represents the student's logit vector. The distillation temperature. ; The loss function during the joint distillation training phase Represented as: , in, For dynamically adjusted different weighting coefficients, This represents the logit distillation loss. Indicates characteristic distillation loss, This indicates the time-frequency mask distillation loss. This indicates the difference in distillation losses. Represents an uncertain regularization term. , , in, This represents the temperature probability distribution output by the teacher model. and Kullback-Leibler divergence, Represents the teacher's logit vector. , in, Represents the square of the L2 norm. This represents the intermediate representation of the teacher model. Represents a linear mapping function. This represents the layer representation of the student model. , in, This represents the mask predicted by the teacher model. This represents the mask predicted by the student model. for and Dice loss, express and Standard cross-entropy loss, in, Indicates student embedding, Indicates teacher embedding, Indicates inner product similarity. This represents the temperature parameter.

7. A processing device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the method as described in any one of claims 1 to 5 when it invokes the computer program in the memory.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Knowledge distillation-based acoustic-electric dual-mode congenital heart disease prediction device

    CN117017310A

  • Heart failure auxiliary diagnosis and treatment knowledge distillation method and system

    CN121260422A