Multi-view cross-domain heart sound segmentation method, device and system and storage medium

By employing a multi-view cross-domain heart sound segmentation method, utilizing an improved wavelet gMLP module and an uncertainty estimation module, combined with an adversarial time-frequency modulation decoupling mechanism, a multi-view cross-domain adaptive heart sound segmentation model is constructed. This addresses the shortcomings of existing technologies in multi-scale feature perception, noise robustness, and cross-domain adaptability, achieving higher accuracy and stability, and enhancing the model's application capability in complex environments.

CN121148432APending Publication Date: 2025-12-16GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511360522.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing heart sound segmentation technologies have limitations in multi-scale feature perception, noise robustness, and cross-domain adaptability, resulting in insufficient accuracy and stability in real-world complex environments, making them difficult to apply effectively to different individuals, devices, and environments.

Method used

A multi-view cross-domain heart sound segmentation method is adopted. By using a parallel dual-UNet architecture and an improved wavelet gMLP module, combined with an uncertainty estimation module and an adversarial time-frequency modulation decoupling mechanism, a multi-view cross-domain adaptive heart sound segmentation model is constructed to achieve decoupling of time-frequency features and knowledge transfer, thereby improving the model's segmentation performance in noisy environments and cross-domain scenarios.

Benefits of technology

It significantly improves the accuracy and stability of the heart sound segmentation model in real-world noise environments and across devices and individuals, enhances the model's cross-domain adaptability and generalization performance, and can better identify the boundary between S1 and S2 heart sounds, overcoming the shortcomings of traditional methods in cross-domain adaptation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121148432A_ABST
    Figure CN121148432A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-view cross-domain heart sound segmentation method, device and system and a storage medium. The multi-view cross-domain heart sound segmentation method comprises the steps that S1, heart sound data are acquired and preprocessed; s2, constructing a multi-view cross-domain adaptive heart sound segmentation model according to the preprocessed heart sound data; and S3, inputting the heart sound data to be processed into the multi-view cross-domain adaptive heart sound segmentation model for heart sound segmentation. By the adoption of the technical scheme, the accuracy and stability of heart sound segmentation in a real noise environment and cross-device and cross-individual application are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information processing technology, and particularly relates to a multi-view cross-domain heart sound segmentation method, device, system, and storage medium. Background Technology

[0002] Cardiovascular diseases (CVDs) are one of the leading causes of death worldwide. Heart sound auscultation, as a non-invasive and low-cost preliminary screening method, has always played a vital role in cardiovascular function assessment. In recent years, with the integration of artificial intelligence and Internet of Things (IoT) technologies, intelligent auscultation devices have enabled long-term, continuous heart sound acquisition in home environments, providing new possibilities for early screening and remote monitoring of cardiovascular diseases.

[0003] Heart sound signals typically consist of four basic components: the first heart sound (S1), the systolic sound, the second heart sound (S2), and the diastolic sound. The goal of automatic heart sound segmentation is to accurately identify the start and end points of S1 and S2 from the raw signal; this is a crucial prerequisite in the entire analysis process. The accuracy of the segmentation results directly affects whether subsequent tasks such as heart sound classification, parameter calculation, and disease diagnosis can be reliably completed.

[0004] From a signal characteristics perspective, heart sounds contain both transient, non-stationary events like S1 and S2, requiring the model to achieve high-precision temporal localization; and they also exhibit a certain degree of quasi-periodicity, relying on the model's understanding of the temporal relationships between events. Furthermore, in real-world acquisition environments, heart sounds are highly susceptible to various noise interferences, including ambient background noise, sensor disturbances, and individual physiological differences, leading to significant fluctuations in signal quality. Therefore, an ideal heart sound segmentation model must first meet two core requirements: first, it must possess multi-scale feature perception capabilities, simultaneously characterizing local transient features and global periodic structures; second, it must have strong noise resistance and robustness, maintaining stable performance under different noise environments.

[0005] However, simply meeting the above requirements is insufficient to support the widespread application of the model in real-world scenarios. Heart sound segmentation models also face another key challenge: insufficient cross-domain adaptability. That is, a model that performs well on one dataset may experience a significant performance drop in new scenarios with different data distributions. These inter-domain differences mainly originate from cross-individual, cross-device, cross-environment, and even cross-species differences. These differences essentially reflect a mismatch between the probability distribution of the training data and the actual application scenario.

[0006] Most current heart sound segmentation methods are still built on a single dataset or implicit identical distribution assumptions, focusing primarily on improving the model's performance within a single domain, without fully considering the domain shift challenges in real-world deployments. This means that while existing methods may achieve good performance on specific datasets, they severely limit the widespread adoption and application of intelligent heart sound systems in the real world.

[0007] To address the cross-domain adaptation problem, domain adaptation methods have been effectively applied in the broader field of time-series signal processing to mitigate distributional disparities. Their core principle is to achieve knowledge transfer from the source domain to the target domain by learning domain-invariant feature representations. On the other hand, multi-view learning has been introduced into cardiac sound analysis, aiming to integrate features from multiple perspectives, such as the time domain, frequency domain, and time-frequency domain, to enhance the model's discriminative ability. However, existing works are typically limited to applications within a single domain and fail to organically combine multi-view representations with cross-domain adaptation mechanisms. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a multi-view cross-domain heart sound segmentation method, device, system, and storage medium, which aims to effectively overcome the limitations of existing heart sound segmentation technology in terms of multi-scale feature perception, noise robustness, and cross-domain adaptability, and improve the accuracy, stability, and generalization performance of heart sound segmentation in real and complex environments.

[0009] To achieve the above objectives, the present invention adopts the following technical solution:

[0010] A multi-view cross-domain heart sound segmentation method includes:

[0011] Step S1: Acquire heart sound data and perform preprocessing;

[0012] Step S2: Construct a multi-view cross-domain adaptive heart sound segmentation model based on the preprocessed heart sound data;

[0013] Step S3: Input the heart sound data to be processed into the multi-view cross-domain adaptive heart sound segmentation model for heart sound segmentation.

[0014] Preferably, step S1 includes:

[0015] Denoise the source and target domain heart sound datasets;

[0016] Temporal and frequency domain features were extracted from the denoised source and target domain heart sound data.

[0017] As a preferred approach, a parallel dual-UNet architecture is used as the backbone network during the training phase of the multi-view cross-domain adaptive heart sound segmentation model. This architecture includes two branches that process time-domain and frequency-domain signals respectively, with the frequency-domain branch serving as an auxiliary path. An improved wavelet gMLP module is embedded at the central bottleneck positions of the time-domain and frequency-domain branches, located between the encoder and decoder. The improved wavelet gMLP module introduces a Morlet wavelet driving mechanism in the spatially gated path. Based on time-frequency domain mutual learning guided by the uncertainty estimation module, the highly discriminative frequency-domain branch is used as the teaching branch in the source domain. The teacher guides the learning of the time-domain branch; in the target domain, the intensity of knowledge transfer is dynamically adjusted through the uncertainty of time-domain prediction; an adversarial time-frequency modulation decoupling mechanism is adopted, and by constructing a dual-path adversarial learning framework, the time-frequency related features are explicitly decoupled into two functionally complementary components: one is a domain-invariant related component, which is used to promote cross-domain knowledge transfer; the other is a domain-specific related component, which is used to retain the discriminative information unique to the source domain and the target domain. During the training process, standard adversarial training is used for the domain-invariant component to strengthen cross-domain consistency, while reverse adversarial training is used for the domain-specific component to enhance its domain discrimination ability.

[0018] As a preferred option, the total loss function L of the multi-view cross-domain adaptive heart sound segmentation model is... total for:

[0019]

[0020] in, For the standard cross-entropy loss of the time-domain branch, Standard cross-entropy loss of frequency domain branch, L s For source domain distillation loss, L t For the target domain distillation loss, L D To counteract the decoupling loss, β1, β2, and β3 are all hyperparameters used to balance the importance of each loss term in the overall optimization objective.

[0021] The present invention also provides a multi-view cross-domain heart sound segmentation device, comprising:

[0022] The first processing unit is used to acquire heart sound data and perform preprocessing.

[0023] The second processing unit is used to construct a multi-view cross-domain adaptive heart sound segmentation model based on the preprocessed heart sound data.

[0024] The third processing unit is used to input the heart sound data to be processed into the multi-view cross-domain adaptive heart sound segmentation model for heart sound segmentation.

[0025] Preferably, the first processing unit includes:

[0026] The first processing component is used to denoise the source and target domain heart sound datasets;

[0027] The second processing component is used to extract time-domain and frequency-domain features from the denoised source and target domain heart sound data.

[0028] As a preferred approach, a parallel dual-UNet architecture is used as the backbone network during the training phase of the multi-view cross-domain adaptive heart sound segmentation model. This architecture includes two branches that process time-domain and frequency-domain signals respectively, with the frequency-domain branch serving as an auxiliary path. An improved wavelet gMLP module is embedded at the central bottleneck positions of the time-domain and frequency-domain branches, located between the encoder and decoder. The improved wavelet gMLP module introduces a Morlet wavelet driving mechanism in the spatially gated path. Based on time-frequency domain mutual learning guided by the uncertainty estimation module, the highly discriminative frequency-domain branch is used as the teaching branch in the source domain. The teacher guides the learning of the time-domain branch; in the target domain, the intensity of knowledge transfer is dynamically adjusted through the uncertainty of time-domain prediction; an adversarial time-frequency modulation decoupling mechanism is adopted, and by constructing a dual-path adversarial learning framework, the time-frequency related features are explicitly decoupled into two functionally complementary components: one is a domain-invariant related component, which is used to promote cross-domain knowledge transfer; the other is a domain-specific related component, which is used to retain the discriminative information unique to the source domain and the target domain. During the training process, standard adversarial training is used for the domain-invariant component to strengthen cross-domain consistency, while reverse adversarial training is used for the domain-specific component to enhance its domain discrimination ability.

[0029] As a preferred option, the total loss function L of the multi-view cross-domain adaptive heart sound segmentation model is... total for:

[0030]

[0031] in, For the standard cross-entropy loss of the time-domain branch, Standard cross-entropy loss of frequency domain branch, L s For source domain distillation loss, L t For the target domain distillation loss, L D To counteract the decoupling loss, β1, β2, and β3 are all hyperparameters used to balance the importance of each loss term in the overall optimization objective.

[0032] The present invention also provides a multi-view cross-domain heart sound segmentation system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a multi-view cross-domain heart sound segmentation method when executed by the processor.

[0033] The present invention also provides a storage medium storing a computer program that executes a multi-view cross-domain heart sound segmentation method when running.

[0034] This invention employs a segmentation model that can effectively integrate multi-scale information and has cross-domain adaptability to improve the accuracy and stability of heart sound segmentation in real noise environments and in cross-device and cross-individual applications. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0036] Figure 1 This is a flowchart of the multi-view cross-domain heart sound segmentation method according to an embodiment of the present invention;

[0037] Figure 2 Data processing flowchart for cross-domain adaptive heart sound signal segmentation model;

[0038] Figure 3 Wavelet gMLP module structure diagram Detailed Implementation

[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] Example 1:

[0042] like Figure 1 , 2 As shown, this embodiment of the invention provides a multi-view cross-domain heart sound segmentation method, including:

[0043] Step S1: Acquire heart sound data and perform preprocessing;

[0044] Step S2: Construct a multi-view cross-domain adaptive heart sound segmentation model based on the preprocessed heart sound data;

[0045] Step S3: Input the heart sound data to be processed into the multi-view cross-domain adaptive heart sound segmentation model for heart sound segmentation.

[0046] As one implementation of an embodiment of the present invention, in step S1, in order to achieve effective modeling and unified input of cross-domain heart sound data, the source domain and target domain heart sound data are first collected and preprocessed, specifically including:

[0047] Step 1.1: Prepare source and target domain heart sound datasets

[0048] The source domain uses the PhysioNet / CinC Challenge 2016 dataset. This dataset contains 3,153 heart sound segments from 315 subjects, sampled at a frequency of 2,000 Hz, single-channel, and covering both normal and abnormal samples. Each heart sound segment is provided with beat-by-beat labels, accurately labeling the signal into four categories: S1, systole, S2, and diastole, effectively supporting fine-grained heart sound segmentation tasks.

[0049] The target domain uses the CirCorDigiScope PCG dataset. This dataset contains 3,163 heart sound segments from 942 subjects, sampled at a frequency of 2,000 Hz, single-channel, and exhibiting true noise characteristics. This dataset only provides segment-level labels, and the labeling quality is inconsistent. Therefore, in this invention, following the unsupervised domain adaptation setting, the labeling information from this dataset is not utilized.

[0050] Step 1.2: Noise Reduction

[0051] For the source and target domain heart sound data in step 1.1, a 25–400Hz bandpass filter is first used to remove low-frequency drift and high-frequency noise, preserving the main spectral components of the heart sounds. Then, a maximum absolute amplitude detection method based on a sliding window is introduced to identify and suppress transient spike interference: the original signal x = {x1, x2, ..., x...} is processed... N} Based on window length L = 0.5 × f s Divide into several sub-segments x (k) Calculate the maximum absolute amplitude for each segment:

[0052] MAA(k) = max(|x (k) |) (1-1)

[0053] Set the threshold to three times the median of the maximum absolute amplitude of all windows:

[0054] θ = 3 × median(MAA(k)) (1-2)

[0055] When MAA(k) > θ for a certain window, the peak position and the zero-crossing points before and after it are found within that window. After determining the peak interval, the signal value of the interval is replaced with a constant ∈ = 0.0001. This process is repeated until all windows satisfy MAA(k) ≤ θ, so as to effectively suppress the abnormal disturbances caused by the peaks. Finally, the signals of the two datasets are uniformly downsampled to 50Hz to facilitate unified modeling and calculation in the subsequent model training process.

[0056] Step 1.3: Extract envelope features

[0057] After completing step 1.2, various envelope features are further extracted in the time domain to characterize the energy variation of the heart sound signal over time. These include the Hilbert envelope, homogeneous envelope, power spectral density envelope, and wavelet envelope. First, the analytical form of the signal is obtained through Hilbert transform.

[0058]

[0059] Then its modulus is calculated as the envelope, that is

[0060]

[0061] Secondly, the homogeneous envelope employs a logarithmic domain linearization strategy. After taking the logarithm of the signal amplitude, it is smoothed by a low-pass filter and finally recovered exponentially, expressed as follows:

[0062] E Homo (t)=exp(LPF(log(|x(t)|+∈))) (1-5)

[0063] Furthermore, the power spectral density envelope is used to obtain the local spectrum through short-time Fourier transform, and the power within the specified frequency band is integrated to obtain...

[0064]

[0065] Finally, the wavelet envelope is used to calculate the time-frequency energy at different scales through continuous wavelet transform, and the squares of the wavelet coefficients are summed, as shown in the formula:

[0066]

[0067] The aforementioned multiple envelopes can provide complementary structural information. These envelopes are then normalized in magnitude and linearly scaled to the interval [-1, 1], serving as the time-domain input features of the model.

[0068]

[0069] Where min(x) and max(x) are the minimum and maximum values ​​of the signal, respectively. This normalization process eliminates individual differences and the influence of dimensions, ensuring that the signal amplitude distribution of different samples remains consistent.

[0070] Step 1.4: Extract frequency domain features

[0071] Frequency domain features are extracted from the heart sound signals processed in step 1.2. First, a Fast Fourier Transform (FFT) is performed on each signal segment x(n) to transform it from the time domain to the frequency domain, and its amplitude spectrum is calculated. The FFT calculation formula is as follows:

[0072]

[0073] Where N is the length of the signal segment. Then, the modulus is taken to obtain the amplitude spectrum M(k)=|X(k)|, and its amplitude is normalized and scaled to the interval [-1,1], which serves as the frequency domain input feature of the model.

[0074] As one implementation of an embodiment of the present invention, in step S2, the present invention constructs a multi-view cross-domain adaptive heart sound segmentation model. While ensuring excellent segmentation performance, this model effectively achieves domain adaptation from the source domain to the target domain, improving its generalization ability on unseen target data. The following will systematically elaborate on this from three aspects: network backbone structure, core module improvement, and cross-domain adaptation strategy.

[0075] Step 2.1: Backbone Network Structure

[0076] The UNet network has achieved remarkable results in heart sound segmentation tasks thanks to its unique encoder-decoder architecture and skip connection mechanism. Its superior performance primarily stems from its ability to collaboratively capture signal contextual information and spatial details. The standard UNet structure comprises five core parts: the encoder path progressively extracts multi-scale features through repeated convolutional layers and downsampling operations to capture global contextual information of the heart sound; the decoder path progressively restores spatial resolution through upsampling and convolutional operations to achieve accurate target localization; skip connections fuse high-resolution feature maps from each stage of the encoder with corresponding levels of the decoder, effectively compensating for details lost during downsampling; the bottleneck layer, located between the encoder and decoder, processes the highest level of feature representation; and the final output layer maps deep features to target categories through 1×1 convolutions, generating pixel-level segmentation results.

[0077] Based on the advantages of UNet mentioned above, this invention employs a parallel dual-UNet architecture as the backbone network during the training phase. This structure includes two branches that process time-domain and frequency-domain signals respectively, with the frequency-domain branch serving as an auxiliary path to enhance the representational power and robustness of the time-domain branch. Notably, only the time-domain branch is retained during the inference phase to ensure the model's efficiency and practicality while maintaining superior segmentation performance.

[0078] Step 2.2: Improvement and Embedding of the gMLP Module

[0079] The original gMLP structure introduces a spatial gating mechanism through channel separation, achieving global modeling and channel-level modulation of sequence features. Essentially, it can be viewed as an adaptive activation function, dynamically generating control gates within the network based on the input to adjust the activation levels at various locations and channels. However, traditional gating methods are typically based only on linear combinations in the time dimension, lacking the ability to model multi-scale periodicity, local oscillations, and abrupt structures, making it difficult to effectively characterize the complex time-frequency variations in heart sound signals.

[0080] To address this issue, this invention proposes using Morlet wavelet groups to structurally standardize and enhance the gMLP module. The overall module structure is as follows: Figure 3 As shown, Morlet wavelets inherently possess time-frequency localization capabilities and are a multi-scale basis function with good analytical representation. By introducing a set of controllable parameters from the Morlet wavelet basis into the gating path, the original linear gating can be transformed into a frequency-domain enhanced selective activation mechanism, thereby enabling the model to respond simultaneously to low-frequency rhythms and high-frequency noise. This mechanism not only makes the boundary detection of heart sound signals more sensitive but also gives the network stronger expressive and generalization capabilities when facing non-stationary signal structures.

[0081] Step 2.2.1 Morlet Wavelet Basis Generation Method

[0082] In this module, Morlet wavelets are used to construct a set of fixed function bases with time-frequency local sensing capabilities to regulate the spatial gating process in gMLP. Morlet wavelets are a type of composite wavelet whose core characteristic is the simultaneous ability to locate in the time domain and resolve frequencies. They are particularly suitable for capturing short-term periodic structures and boundary transition features, which is crucial for the accurate segmentation of S1 and S2 in heart sound signals.

[0083] The Morlet wavelet consists of a Gaussian envelope function and a sinusoidal carrier function, and its basic form is as follows:

[0084]

[0085] Here, t is the time variable, s is the wavelet width parameter, which controls the expansion range of the Gaussian window and determines the temporal locality of the wavelet; f0 is the center frequency, which controls the oscillation rate of the sine wave and determines the frequency selectivity of the wavelet. By adjusting the values ​​of f0 and s, a series of wavelet functions with different frequency responses can be generated.

[0086] To accommodate discrete sequence inputs of fixed length, this invention employs a symmetric central design when defining the time axis, namely:

[0087]

[0088] Where L is the length of the input sequence. Then, multiple wavelet functions with different frequencies are constructed and stacked along the frequency dimension to generate a fixed set of Morlet wavelet bases, denoted as . Where N represents the number of wavelet bases.

[0089] This set of wavelet bases remains constant throughout the model training process, serving as structural priors in subsequent gating modeling operations. Each wavelet base is treated as a predefined function activation template, used for pattern matching projection of intermediate features to extract the dynamic response of the sequence at that scale. This approach avoids the instability caused by channel features directly participating in learning in the temporal dimension. By leveraging the smoothness and bandpass properties of wavelets, spatial gating becomes more expressive and stable when handling periodicity and local transitions.

[0090] Step 2.2.2 Design of wavelet-driven spatial gating mechanism

[0091] The core innovation of this step lies in designing a gating mechanism that combines frequency domain feature perception by embedding the Morlet wavelet basis constructed in step 2.2.1 into the spatial gating path of the gMLP module. This enhances the model's ability to model different time scales and structural features in heart sound signals. This mechanism replaces the linear operation-based gating module in the original gMLP, thereby introducing more expressive frequency-selective modulation.

[0092] Input feature tensor It is first divided into two parts along the channel dimension. Retained as residual branches Entering the gating path, it is responsible for generating the activation modulation factor.

[0093] After entering the gated path, z2 is first normalized to improve numerical stability. Then, z2 for each channel is projected onto a predefined Morlet wavelet basis to extract the channel's response at multiple frequency scales. This process can be represented as:

[0094]

[0095] in The result obtained is Where N represents the number of wavelets and represents the frequency dimension. Unlike conventional gating operations, this invention introduces a learnable weighting matrix. The frequency domain response is remapped back to the original time dimension. This weighting operation essentially introduces a dynamic fusion of multiple frequency responses at each time step, and the mapping process is implemented as follows:

[0096]

[0097] Multiplication enables cross-dimensional feature combination. Furthermore, to increase model flexibility, a learnable bias term corresponding to the time step is introduced. The output results are gradually corrected:

[0098]

[0099] The final gated output is multiplied point-by-point by the wavelet-weighted activation factor through the residual branch to form the final spatial modulation output:

[0100]

[0101] This structure fully integrates the multi-scale time-frequency analytical capabilities of Morlet wavelets with the dynamic selectivity of the gating mechanism, enabling the model to possess stronger discriminative and generalization abilities when dealing with complex structures such as blurred boundaries, periodic transitions, and heart sound jumps. In particular, this gating module introduces a controllable structural inductive bias using parameterized wavelet responses as input signals. This allows gMLP to move beyond relying solely on empirically learned gating patterns and instead uses wavelet function priors to guide its activation selection in the frequency domain, thereby enhancing representation without significantly increasing the number of parameters.

[0102] Step 2.2.3 Wavelet gMLP Module Structure and Embedding Method

[0103] The input to the wavelet gMLP module is a shape of The module first normalizes the input features along the channel dimension to improve numerical stability and avoid gradient explosion. The normalized tensor then enters the first channel mapping layer, where a linear transformation increases the channel dimension from D to Di. ffn To enhance the ability to express features.

[0104] After dimensionality upscaling, the features undergo a nonlinear transformation using the GELU activation function, introducing nonlinearity into the model's representation. They then enter a wavelet-driven spatial gating module (see step 2.2.2). This module constructs multi-frequency modulation factors using a predefined Morlet wavelet basis and a learnable weighting mechanism, thereby selectively controlling the activation of different time steps and feature channels. The gating output is multiplied point-by-point with the channel branch-preserved path to complete dynamic modulation.

[0105] After gating, the tensor is remapped through another layer of channels, reducing the channel dimension from D. ffnThe output is returned to the original dimension D and residually connected to the initial input tensor to form the final output. While maintaining a compact structure, the entire module introduces a wavelet-driven selective activation path, enabling accurate modeling of non-stationary signals such as multi-scale periodic changes and structural jumps.

[0106] In the network architecture, wavelet gMLP modules are embedded at the central bottleneck positions of the time-domain and frequency-domain branches, specifically between the encoder and decoder. The feature tensors at this location possess the highest level of semantic abstraction and the lowest temporal resolution, converging the most discriminative information. This deep, abstract feature representation is highly suitable for modeling long-range global dependencies and performing high-level representation learning, maximizing its global perception and adaptive gating capabilities for complex dynamic structures between sequences.

[0107] Step 2.3: Time-Frequency Domain Mutual Learning Guided by the Uncertainty Estimation Module

[0108] In cross-domain adaptation, discriminability and transferability are two key criteria for evaluating the quality of feature representations. Discriminability requires features to be sufficiently class-distinguishing within the source domain, while transferability requires features to remain stable across domains. Existing research indicates that frequency domain features typically possess stronger discriminability within the source domain, which can be verified with labeled data in the source domain. However, whether time domain features are more transferable in cross-domain scenarios is difficult to assess directly due to the lack of labels in the target domain. Therefore, this paper proposes a time-frequency domain mutual learning mechanism guided by an uncertainty estimation module.

[0109] The core idea of ​​this mechanism is to borrow from the knowledge distillation paradigm, treating inter-domain learning as a prediction alignment problem. Specifically, the time-domain branch and the frequency-domain branch generate prediction results respectively, and knowledge transfer is achieved by minimizing the KL divergence between them. Its basic form is:

[0110]

[0111] Here, P and Q represent two prediction distributions. Since the KL divergence is asymmetric, the choice of direction directly affects the way knowledge is transferred. In the source domain, this embodiment of the invention treats the frequency domain model as the teacher branch, using its prediction distribution as a reference to guide the time domain model to learn more discriminative features, i.e., by minimizing...

[0112]

[0113] in and These represent the frequency domain prediction and time domain prediction of the source domain, respectively. This strategy enables the time domain branch to obtain more stable discrimination information from the frequency domain branch.

[0114] Within the target domain, the time-domain branch possesses potential transferability, but its reliability varies depending on the sample. Therefore, this embodiment of the invention utilizes prediction entropy to weight the distillation process: when the time-domain prediction... When the entropy is low, it indicates that the model has high confidence in the sample, and thus it serves as the teacher to guide the frequency domain branch learning; when the entropy is high, it indicates that there is significant uncertainty in the prediction, and the corresponding distillation intensity will automatically decrease. The target domain distillation loss is defined as:

[0115]

[0116] The weighting function w(x) is controlled by the prediction entropy:

[0117]

[0118] α is an adjustment factor used to control the effect of uncertainty on distillation intensity.

[0119] Step 2.4: Adversarial Time-Frequency Modulation Decoupling Mechanism

[0120] To achieve coordinated alignment between the time and frequency domains, this invention proposes an adversarial time-frequency modulation decoupling mechanism. The time-frequency correlation subspace not only inherits the statistical properties of the original time-domain and frequency-domain features, but more importantly, it can characterize the intrinsic correlation between these features. However, directly performing adversarial learning on the original time-frequency correlation subspace W carries certain risks: its discriminative features may be weakened or lost during forced domain alignment. To address this issue, the core innovation of this method lies in decoupling the subspace, separating it into two complementary components: one is the domain-invariant correlation W. inv The first is used to capture common information shared across domains, serving as the basis for domain alignment to facilitate knowledge transfer; the second is domain-specific relevance W. spe This decomposition method is used to preserve discriminative features related to the source or target domain, avoiding information loss caused by forced alignment. Through this decomposition, embodiments of the present invention can effectively maintain the discriminative power of features while promoting inter-domain alignment, thereby improving the generalization performance of the model on the target domain.

[0121] First, the original time-frequency correlation subspace is constructed by calculating the outer product of the time-domain feature f and the frequency-domain feature z and then performing average pooling dimensionality reduction.

[0122]

[0123] Subsequently, to separate domain-specific correlations, the method applies a convolutional layer to the original features to obtain f' and z', and constructs a modulated signal in the same manner.

[0124]

[0125] Domain-invariant correlation components are extracted by subtracting a domain-specific component scaled by a learnable parameter λ from the original correlation.

[0126] W inv =W-λ·ΔW (2-14)

[0127] During the adversarial training phase, this embodiment of the invention employs a shared domain discriminator g. D To process two components simultaneously. For the domain-invariant component W inv In this embodiment of the invention, standard adversarial training is performed to confuse the discriminator and enhance its cross-domain invariance; for a domain-specific component ΔW, reverse adversarial training is performed to strengthen its domain discriminability, enabling the discriminator to clearly identify its source domain. This strategy is achieved through the following combined loss function:

[0128]

[0129] The hyperparameter α is used to balance the importance of two tasks during training: domain invariant alignment and domain-specific reinforcement.

[0130] Step 2.5: Training Strategy

[0131] To ensure the model's discriminativeness in the source domain and its transferability in the target domain, the entire model is trained using a multi-objective loss function that integrates multiple objectives, including supervised learning, cross-distillation learning, and adversarial learning.

[0132] Step 2.5.1: Monitor Losses

[0133] On the source domain data, both the time-domain and frequency-domain branches are trained under supervised conditions using standard cross-entropy loss to ensure their basic class discrimination ability. The loss function is defined as follows:

[0134]

[0135] in, and Representing source domain samples x respectively s The predicted probability distributions in the time and frequency domain branches, y s These are the corresponding real tags.

[0136] Step 2.5.2: Multi-objective joint optimization

[0137] The model's overall objective function is composed of the basic supervision loss, the inter-distillation loss proposed in step 2.3, and the adversarial decoupling loss proposed in step 2.4. The overall loss function L... total Defined as the weighted sum of all loss terms:

[0138]

[0139] in, For the standard cross-entropy loss of the time-domain branch, Standard cross-entropy loss of frequency domain branch, L s For source domain distillation loss, L t For the target domain distillation loss, L D To counteract the decoupling loss, β1, β2, and β3 are all hyperparameters used to balance the importance of each loss term in the overall optimization objective.

[0140] As one embodiment of the present invention, the backbone network in step 2.1 can be a CNN-based network or an RNN-based network to complete the structural modeling and segmentation of the heart sound signal.

[0141] As one embodiment of the present invention, the uncertainty estimation module in step 2.3 can use uncertainty estimation mechanisms such as Monte Carlo dropout or deep ensemble to guide time-frequency domain mutual learning.

[0142] The innovation of this invention lies in:

[0143] (1) This invention proposes an improved gMLP module to enhance the modeling ability of neural networks for non-stationary and structurally complex time-series signals. Addressing the problem that traditional gMLP modules rely solely on spatial projection for their gating mechanisms, lacking time-frequency local perception capabilities, the core improvement of this invention lies in introducing a Morlet wavelet-driven mechanism into the spatial gating path. This mechanism uses a set of predefined Morlet wavelet bases as structural priors to perform time-frequency localization transformation on the input sequence, thereby effectively extracting frequency response patterns at multiple scales. Furthermore, an adaptive fusion of wavelet responses at different frequencies is performed using a learnable time-weighted matrix, dynamically generating activation control factors in the time dimension to achieve more expressive and discriminative gating operations. The improved module significantly enhances multi-scale feature extraction capabilities and noise resistance while introducing minimal parameter overhead. This module can be flexibly embedded into mainstream network architectures such as U-Net, and is particularly suitable as an enhancement component for semantic bottleneck layers, significantly improving the model's ability to perceive the instantaneous changes of S1 and S2 peaks.

[0144] (2) This invention proposes a time-frequency domain mutual learning mechanism based on uncertainty estimation. The core innovation of this mechanism lies in introducing prediction entropy as a proxy index for the transferability confidence of the time-domain branch, thereby achieving bidirectional knowledge distillation with sample adaptation. Specifically, in the source domain, the highly discriminative frequency domain branch is still used as the teacher to guide the learning of the time-domain branch; while in the target domain, the intensity of knowledge transfer is dynamically adjusted through the uncertainty of the time-domain prediction to avoid low-quality predictions misleading the frequency domain model.

[0145] (3) This invention proposes an adversarial time-frequency modulation decoupling mechanism to address the difficulty of distinguishing between domain-invariant and domain-specific features in the time-frequency correlated subspace using traditional methods. This mechanism constructs a dual-path adversarial learning framework, explicitly decoupling the time-frequency correlated features into two complementary components: a domain-invariant correlated component to facilitate cross-domain knowledge transfer, and a domain-specific correlated component to preserve the discriminative information unique to the source and target domains. During training, standard adversarial training is used for the domain-invariant component to enhance cross-domain consistency, while reverse adversarial training is used for the domain-specific component to enhance its domain discriminative ability. This dual-path collaborative optimization strategy fundamentally solves the inherent feature confusion problem in adversarial training within a mixed feature space using traditional methods, achieving more refined feature alignment.

[0146] Compared with the prior art, the present invention has the following technical effects:

[0147] (1) Traditional FFN or MLP-based methods often struggle to effectively capture the time-frequency structures with significant variations in heart sound signals, especially when dealing with transient spikes such as S1 and S2. This invention, by introducing Morlet wavelet basis sets into the spatially gated path, achieves time-frequency localization analysis and multi-scale frequency response extraction of the input signal, enabling the model to simultaneously perceive high-frequency transient components and low-frequency rhythmic patterns. This also enhances the model's nonlinear expressiveness and adaptability, allowing for more flexible fitting of complex time-varying patterns in heart sound signals, and significantly improving the representation and generalization ability of non-stationary heart sound structures.

[0148] (2) Compared with existing technologies, this invention overcomes the fundamental limitation that "the transferability of time-domain features cannot be directly verified." By introducing uncertainty estimation, it achieves bidirectional knowledge distillation for sample adaptation, significantly improving the accuracy and reliability of cross-domain knowledge transfer. This invention abandons the traditional approach of indirectly inferring transferability based on posterior performance and innovatively constructs a collaborative learning mechanism based on uncertainty perception. This mechanism effectively solves the model robustness problem caused by ignoring sample heterogeneity, significantly improving the adaptability and generalization performance in cross-domain scenarios, and providing a more reliable technical path for time series domain adaptation.

[0149] (3) Existing methods typically perform adversarial alignment in the time-frequency fusion feature space. Because they fail to effectively distinguish between domain-invariant and domain-specific features, feature confusion often occurs during cross-domain alignment, thus affecting transfer performance. This invention constructs a dual-path adversarial learning framework, explicitly decoupling time-frequency related features into two complementary components: domain-invariant and domain-specific. Forward alignment and reverse adversarial constraints are applied to each component, effectively enhancing cross-domain consistency while preserving the discriminative information unique to each domain. This structure fundamentally overcomes the shortcomings of traditional methods in adversarial training in a mixed feature space, achieving more accurate and stable cross-domain feature alignment and significantly improving the model's recognition performance and generalization ability in the target domain.

[0150] Example 2:

[0151] This invention also provides a multi-view cross-domain heart sound segmentation device, comprising:

[0152] The first processing unit is used to acquire heart sound data and perform preprocessing.

[0153] The second processing unit is used to construct a multi-view cross-domain adaptive heart sound segmentation model based on the preprocessed heart sound data.

[0154] The third processing unit is used to input the heart sound data to be processed into the multi-view cross-domain adaptive heart sound segmentation model for heart sound segmentation.

[0155] As one embodiment of the present invention, the first processing unit includes:

[0156] The first processing component is used to denoise the source and target domain heart sound datasets;

[0157] The second processing component is used to extract time-domain and frequency-domain features from the denoised source and target domain heart sound data.

[0158] In one embodiment of the present invention, a parallel dual-UNet architecture is used as the backbone network during the training phase of the multi-view cross-domain adaptive heart sound segmentation model. This architecture includes two branches that process time-domain and frequency-domain signals respectively, with the frequency-domain branch serving as an auxiliary path. An improved wavelet gMLP module is embedded at the central bottleneck positions of the time-domain and frequency-domain branches, i.e., between the encoder and decoder. The improved wavelet gMLP module introduces a Morlet wavelet driving mechanism in the spatially gated path. Based on time-frequency domain mutual learning guided by the uncertainty estimation module, in the source domain, the highly discriminative frequency domain... The branch acts as a teacher, guiding the learning of the time-domain branch. In the target domain, the intensity of knowledge transfer is dynamically adjusted based on the uncertainty of time-domain prediction. An adversarial time-frequency modulation decoupling mechanism is adopted, and a dual-path adversarial learning framework is constructed to explicitly decouple the time-frequency related features into two complementary components: one is a domain-invariant related component, used to promote cross-domain knowledge transfer; the other is a domain-specific related component, used to retain the discriminative information unique to the source and target domains. During training, standard adversarial training is used for the domain-invariant component to strengthen cross-domain consistency, while reverse adversarial training is used for the domain-specific component to enhance its domain discrimination ability.

[0159] As one embodiment of the present invention, the total loss function L of the multi-view cross-domain adaptive heart sound segmentation model is... total for:

[0160]

[0161] in, For the standard cross-entropy loss of the time-domain branch, Standard cross-entropy loss of frequency domain branch, L s For source domain distillation loss, L t For the target domain distillation loss, L D To counteract the decoupling loss, β1, β2, and β3 are all hyperparameters used to balance the importance of each loss term in the overall optimization objective.

[0162] Example 3:

[0163] This invention also provides a multi-view cross-domain heart sound segmentation system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a multi-view cross-domain heart sound segmentation method when executed by the processor.

[0164] Example 4:

[0165] This invention also provides a storage medium storing a computer program that executes a multi-view cross-domain heart sound segmentation method during runtime.

[0166] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A multi-view cross-domain heart sound segmentation method, characterized in that, include: Step S1: Acquire heart sound data and perform preprocessing; Step S2: Construct a multi-view cross-domain adaptive heart sound segmentation model based on the preprocessed heart sound data; Step S3: Input the heart sound data to be processed into the multi-view cross-domain adaptive heart sound segmentation model for heart sound segmentation.

2. The multi-view cross-domain heart sound segmentation method as described in claim 1, characterized in that, Step S1 includes: Denoise the source and target domain heart sound datasets; Temporal and frequency domain features were extracted from the denoised source and target domain heart sound data.

3. The multi-view cross-domain heart sound segmentation method as described in claim 2, characterized in that, In the training phase of the multi-view cross-domain adaptive heart sound segmentation model, a parallel dual-UNet architecture is used as the backbone network, which includes two branches that process time-domain signals and frequency-domain signals respectively, with the frequency-domain branch serving as an auxiliary path. An improved wavelet gMLP module is embedded in the central bottleneck position of the time-domain and frequency-domain branches respectively, that is, between the encoder and the decoder. The improved wavelet gMLP module introduces a Morlet wavelet driving mechanism in the spatial gating path. Based on the uncertainty estimation module-guided time-frequency domain mutual learning, in the source domain, the highly discriminative frequency domain branch serves as the teacher to guide the time domain branch learning; in the target domain, the intensity of knowledge transfer is dynamically adjusted through the uncertainty of time domain prediction; an adversarial time-frequency modulation decoupling mechanism is adopted, and by constructing a dual-path adversarial learning framework, the time-frequency correlation features are explicitly decoupled into two functionally complementary components: one is a domain-invariant correlation component, used to promote cross-domain knowledge transfer; the other is a domain-specific correlation component, used to retain the discriminative information unique to the source and target domains. During training, standard adversarial training is used for the domain-invariant component to strengthen cross-domain consistency, while reverse adversarial training is used for the domain-specific component to enhance its domain discrimination ability.

4. The multi-view cross-domain heart sound segmentation method as described in claim 3, characterized in that, The total loss function L of the multi-view cross-domain adaptive heart sound segmentation model total for: in, For the standard cross-entropy loss of the time-domain branch, Standard cross-entropy loss of frequency domain branch, L s For source domain distillation loss, L t For the target domain distillation loss, L D To counteract the decoupling loss, β1, β2, and β3 are all hyperparameters used to balance the importance of each loss term in the overall optimization objective.

5. A multi-view cross-domain heart sound segmentation device, characterized in that, include: The first processing unit is used to acquire heart sound data and perform preprocessing. The second processing unit is used to construct a multi-view cross-domain adaptive heart sound segmentation model based on the preprocessed heart sound data. The third processing unit is used to input the heart sound data to be processed into the multi-view cross-domain adaptive heart sound segmentation model for heart sound segmentation.

6. The multi-view cross-domain heart sound segmentation device as described in claim 5, characterized in that, The first processing unit includes: The first processing component is used to denoise the source and target domain heart sound datasets; The second processing component is used to extract time-domain and frequency-domain features from the denoised source and target domain heart sound data.

7. The multi-view cross-domain heart sound segmentation device as described in claim 6, characterized in that, In the training phase of the multi-view cross-domain adaptive heart sound segmentation model, a parallel dual-UNet architecture is used as the backbone network, which includes two branches that process time-domain signals and frequency-domain signals respectively, with the frequency-domain branch serving as an auxiliary path. An improved wavelet gMLP module is embedded in the central bottleneck position of the time-domain and frequency-domain branches respectively, that is, between the encoder and the decoder. The improved wavelet gMLP module introduces a Morlet wavelet driving mechanism in the spatial gating path. Based on the uncertainty estimation module-guided time-frequency domain mutual learning, in the source domain, the highly discriminative frequency domain branch serves as the teacher to guide the time domain branch learning; in the target domain, the intensity of knowledge transfer is dynamically adjusted through the uncertainty of time domain prediction; an adversarial time-frequency modulation decoupling mechanism is adopted, and by constructing a dual-path adversarial learning framework, the time-frequency correlation features are explicitly decoupled into two functionally complementary components: one is a domain-invariant correlation component, used to promote cross-domain knowledge transfer; the other is a domain-specific correlation component, used to retain the discriminative information unique to the source and target domains. During training, standard adversarial training is used for the domain-invariant component to strengthen cross-domain consistency, while reverse adversarial training is used for the domain-specific component to enhance its domain discrimination ability.

8. The multi-view cross-domain heart sound segmentation device as described in claim 7, characterized in that, The total loss function L of the multi-view cross-domain adaptive heart sound segmentation model total for: in, For the standard cross-entropy loss of the time-domain branch, Standard cross-entropy loss of frequency domain branch, L s For source domain distillation loss, L t For the target domain distillation loss, L D To counteract the decoupling loss, β1, β2, and β3 are all hyperparameters used to balance the importance of each loss term in the overall optimization objective.

9. A multi-view cross-domain heart sound segmentation system, characterized in that, include: A memory and a processor, wherein the memory stores a computer program executed by the processor, the computer program performing the multi-view cross-domain heart sound segmentation method as described in any one of claims 1-4 when executed by the processor.

10. A storage medium, characterized in that, The storage medium stores a computer program that, when executed, performs the multi-view cross-domain heart sound segmentation method as described in any one of claims 1-4.